📧 hipravat@gmail.com Support

AWS AI Practitioner Certification Prep

AI Introduction

Artificial Intelligence (AI) is a branch of computer science dedicated to solving problems that typically require human intelligence, such as image recognition, speech understanding, content generation, and decision-making. At its essence, AI utilizes large datasets to train computational models—sophisticated algorithms with statistical capabilities—that learn patterns and make predictions or classifications based on new, unseen data. For instance, an AI model trained on thousands of images can accurately identify objects like fruits or convert spoken language into text.

The evolution of AI has been remarkable, beginning in the 1950s with foundational concepts like the Turing Test proposed by Alan Turing and the formal establishment of the field by John McCarthy. It has progressed through rule-based expert systems to the modern innovations of machine learning and deep learning techniques. Today, AI is seamlessly integrated into our daily lives, driving applications such as autonomous vehicles, healthcare diagnostics, fraud detection, intelligent document processing, and generative tools like ChatGPT.

In summary, AI represents a broad domain that encompasses machine learning, deep learning, and generative AI, all of which contribute to the development of systems that learn, adapt, and assist humans in increasingly complex tasks.

Generative Artificial Intelligence (Gen AI)

Generative Artificial Intelligence (Gen AI) is a sophisticated subset of artificial intelligence that resides within the realms of deep learning and machine learning. It is specifically designed to produce new data—including text, images, audio, code, and video—that closely mirrors the data on which it was trained. In contrast to traditional AI systems, which primarily focus on analysis and classification, Gen AI models identify patterns from vast datasets and leverage this knowledge to generate entirely new and unique outputs.

Typically, these models are built using large-scale foundation models, trained on extensive and diverse datasets, enabling them to perform a wide array of tasks such as content generation, summarization, translation, and question answering. A prominent example of this technology is ChatGPT, which utilizes large language models (LLMs) to generate human-like text based on user prompts.

A noteworthy characteristic of Gen AI is its non-deterministic nature, meaning that identical inputs can yield slightly different outputs on each occasion due to probabilistic decision-making. Furthermore, Gen AI encompasses more than just text generation; it also includes image generation techniques like diffusion models, where systems learn to transform noise into coherent visuals based on user-provided prompts.

Overall, Generative AI signifies a remarkable advancement in AI capabilities, empowering machines not only to comprehend data but also to creatively generate novel content and solutions across various domains.

Amazon Bedrock

Introduction

Amazon Bedrock is a fully managed service provided by Amazon Web Services that allows developers to build and scale generative AI applications without the burden of managing the underlying infrastructure. It offers a unified interface and API for accessing a diverse array of foundation models from leading AI companies, including Anthropic, Meta, Cohere, and others. This enables users to experiment, configure, and deploy models with ease.

One of the fundamental advantages of Amazon Bedrock is its commitment to data privacy and security; all user data remains within their AWS account and is not shared with model providers. The service features a pay-as-you-go pricing model and incorporates advanced functionalities such as Retrieval-Augmented Generation (RAG), fine-tuning with custom datasets, and the creation of intelligent agents. Furthermore, Bedrock facilitates integration by providing a standardized API across all models, allowing organizations to develop applications like chatbots, content generation systems, and knowledge-based assistants, all while upholding strong governance, compliance, and responsible AI practices.

Amazon Bedrock – Console Steps

Amazon Bedrock offers an intuitive console experience that enables users to effortlessly explore, test, and interact with a diverse array of foundation models directly within the AWS environment. Upon accessing the Bedrock console, users can navigate to the Model Catalog to browse available models by provider—such as Amazon or Anthropic—and filter them based on capabilities like text or image generation. Each model includes comprehensive information regarding its features, pricing, and supported use cases.

To experiment with these models, users can utilize the Playground, where they can submit prompts and compare responses across different models through a unified interface. For instance, a straightforward query such as "What is AWS?" may generate varying outputs depending on the chosen model, thereby showcasing differences in formatting, response style, token usage, and latency.

Furthermore, Bedrock supports multimodal capabilities, enabling users to upload files for analysis or create images using text prompts with tools like image playgrounds. This hands-on approach illustrates how easily developers can evaluate and select the most suitable models for their applications, positioning Amazon Bedrock as a powerful and accessible platform for building generative AI solutions.

Prompt Engineering

Introduction

Prompt engineering involves carefully crafting inputs (prompts) for large language models (LLMs) to achieve more accurate and relevant outputs. Simple prompts, like asking an AI to "summarize what AWS is," may produce vague responses. To improve results, prompt engineering uses four key components: clear instructions, context, input data, and an output indicator. This structured approach enhances AI-generated responses, especially with tools like ChatGPT.

Additionally, negative prompting directs the model on what to avoid, helping maintain focus and clarity by limiting irrelevant details or jargon when addressing specific audiences. Ultimately, prompt engineering is crucial for effectively using generative AI, allowing users to enhance response quality and tailor content to real-world needs.

Prompt Performance Optimization

To enhance the performance of prompts within a generative AI model, it is essential to grasp how Large Language Models (LLMs) produce text and the ways in which you can influence this process through configuration and prompt design. LLMs generate the next word based on probabilities, allowing for guided behavior by fine-tuning specific parameters and structuring prompts effectively. Platforms such as Amazon Bedrock offer various controls to optimize output quality, consistency, and creativity.

Key Parameters:

  • System Prompts – Define the model's role and tone (e.g., "respond like an AWS cloud expert"), which greatly enhances relevance and style.
  • Temperature – Governs creativity. Lower values (e.g., 0.2) yield more predictable and focused responses, while higher values (e.g., 1.0) foster creativity, though they may compromise coherence.
  • Top P (Nucleus Sampling) – Constrains word selection based on cumulative probability. Lower values ensure more precise responses, while higher values promote greater diversity.
  • Top K – Limits the number of candidate words considered. Smaller values enhance consistency, while larger values boost variation.
  • Maximum Length – Determines the response's length.
  • Stop Sequences – Dictate when the model should cease generating text.

Latency Factors:

Prompt performance can be influenced by model size, model type, and the number of input/output tokens—all of which directly affect latency. Generally, larger models and longer prompts result in increased latency, while parameters like temperature, Top P, and Top K do not significantly impact response speed.

In practice, improving prompt performance involves striking a balance between clarity, control, and experimentation—combining well-structured prompts with appropriate parameter tuning to achieve the desired levels of accuracy, creativity, and efficiency.

Prompt Engineering Techniques

To improve prompt effectiveness in generative AI systems, several advanced prompt engineering techniques can be applied:

Zero-Shot Prompting – The simplest approach, where you provide a task without any examples—such as asking the model to "write a short story about a dog solving a mystery." The model relies entirely on its pre-trained knowledge.

Few-Shot Prompting – Introduces sample examples before the actual request. By showing the model how similar tasks should be handled, you guide its style, tone, and structure, resulting in more predictable and aligned outputs. Even providing a single example (one-shot prompting) can significantly improve results.

Chain-of-Thought Prompting – Explicitly instructs the model to think step by step. This is especially useful for complex reasoning tasks, as it encourages structured and logical responses by breaking the problem into sequential steps.

Retrieval-Augmented Generation (RAG) – Enhances prompts by incorporating external data sources. Instead of relying solely on the model's internal knowledge, RAG retrieves relevant information (such as documents or facts) and injects it into the prompt, leading to more accurate, context-aware, and up-to-date responses.

In practice, these techniques can be combined—using structured instructions, examples, step-by-step reasoning, and external data—to significantly enhance the quality, reliability, and relevance of AI-generated outputs.

Prompt Templates

Prompt templates are a powerful way to standardize and simplify how prompts are created and used in generative AI systems. Instead of writing a new prompt from scratch each time, you define a reusable structure with placeholders (e.g., for text, questions, or options) that users can fill in. This ensures consistency in how inputs are formatted and how outputs are generated, which is especially useful when working with platforms like Amazon Bedrock.

Prompt templates can also embed advanced techniques like few-shot prompting, detailed instructions, and formatting rules—all hidden from the end user—making them highly flexible and scalable.

Prompt Injection Attacks: One important challenge is prompt injection attacks, where malicious inputs attempt to override the original intent of the template. To mitigate this, templates should include safeguards—such as explicit instructions for the model to ignore irrelevant or harmful content and strictly follow the intended context.

Overall, prompt templates enhance reliability, security, and efficiency, making them essential for building robust and production-ready generative AI applications.

Amazon Q

Q Business

Amazon Q Business is a fully managed generative AI assistant designed specifically for enterprise use, enabling employees to interact with their organization's internal knowledge in a secure and efficient way. Built on top of Amazon Bedrock, it leverages multiple foundation models to provide capabilities such as answering questions, summarizing documents, generating content, and automating routine tasks like creating tickets or scheduling meetings.

A key strength of Amazon Q Business lies in its ability to connect to a wide range of enterprise data sources through built-in connectors, including services like Amazon S3, Amazon RDS, and external platforms such as Microsoft 365, Google Drive, and Slack. It uses a retrieval-augmented generation (RAG) approach to fetch relevant information from these sources and provide accurate, source-backed answers. Additionally, plugins enable the assistant to take actions in third-party systems—such as creating issues in Jira or managing workflows in ServiceNow.

Security and governance are central to Amazon Q Business. Users authenticate through IAM Identity Center, ensuring that responses are restricted based on their access permissions. Administrative controls (guardrails) allow organizations to enforce policies, restrict topics, and ensure that responses align with internal guidelines.

Q Apps

Amazon Q Apps is a feature within Amazon Q Business that enables users to create generative AI–powered applications using natural language—without writing any code. Through a web-based interface called the Q Apps Creator, users simply describe the type of application they want (for example, a document analyzer, content generator, or internal knowledge assistant), and the system automatically generates a functional web app tailored to that request.

A key advantage of Amazon Q Apps is its accessibility—it empowers non-developers across the organization to build useful tools quickly, such as apps that allow users to upload documents, run prompts, generate summaries, or extract insights. Additionally, these apps can integrate with enterprise plugins and data sources, enabling them to not only retrieve information but also perform actions within business systems.

Q Developer

Amazon Q Developer is an AI-powered assistant from Amazon Web Services designed to help developers and cloud engineers work more efficiently across AWS environments. It combines two major capabilities:

Operational Assistance – Interprets natural language requests and generates accurate AWS CLI commands (e.g., updating an AWS Lambda timeout). It can also analyze account data, identify highest-cost services, and troubleshoot errors.

AI-Driven Coding Support – Acts as an intelligent coding companion (similar to GitHub Copilot). It can generate code in multiple languages (Python, Java, JavaScript), provide real-time suggestions, assist with debugging, and scan code for security vulnerabilities. It supports popular IDEs like VS Code and JetBrains tools.

Additionally, it can help bootstrap new projects, generate documentation, and optimize code performance, reducing manual effort and simplifying AWS-specific development tasks.

Q for AWS Services

Amazon Q is increasingly being embedded as an intelligent layer across multiple AWS services. A great example is Amazon Q for QuickSight, which brings natural language capabilities into Amazon QuickSight—a tool traditionally used for building dashboards and visualizing data.

With this integration, users can simply ask questions in natural language—such as "show sales by city and product"—and Amazon Q will automatically interpret the request, analyze the dataset, and generate the appropriate visualizations (like maps, charts, or graphs). This significantly lowers the barrier for non-technical users, enabling faster insights and more intuitive data exploration.

PartyRock

PartyRock is a lightweight, no-code playground that lets anyone experiment with building generative AI applications without needing an AWS account. Powered behind the scenes by Amazon Bedrock, it provides a simple web interface where users can create apps using natural language prompts and prebuilt widgets.

The experience is like Amazon Q Apps, but much simpler. You define inputs (like location, cuisine, or ingredients), and the system automatically generates outputs such as text, images, or recommendations. Under the hood, it uses prompt templates and connects widgets together—for example, one widget might generate a restaurant recommendation, and another might expand on it with more details.

Overall, PartyRock is a great hands-on sandbox to explore how prompts, models, and workflows come together in generative AI, especially helpful for quickly testing ideas before moving into production-ready tools in AWS.

Artificial Intelligence (AI) & Machine Learning (ML)

Overview of AI, ML, Deep Learning, Gen AI

Artificial Intelligence (AI) is an expansive field creating intelligent systems that perform tasks associated with human intelligence—perception, reasoning, learning, problem-solving, and decision-making. AI encompasses a hierarchy: AI → Machine Learning (ML) → Deep Learning → Generative AI.

Four Fundamental Layers of AI Systems: 1. Data Layer – Collection of vast amounts of information 2. Algorithm Layer – Frameworks are defined 3. Model Layer – Dedicated to training 4. Application Layer – Delivers the model to users

Machine Learning enables machines to identify patterns from data rather than being explicitly programmed. Techniques include regression (trend prediction) and classification (data categorization).

Deep Learning employs neural networks inspired by the human brain. It is termed "deep" due to multiple hidden layers between input and output, enabling processing of intricate patterns. Requires large datasets and powerful GPUs.

Generative AI utilizes foundation models based on the transformer architecture, processing entire sentences efficiently. Powers models like BERT and ChatGPT, and includes diffusion models for image generation and multi-modal models for various input/output types.

Key Machine Learning Terms

Term Purpose Use Case
GPT (Generative Pre-trained Transformer) Generates human-like text or code Text generation, coding
BERT (Bidirectional Encoder Representations) Processes text bidirectionally Translation tasks
RNN (Recurrent Neural Network) Handles sequential data Speech recognition, time series
ResNet (Residual Network) Deep CNN for visual tasks Image recognition, facial recognition
SVM (Support Vector Machine) Classification and regression Data categorization
WaveNet Generates audio waveforms Speech synthesis
GAN (Generative Adversarial Network) Produces synthetic data Data augmentation
XGBoost (Extreme Gradient Boosting) Regression tasks Predictions

Exam Recall: GPT & BERT → language, ResNet → images, WaveNet → audio, GAN → data augmentation.

Training Data

In machine learning, data quality is paramount—if you input poor-quality data, the output will similarly be poor.

Labeled vs. Unlabeled Data: - Labeled data contains input features and corresponding output labels (e.g., images tagged as "dog" or "cat"). Used for supervised learning. - Unlabeled data consists solely of input features without labels. The algorithm identifies patterns independently through unsupervised learning.

Structured vs. Unstructured Data: - Structured data is organized in rows and columns (spreadsheets, databases, time series). - Unstructured data lacks specific structure (text, social media posts, images, audio).

Both data types are valid for ML but require different algorithms for effective processing.

Supervised Learning

Supervised learning uses labeled data to establish a mapping function that predicts outcomes for new, unseen inputs.

Regression vs. Classification: - Regression predicts continuous numeric values (house prices, stock values, weather) - Classification predicts discrete categorical labels—can be binary (spam/not spam), multi-class (mammal/bird/reptile), or multi-label (action + comedy movie)

Data Splitting: - Training set (60-80%) – for model training - Validation set (10-20%) – for tuning parameters - Test set (10-20%) – for evaluating final accuracy

Feature Engineering: Transforming raw data into meaningful features: - Feature extraction (deriving age from birth date) - Feature selection (choosing relevant features) - Feature transformation (normalizing values for faster convergence)

Unsupervised Learning

Unsupervised Learning utilizes unlabeled data to uncover inherent patterns, structures, or relationships.

Key Techniques: - Clustering – Organizes similar data points into groups (e.g., customer segmentation: students, new parents, vegetarians) - Association Rule Learning – Discovers items frequently purchased together (bread and butter) - Anomaly Detection – Identifies outliers deviating from normal patterns (fraud detection)

Semi-Supervised Learning: Bridges supervised and unsupervised approaches. Starts with small labeled data + larger unlabeled pool. Trains initial model → assigns pseudo-labels to unlabeled data → retrains on expanded dataset.

Self-Supervised Learning

Self-supervised learning is a methodology where a model creates its own pseudo-labels from unlabeled data, eliminating the need for human annotation.

How it works: Uses pretext tasks—simple challenges the model solves to uncover patterns: - Predicting the next word ("Amazon Web ___" → "Services") - Completing sentences - Reconstructing hidden elements from visible ones

By engaging in numerous pretext tasks, the model develops an internal representation of data—understanding grammar, word meanings, and relationships. This approach powers groundbreaking models like GPT and sophisticated image recognition systems.

Reinforcement Learning

Reinforcement learning is a subset of ML where an agent learns to make decisions by taking actions within an environment to maximize cumulative rewards.

Key Concepts: - State – Current situation within the environment - Policy – Strategy for selecting actions based on state - Reward – Feedback from the environment - Agent – The learner/decision maker

Learning Cycle: Observe state → Choose action → Environment transitions → Receive reward → Update policy → Repeat

Use Cases: Gaming (chess, Go), robotics (navigation), finance (trading strategies), healthcare (treatment optimization), autonomous vehicles (path planning).

Reinforcement Learning from Human Feedback (RLHF)

RLHF is a technique used to fine-tune large language models by incorporating human preferences into the training process. It bridges the gap between what a model can generate and what humans actually find helpful, harmless, and honest.

How RLHF Works:

  1. Initial Training – A base LLM is pre-trained on large text corpora using self-supervised learning
  2. Human Ranking – Human evaluators rank multiple model outputs for the same prompt based on quality, helpfulness, and safety
  3. Reward Model Training – A separate reward model is trained on these human preferences to predict which outputs humans would prefer
  4. Policy Optimization – The original LLM is fine-tuned using reinforcement learning (typically PPO – Proximal Policy Optimization) with the reward model providing the reward signal

Why RLHF Matters: - Aligns model outputs with human values and expectations - Reduces harmful, biased, or misleading responses - Improves helpfulness and instruction-following ability - Used by OpenAI (ChatGPT), Anthropic (Claude), and others to make models safer and more useful

Key Exam Points: RLHF combines supervised learning (human rankings) with reinforcement learning (policy optimization) to align AI behavior with human preferences.

Model Fit, Bias, and Variance

Understanding model fit is crucial for building effective ML models:

Underfitting (High Bias): - Model is too simple to capture patterns in the data - Performs poorly on both training and test data - Example: Using a straight line to fit curved data - Fix: Use more complex models, add features, reduce regularization

Overfitting (High Variance): - Model memorizes training data including noise - Performs well on training data but poorly on new/test data - Example: A model that perfectly fits every training point but fails on unseen data - Fix: Add more training data, use regularization, simplify model, use dropout

Good Fit: - Model generalizes well to unseen data - Balanced performance on both training and test sets

Bias-Variance Tradeoff: - Bias – Error from oversimplifying assumptions. High bias = underfitting - Variance – Error from sensitivity to training data fluctuations. High variance = overfitting - Goal: Find the sweet spot that minimizes total error (bias² + variance)

Model Evaluation Metrics

Classification Metrics:

Metric Formula Use Case
Accuracy (TP + TN) / Total Overall correctness (use when classes are balanced)
Precision TP / (TP + FP) When false positives are costly (spam detection)
Recall (Sensitivity) TP / (TP + FN) When false negatives are costly (disease detection)
F1 Score 2 × (Precision × Recall) / (Precision + Recall) Balance between precision and recall
AUC-ROC Area under ROC curve Model's ability to distinguish between classes

Regression Metrics: - MAE (Mean Absolute Error) – Average absolute difference between predicted and actual values - MSE (Mean Squared Error) – Average squared difference (penalizes large errors more) - RMSE (Root Mean Squared Error) – Square root of MSE (same units as target variable) - R² Score – Proportion of variance explained by the model (1.0 = perfect fit)

Confusion Matrix: A table showing True Positives, True Negatives, False Positives, and False Negatives—the foundation for classification metrics.

Machine Learning – Inferencing

Inferencing (or inference) is the process of using a trained ML model to make predictions on new, unseen data in production.

Key Concepts:

  • Batch Inference – Processing large volumes of data at once (e.g., scoring all customers overnight for churn risk). Cost-effective but not real-time.
  • Real-Time Inference – Processing individual requests as they arrive with low latency (e.g., fraud detection on each credit card transaction). Requires always-on endpoints.
  • Edge Inference – Running models on edge devices (IoT, mobile) for ultra-low latency without internet connectivity.

Inference Optimization: - Model compression (quantization, pruning, distillation) - Hardware acceleration (GPUs, AWS Inferentia chips) - Caching frequent predictions - Model serving frameworks (TensorFlow Serving, TorchServe)

AWS Inference Services: Amazon SageMaker endpoints, AWS Inferentia/Trainium chips, Amazon Elastic Inference.

Phases of a Machine Learning Project

  1. Business Problem Definition – Identify the problem, define success metrics, determine if ML is the right approach
  2. Data Collection & Integration – Gather data from various sources (databases, APIs, logs, third-party datasets)
  3. Data Preprocessing & Cleaning – Handle missing values, remove duplicates, fix inconsistencies, transform features
  4. Exploratory Data Analysis (EDA) – Visualize data, identify patterns, understand distributions, detect outliers
  5. Feature Engineering – Create new features, select relevant ones, transform variables for better model performance
  6. Model Selection – Choose appropriate algorithms based on problem type, data characteristics, and requirements
  7. Model Training – Train the model on prepared data, tune hyperparameters using validation set
  8. Model Evaluation – Assess performance using appropriate metrics, compare against baselines
  9. Model Deployment – Deploy to production (endpoint, batch, edge), set up monitoring
  10. Monitoring & Maintenance – Track model performance over time, detect data drift, retrain as needed

Hyperparameters

Hyperparameters are configuration settings that control the training process itself—they are set BEFORE training begins and are not learned from data.

Key Hyperparameters:

Hyperparameter Purpose Effect
Learning Rate Controls step size during optimization Too high = overshooting; too low = slow convergence
Batch Size Number of samples processed before updating weights Larger = stable but slower; smaller = noisy but faster
Number of Epochs How many times the model sees the entire dataset Too few = underfitting; too many = overfitting
Regularization (L1/L2) Prevents overfitting by penalizing large weights Higher = simpler model; lower = more complex
Number of Layers/Neurons Network architecture depth and width More = greater capacity but risk of overfitting
Dropout Rate Fraction of neurons randomly disabled during training Prevents overfitting in neural networks

Hyperparameter Tuning Methods: - Grid Search – Exhaustively tries all combinations (thorough but expensive) - Random Search – Randomly samples combinations (faster, often sufficient) - Bayesian Optimization – Intelligently explores promising regions (efficient) - AWS SageMaker Automatic Model Tuning – Managed hyperparameter optimization service

Parameters vs. Hyperparameters: - Parameters are learned during training (weights, biases) - Hyperparameters are set before training (learning rate, epochs)

AWS Managed AI Services

Introduction

AWS provides a comprehensive suite of pre-built, fully managed AI/ML services that enable developers to add intelligence to applications without requiring deep ML expertise. These services cover natural language processing, computer vision, speech, personalization, and document analysis—all accessible via simple API calls with pay-per-use pricing.

Amazon Comprehend

Amazon Comprehend is a natural language processing (NLP) service that uses machine learning to extract insights from text. It can identify the language, extract key phrases, entities (people, places, organizations), determine sentiment (positive, negative, neutral, mixed), and classify documents into custom categories.

Key Use Cases: Customer feedback analysis, content categorization, compliance document scanning, social media monitoring.

Key Feature: Comprehend Medical is a specialized version that extracts medical information (conditions, medications, dosages) from clinical text.

Amazon Translate

Amazon Translate is a neural machine translation service that provides fast, high-quality, and affordable language translation. It supports 75+ languages and automatically detects the source language.

Key Use Cases: Translating user-generated content, localizing applications, enabling multilingual communication, translating large volumes of documents.

Key Feature: Custom Terminology lets you define specific translations for brand names or industry terms that should not be translated generically.

Amazon Transcribe

Amazon Transcribe is an automatic speech recognition (ASR) service that converts speech to text. It supports real-time and batch transcription with features like speaker identification, custom vocabularies, and automatic punctuation.

Key Use Cases: Meeting transcription, call center analytics, subtitle generation, medical dictation (Transcribe Medical).

Key Features: Speaker diarization (identifying who said what), custom vocabulary for domain-specific terms, content redaction for PII.

Amazon Polly

Amazon Polly is a text-to-speech service that converts text into lifelike speech using deep learning. It supports multiple languages and offers both standard and neural (more natural-sounding) voices.

Key Use Cases: Voice-enabled applications, audiobook generation, announcement systems, accessibility tools.

Key Features: SSML (Speech Synthesis Markup Language) support for controlling pronunciation, emphasis, and pauses. Neural TTS for more human-like speech.

Amazon Rekognition

Amazon Rekognition is a computer vision service that can analyze images and videos to detect objects, scenes, faces, text, celebrities, and inappropriate content. It can also perform facial comparison and search.

Key Use Cases: Content moderation, identity verification, people counting, celebrity recognition, PPE detection, label detection.

Key Features: Custom Labels allows training custom models to detect domain-specific objects (e.g., specific product defects on a manufacturing line).

Amazon Lex

Amazon Lex is a service for building conversational interfaces (chatbots and voice bots) using the same technology that powers Amazon Alexa. It provides automatic speech recognition (ASR) and natural language understanding (NLU).

Key Use Cases: Customer service chatbots, IVR systems, virtual assistants, order status bots.

Key Features: Multi-turn conversations, slot filling, integration with AWS Lambda for fulfillment, deployment across multiple channels (web, mobile, messaging platforms).

Amazon Personalize

Amazon Personalize is a machine learning service that enables developers to create real-time personalized recommendations—the same technology used by Amazon.com. No ML expertise required.

Key Use Cases: Product recommendations, personalized search results, customized marketing emails, content curation.

Key Features: Real-time personalization, user segmentation, automatic retraining, support for multiple recommendation types (user personalization, similar items, personalized ranking).

Amazon Textract

Amazon Textract is a document analysis service that goes beyond simple OCR. It automatically extracts text, handwriting, tables, and form data from scanned documents while maintaining the structure and relationships between data.

Key Use Cases: Invoice processing, form extraction, ID document verification, mortgage processing, medical records digitization.

Key Features: Forms extraction (key-value pairs), table extraction, query-based extraction (ask specific questions about a document).

Amazon Kendra

Amazon Kendra is an intelligent enterprise search service powered by machine learning. Unlike keyword-based search, it understands natural language queries and returns precise answers extracted from documents.

Key Use Cases: Internal knowledge bases, IT help desks, customer support portals, research document search.

Key Features: FAQ matching, document ranking, incremental learning from user feedback, connectors for 40+ data sources (S3, SharePoint, Salesforce, databases).

Amazon Augmented AI (A2I)

Amazon Augmented AI (A2I) enables human review workflows for ML predictions. When a model's confidence is below a threshold, the prediction is routed to human reviewers for verification.

Key Use Cases: Content moderation review, document processing verification, medical image analysis validation, any ML prediction requiring human oversight.

Key Features: Integration with Amazon Textract and Rekognition, custom ML model support, built-in reviewer workforce (Amazon Mechanical Turk, private teams, or third-party vendors).

AWS HealthScribe

AWS HealthScribe is an AI-powered service that automatically generates clinical documentation from patient-clinician conversations. It uses speech recognition and generative AI to create structured clinical notes.

Key Use Cases: Automating medical note-taking, reducing physician documentation burden, creating structured clinical summaries.

Key Features: Speaker identification (doctor vs. patient), medical terminology recognition, structured output (chief complaint, history, assessment, plan), integration with EHR systems.

Amazon's Hardware for AI

AWS offers purpose-built chips optimized for AI/ML workloads:

Hardware Purpose Key Benefit
AWS Trainium Training ML models Up to 50% cost savings vs. GPU-based training
AWS Inferentia Running inference (predictions) Lowest cost per inference in the cloud
NVIDIA GPUs (P4d, P5 instances) General ML training & inference Highest performance for complex models
AWS Graviton General compute with ML capabilities Best price-performance for general workloads

Key Exam Points: - Trainium → Training models (SageMaker integration) - Inferentia → Running inference at scale (cost-optimized) - Use EC2 instances with these chips (Trn1 for Trainium, Inf2 for Inferentia)

Amazon SageMaker

Amazon SageMaker is a fully managed platform that enables developers and data scientists to build, train, and deploy machine learning models at scale. It removes the heavy lifting from each step of the ML workflow.

Key Components:

  • SageMaker Studio – Integrated IDE for ML development with notebooks, experiments, and debugging
  • SageMaker Canvas – No-code ML tool for business analysts to build models visually
  • SageMaker Autopilot – AutoML that automatically explores data, selects algorithms, and tunes hyperparameters
  • SageMaker Training – Managed training infrastructure with spot instance support for cost savings
  • SageMaker Endpoints – Real-time inference hosting with auto-scaling
  • SageMaker Pipelines – MLOps workflow automation (CI/CD for ML)
  • SageMaker Model Monitor – Detects data drift and model quality degradation in production
  • SageMaker Feature Store – Centralized repository for ML features
  • SageMaker Ground Truth – Data labeling service with human and automated workflows
  • SageMaker JumpStart – Pre-built ML solutions and foundation models ready to deploy

When to Use SageMaker vs. Bedrock: - Use SageMaker when you need full control over model training, custom algorithms, or traditional ML workloads - Use Bedrock when you want to use pre-built foundation models for generative AI without managing infrastructure