1/57
Comprehensive vocabulary flashcards covering key GenAI concepts and AWS services for the AIP-C01 study guide.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Foundation model (FM)
A large model pre-trained on broad data that you adapt to many tasks through prompting or light customization.
Large language model (LLM)
A foundation model specialized for text. It predicts the next token to generate language.
Token
A chunk of text a model processes, roughly a word piece. Cost and limits are measured in tokens.
Context window
The maximum number of tokens a model can read and generate in one request. Overflowing it truncates content.
Prompt
The input text and instructions you send to a model to shape its output.
Prompt engineering
The practice of writing and refining prompts to get reliable, accurate outputs.
System prompt
A high-level instruction that sets the model role and rules for a conversation.
Chain-of-thought
A prompting pattern that asks the model to reason step by step before answering.
Temperature
A setting that controls randomness. Low values give focused, repeatable output; high values give varied output.
Top-p (nucleus sampling)
A setting that limits token choices to the smallest set whose probabilities add up to p. Lower p is more focused.
Top-k
A setting that limits token choices to the k most likely tokens at each step.
Embedding
A numeric vector that represents the meaning of text or other data. Similar meanings sit close together.
Vector store
A database that indexes embeddings and finds the nearest vectors to a query. The retrieval half of RAG.
Retrieval Augmented Generation (RAG)
A pattern that retrieves relevant documents and adds them to the prompt so the model answers from your data.
Chunking
Splitting documents into smaller pieces before embedding, so retrieval returns focused, relevant context.
Semantic search
Search by meaning using embeddings, rather than exact keyword matching.
Hybrid search
A search that combines keyword matching and vector similarity to improve relevance.
Reranking
A second-pass model that reorders retrieved results by relevance before they reach the main model.
Grounding
Tying a model answer to retrieved source data so it stays factual and can cite sources.
Hallucination
A confident model output that is false or unsupported by the source data.
Fine-tuning
Further training a model on your own labelled data to specialize it for a task or domain.
LoRA (low-rank adaptation)
A parameter-efficient fine-tuning method that trains small adapter weights instead of the whole model.
RLHF (reinforcement learning from human feedback)
Training that uses human preference ratings to align a model with what people want.
Agent
A system where a model plans steps and calls tools or APIs to complete a task, going beyond plain text output.
Tool use / function calling
A model ability to invoke defined functions or APIs with structured arguments to act in the world.
Model Context Protocol (MCP)
An open standard that defines how an agent connects to external tools and data sources.
Multi-agent system
Several agents that cooperate, each handling part of a task. Built on AWS with Strands Agents or Agent Squad.
Guardrails
Configurable filters that block harmful content, denied topics, and PII on model inputs and outputs.
Prompt injection
An attack where crafted input tricks a model into ignoring its instructions or leaking data.
Jailbreak
A prompt that bypasses a model safety rules to force disallowed output.
Model distillation
Training a smaller model to mimic a larger one, cutting cost and latency while keeping much of the quality.
Inference profile
A Bedrock configuration that routes model calls, including across regions, for capacity and resilience.
Cross-Region inference
Serving requests from more than one region to improve availability and capacity.
Provisioned throughput
Reserved, steady model capacity in Bedrock for predictable, high-volume workloads. Contrast with on-demand.
On-demand inference
Pay-per-request model calls with no reserved capacity. Good for spiky or low-volume traffic.
Model cascading
Sending easy queries to a cheap, small model and escalating hard ones to a larger model to save cost.
Prompt caching
Reusing model work for repeated prompt prefixes to cut latency and cost.
Semantic caching
Returning a stored answer when a new query is close in meaning to a previous one.
LLM-as-a-judge
Using a capable model to score another model outputs against criteria, for automated evaluation.
Drift
A gradual change in inputs or model behaviour over time that degrades quality. Monitored to catch regressions.
Model card
A document that records a model purpose, data, limits, and risks for governance and compliance.
Data lineage
A record of where data came from and how it moved and changed, used for audit and traceability.
Amazon Bedrock
Managed service for calling foundation models through one API. The core of most exam scenarios. Supports on-demand and provisioned throughput.
Amazon Bedrock Knowledge Bases
Managed RAG service that ingests documents, chunks them, creates embeddings, stores vectors, and returns grounded answers with citations.
Amazon Bedrock Guardrails
Policy layer that filters harmful content, blocks denied topics, and redacts PII on inputs and outputs.
Amazon Bedrock Prompt Management
Stores, versions, and parameterizes prompt templates with approval workflows for prompt governance.
Amazon Bedrock Prompt Flows
Visual, low-code way to chain prompts with branching and pre and post processing.
Amazon Bedrock Model Evaluations
Built-in way to score models on quality, including automated, LLM-as-a-judge, and human evaluation.
Amazon Titan
AWS family of foundation models, including Titan Text Embeddings for RAG.
Amazon SageMaker AI
Build, train, and host custom or fine-tuned models on managed endpoints when Bedrock does not host the model you need.
SageMaker Model Registry
Versions and stages models for deployment and supports rollback of customized models.
SageMaker Clarify
Detects bias and explains model predictions for fairness and Responsible AI.
SageMaker Model Monitor
Watches deployed models for data and quality drift in production.
Amazon OpenSearch Service
Search and vector database supporting k-NN vector search and hybrid search.
Amazon Aurora with pgvector
PostgreSQL-compatible database with the pgvector extension for embeddings, used when vectors live beside relational data.
Amazon Comprehend
NLP service that extracts entities and detects PII and sentiment.
Amazon Macie
Discovers and classifies sensitive data such as PII in Amazon S3.
Strands Agents and AWS Agent Squad
AWS-native frameworks for building single and multi-agent systems.