E077 Huyen Ch1 — Practice Test

0.0(0)
Studied by 0 people
call kaiCall Kai
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/25

encourage image

There's no tags or description

Looks like no tags are added yet.

Last updated 6:20 AM on 10/8/26
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

26 Terms

1
New cards
Q1. According to Huyen, what is AI engineering? A) Training new foundation models from scratch B) Operating GPU clusters for model providers C) Labeling data for supervised learning D) Building applications on top of readily available foundation models
D) Building applications on top of readily available foundation models. Huyen defines AI engineering as building applications on top of foundation models; traditional ML engineering develops models, AI engineering leverages existing ones. (p. 1, 12)
2
New cards
Q2. Why could language models scale into LLMs when many other model types could not? A) They are unsupervised and need no data B) They need no compute C) They can be trained with self-supervision, avoiding the data labeling bottleneck D) They use only labeled ImageNet data
C) They can be trained with self-supervision, avoiding the data labeling bottleneck. Each text sequence supplies its own labels (the next tokens), so massive training sets need no manual labeling. (p. 6-8)
3
New cards
Q3. Which statement correctly distinguishes self-supervised from unsupervised learning? A) They are identical B) Self-supervised learning needs human labels; unsupervised does not C) In self-supervised learning labels are inferred from the input data; unsupervised learning needs no labels at all D) Unsupervised learning infers labels from images only
C) In self-supervised learning labels are inferred from the input data; unsupervised learning needs no labels at all. Huyen's note: self-supervision infers labels from the input; unsupervised learning doesn't use labels. (p. 8)
4
New cards
Q4. Per Table 1-1, how many training samples does the sentence "I love street food." produce for language modeling? A) Four B) Five C) Six D) Seven
C) Six. Contexts from through "food ." each predict one next token, ending with : six samples. (p. 7)
5
New cards
Q5. A model is trained to fill in "My favorite __ is blue" using context on both sides of the blank. What type of language model is it? A) Diffusion B) Causal C) Autoregressive D) Masked
D) Masked. Masked language models (e.g., BERT) predict missing tokens using context before and after; autoregressive (causal) models use only preceding tokens. (p. 4)
6
New cards
Q6. For GPT-4, roughly how many words do 100 tokens represent? A) About 25 B) About 50 C) About 133 D) About 75
D) About 75. An average GPT-4 token is about 3/4 the length of a word. (p. 3)
7
New cards
Q7. Which is NOT one of Huyen's reasons for using tokens instead of words or characters? A) Fewer unique tokens than words shrinks the vocabulary B) Tokens help process unknown or made-up words C) Tokens guarantee the model's output is correct D) Tokens break words into meaningful components
C) Tokens guarantee the model's output is correct. The three reasons are meaningful sub-word parts, smaller vocabulary/efficiency, and handling unknown words. Outputs are probabilistic, never guaranteed correct. (p. 4-5)
8
New cards
Q8. Huyen uses the term "foundation models" to refer to: A) Only embedding models like CLIP B) Only text-only LLMs C) Only task-specific classifiers D) Both large language models and large multimodal models
D) Both large language models and large multimodal models. She explicitly uses foundation models for both LLMs and LMMs; "foundation" reflects their importance and that they can be built upon. (p. 9-10)
9
New cards
Q9. What is true of CLIP? A) It is a generative text model trained on labeled ImageNet B) It is an embedding model trained with natural language supervision on 400 million (image, text) pairs C) It was trained on 1,000 hand-labeled categories D) It is an autoregressive LMM
B) It is an embedding model trained with natural language supervision on 400 million (image, text) pairs. CLIP used co-occurring (image, text) pairs from the internet (400x ImageNet) and produces joint embeddings; it isn't generative. (p. 10)
10
New cards
Q10. A team writes detailed instructions plus examples into the prompt, without changing the model. This is: A) Prompt engineering B) Pre-training C) Finetuning D) Quantization
A) Prompt engineering. Prompt engineering adapts a model through instructions and context, with no weight updates. (p. 11, 40)
11
New cards
Q11. Supplementing a model's instructions with information retrieved from a database (e.g., customer reviews) is called: A) Finetuning B) Retrieval-augmented generation (RAG) C) Post-training D) Dataset engineering
B) Retrieval-augmented generation (RAG). Using a database to supplement the instructions is RAG. (p. 11)
12
New cards
Q12. Which is one of the three factors Huyen says created ideal conditions for AI engineering's growth? A) Low entrance barrier to building AI applications B) Stricter AI regulation C) Cheaper data labeling D) The end of GPU shortages
A) Low entrance barrier to building AI applications. The three factors: general-purpose AI capabilities, increased AI investments, and a low entrance barrier (model as a service, minimal coding). (p. 12-14)
13
New cards
Q13. In Eloundou et al. (2023), a task counts as "exposed" to AI when AI can: A) Reduce the time needed to complete it by at least 50% B) Perform it with 100% accuracy C) Be legally used for it D) Fully replace the worker
A) Reduce the time needed to complete it by at least 50%. Exposure is defined as at least a 50% reduction in task time. (p. 17)
14
New cards
Q14. Why are enterprises faster to deploy internal-facing AI apps (e.g., knowledge management) than external-facing ones? A) External apps are illegal B) They build expertise while limiting data privacy, compliance, and catastrophic-failure risk C) Internal apps need no evaluation D) Internal apps always use larger models
B) They build expertise while limiting data privacy, compliance, and catastrophic-failure risk. The a16z data shows enterprises prefer lower-risk applications first. (p. 19)
15
New cards
Q15. Which business reason for building an AI application does Huyen rank as the highest risk? A) Competitors with AI could make you obsolete B) Curiosity about the technology C) Missing opportunities to boost profits D) Wanting to upskill the team
A) Competitors with AI could make you obsolete. An existential threat (business continuity) ranks highest; missing profit opportunities is second; not wanting to be left behind is third. (p. 29)
16
New cards
Q16. Gmail would still work without Smart Compose, but Face ID would not work without facial recognition. Smart Compose is therefore: A) Dynamic AI B) Critical AI C) Proactive AI D) Complementary AI
D) Complementary AI. If the app works without the AI, the AI is complementary. The more critical AI is, the more accurate and reliable it must be. (p. 30)
17
New cards
Q17. Why do proactive AI features typically need a higher quality bar than reactive ones? A) Users didn't ask for them, so low-quality output feels intrusive B) They must be real-time C) They run on larger models D) They are always critical to the app
A) Users didn't ask for them, so low-quality output feels intrusive. Proactive features can be precomputed (latency matters less), but unrequested low-quality output annoys users. (p. 30)
18
New cards
Q18. A support team uses AI only to suggest replies that human agents must review. In Microsoft's Crawl-Walk-Run framework this is: A) Crawl B) Walk C) Run D) Sprint
A) Crawl. Crawl means human involvement is mandatory; Walk means AI interacts with internal employees; Run means more automation, possibly with external users. (p. 31)
19
New cards
Q19. According to Huyen, which competitive advantage is most realistic for a startup building on foundation models? A) None; startups cannot compete B) Data, by getting to market first and using usage data to improve C) Technology, since core tech is unique D) Distribution, since startups reach users fastest
B) Data, by getting to market first and using usage data to improve. Core technology converges and big companies usually own distribution; usage data can become a startup's moat. (p. 32)
20
New cards
Q20. TTFT and TPOT are examples of which group of usefulness-threshold metrics? A) Fairness B) Latency C) Cost D) Quality
B) Latency. Latency metrics include TTFT (time to first token), TPOT (time per output token), and total latency. (p. 33)
21
New cards
Q21. LinkedIn reached 80% of its desired experience in one month but needed four more months to pass 95%. This illustrates: A) The data flywheel B) Model as a service C) The last mile challenge D) Crawl-Walk-Run
C) The last mile challenge. Initial demos mislead; going from 60 to 100 is far harder than 0 to 60. (p. 33-34)
22
New cards
Q22. Which correctly orders the three layers of the AI stack from top to bottom? A) Infrastructure, model development, application development B) Application development, model development, infrastructure C) Model development, application development, infrastructure D) Application development, infrastructure, model development
B) Application development, model development, infrastructure. Builders usually start at the top layer (application development) and move down as needed. (p. 37)
23
New cards
Q23. Which is one of the three major ways AI engineering differs from traditional ML engineering? A) It uses smaller models with lower latency B) It focuses mainly on feature engineering of tabular data C) It never needs GPUs D) Its open-ended outputs make evaluation a much bigger problem
D) Its open-ended outputs make evaluation a much bigger problem. The three differences: model adaptation over training, bigger and more compute-hungry models, and open-ended outputs that are harder to evaluate. (p. 39-40)
24
New cards
Q24. An engineer reduces the numerical precision of a model's weights so it runs faster. Which statement is correct? A) This is pre-training B) This is prompt engineering C) This is finetuning because the weights change D) This is quantization, an inference optimization that changes weights but isn't training
D) This is quantization, an inference optimization that changes weights but isn't training. Training always changes weights, but not every weight change is training; quantization is an example. (p. 41, 44)
25
New cards
Q25. Gemini Ultra scored 90.04% on MMLU with CoT@32 but 83.7% with 5-shot, while GPT-4 scored 86.4% with 5-shot. What lesson does Huyen draw? A) Gemini Ultra is clearly better in every setting B) MMLU is a useless benchmark C) Prompt technique can change results so much that comparisons need the same evaluation setup D) 5-shot prompting always beats CoT
C) Prompt technique can change results so much that comparisons need the same evaluation setup. Table 1-5 shows prompting technique changes rankings, which is part of why evaluation is harder with foundation models. (p. 44-45)
26
New cards

Q26. Per Table 1-6, how has the importance of prompt engineering changed from traditional ML to foundation models? A) Important → less important B) Not applicable → important C) Less important → not applicable D) Unchanged

B) Not applicable → important. Table 1-6: AI interface less important → important; prompt engineering not applicable → important; evaluation important → more important. (p. 46)