1/233
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
[T4] What is the main topic of Lecture 4?
Language models.
[T4] What three main topics are covered in Lecture 4?
N-gram language models (count-based), neural language models, and language model evaluation.
[T4] What shift in focus does Lecture 4 make relative to earlier material?
From discussing words mostly in isolation to discussing sequences of words.
[T4] What previous topic is reviewed at the beginning of Lecture 4?
Real-word spelling errors.
[T4] What is the first step for correcting real-word spelling errors?
Collect common confusion sets C={C1,…,Cn}.
[A4] What confusion-set examples are reviewed at the beginning of Lecture 4?
{their, they’re, there}, {to, too, two}, and {weather, whether}.
[T4] When is the real-word spelling-correction procedure triggered?
Whenever an encountered word c' belongs to one of the confusion sets Ci.
[T4] After encountering c' ∈ Ci, what is computed first?
The probability of the sentence containing c'.
[T4] What is done with the other words c ∈ Ci where c ≠ c'?
Substitute each into the sentence and compute the probability of the resulting sentence.
[T4] How is the final correction selected for a real-word spelling error?
Choose the candidate yielding the highest-probability sentence.
[A4] In the lecture example, what two sentences are compared?
“The whether forecast for today is sunny” and “The weather forecast for today is sunny.”
[T4] What is a probabilistic language model used to compute?
The probability of a sentence or sequence of words.
[T4] What other fundamental task can a language model perform?
Predict the next word given the preceding words.
[T4] What notation is used for a sequence of words from w1 through wn?
w1^n.
[T4] What probability does a language model assign to a complete sequence?
P(w1,w2,…,wn).
[T4] What conditional probability represents next-word prediction?
P(wn | w1,…,wn−1).
[T4] How can language models help spelling correction?
Prefer a more probable word sequence over a less probable one.
[A4] What spelling-correction comparison is given in the lecture?
P(about fifteen minutes from) > P(about fifteen minuets from).
[T4] How can language models help machine translation?
Prefer a more probable translation sequence.
[A4] What machine-translation probability comparison is given?
P(high winds tonight) > P(large winds tonight).
[T4] How can language models help speech recognition?
Prefer the more probable word sequence among acoustically plausible alternatives.
[A4] What speech-recognition comparison is given?
P(I saw a van) >> P(eyes awe of an).
[T4] What additional applications of sentence probabilities are mentioned?
Summarization and question answering.
[T4] What applications specifically use next-word prediction according to the lecture?
Text generation and autocomplete.
[T4] What is the lecture's first proposed method for estimating the probability of a whole sentence?
Count the sentence and divide by the total number of sentences.
[T4] Write the direct whole-sentence frequency estimator.
P(w1^n) = count(w1^n) / #sentences.
[T4] Why is estimating complete sentence probabilities directly from sentence counts problematic?
Most possible sentences will never appear in the corpus or will appear only once.
[T4] What key point about statistical NLP is stated in the lecture?
Your corpus should be much larger than your sample space.
[A4] If there are 10^5 word types and average sentence length is 10 words, approximately how large is the space of possible 10-word sequences?
Greater than 10^50.
[T4] What large-corpus example is given in the lecture?
The Colossal Clean Crawled Corpus (C4), with about 1.4 trillion tokens and over 100 billion sentences.
[T4] Why does even a huge corpus such as C4 fail to cover the complete sentence sample space?
The number of possible sentences is vastly larger than the number of observed sentences.
[T4] What is Method #2 for computing sentence probabilities?
Decompose the joint probability using the chain rule.
[T4] Write the chain-rule decomposition for a sequence w1…wn.
P(w1…wn) = P(w1)P(w2|w1)P(w3|w1w2)…P(wn|w1…wn−1).
[T4] What does the chain rule do to a sentence probability?
Expresses it as a product of conditional probabilities.
[A4] Write the chain rule for three variables A, B, and C.
P(ABC) = P(C|AB)P(B|A)P(A).
[A4] Expand P(the big red dog barks) using the full chain rule.
P(the) × P(big|the) × P(red|the big) × P(dog|the big red) × P(barks|the big red dog).
[T4] Does applying the chain rule by itself solve the data sparsity problem?
No.
[T4] What assumption is introduced to simplify the chain-rule probabilities?
The Markov assumption.
[T4] What is the basic idea of the Markov assumption in language modeling?
The entire prefix history is not necessary; approximate the next-word probability using only a limited recent history.
[T4] In a unigram model, what is P(wn | history) approximated by?
P(wn).
[T4] In a bigram model, what is P(wn | history) approximated by?
P(wn | wn−1).
[T4] In a trigram model, what is P(wn | history) approximated by?
P(wn | wn−2, wn−1).
[T4] How many previous words does a unigram model condition on?
Zero.
[T4] How many previous words does a bigram model condition on?
One.
[T4] How many previous words does a trigram model condition on?
Two.
[T4] In general, how many previous words does an N-gram model use to predict the next word?
N−1.
[A4] Under a bigram model, how is P(the big red dog barks) decomposed?
P(the|) × P(big|the) × P(red|big) × P(dog|red) × P(barks|dog).
[T4] What does represent?
The beginning of a sentence.
[T4] Why is P(the|) not the same as the unigram probability P(the)?
P(the|) specifically measures the probability that “the” begins a sentence.
[T4] What special symbols are shown for a trigram model at the beginning of a sentence?
[A4] Under a trigram model, how is P(the big red dog barks) decomposed?
P(the|
[IT4 — PAGE 31] Compare the sentence-probability formulas for unigram, bigram, and trigram language models.
Unigram multiplies P(wk); bigram uses P(w1|) followed by P(wk|wk−1); trigram uses two start symbols and conditions each later word on the previous two words.
[T4] Write the unigram approximation to sentence probability.
P(w1…wn) = ∏_{k=1}^n P(wk).
[T4] Write the bigram approximation to sentence probability.
P(w1…wn) = P(w1|) ∏_{k=2}^n P(wk|wk−1).
[T4] Write the trigram approximation to sentence probability.
P(w1…wn) = P(w1|
[T4] How is a bigram conditional probability estimated from corpus counts?
P(wn|wn−1) = C(wn−1,wn) / C(wn−1).
[T4] What is the numerator when estimating P(wn|wn−1)?
The count of the bigram wn−1 wn.
[T4] What is the denominator when estimating P(wn|wn−1)?
The count of the previous word wn−1.
[T4] How can C(wn−1) be expressed in terms of bigram counts?
C(wn−1) = Σw C(wn−1,w).
[A4] In the lecture corpus " The big red dog barks at the big pink dog", what is C(big,red)?
1.
[A4] In that same corpus, what is C(big)?
2.
[A4] In that corpus, what is P(red|big)?
C(big,red)/C(big) = 1/2.
[T4] How many word tokens are reported for the example corpus?
10.
[T4] How many word types are reported for the example corpus?
7.
[T4] What is the general maximum-likelihood N-gram estimate shown in the lecture?
P(wk | wk−N+1…wk−1) = C(wk−N+1…wk) / C(wk−N+1…wk−1).
[T4] What is ?
The end-of-sentence token.
[A4] What three sentences are used in the second bigram-counting example?
[T4] What information does an N-gram frequency table contain?
Counts of one word following another word.
[IT4 — PAGE 44] In the lecture's frequency table, what do rows and columns jointly represent?
A cell gives the count C(wn−1,wn) for a particular previous-word/next-word pair.
[T4] How is a frequency table converted into a bigram probability table?
Normalize each next-word count by the total count of the corresponding previous word to obtain P(wn|wn−1).
[IT4 — PAGE 47] Read the highlighted probabilities from the bigram table.
P(want|I)=.32; P(to|want)=.65; P(food|Chinese)=.56; P(lunch|eat)=.055.
[A4] What is P(want|I) in the lecture's probability table?
.32.
[A4] What is P(to|want) in the lecture's probability table?
.65.
[A4] What is P(food|Chinese) in the lecture's probability table?
.56.
[A4] What is P(lunch|eat) in the lecture's probability table?
.055.
[A4] Is P(I|want) necessarily equal to P(want|I)?
No.
[A4] What does P(I|want) ask?
Given that the previous word is “want,” what is the probability that the next word is “I”?
[A4] What does P(want|I) ask?
Given that the previous word is “I,” what is the probability that the next word is “want”?
[T4] Why don't the displayed cells in each row of the shown probability matrix necessarily sum to 1?
The slide is not showing the entire matrix.
[T4] What is the basic idea of language generation using an N-gram model?
Choose N-grams according to their probabilities and string them together.
[T4] What is greedy decoding?
At each step, choose the highest-probability next word.
[T4] What sequence does the lecture's greedy-decoding example generate?
[T4] What is sampling in language generation?
Choose the next word according to the model's probability distribution rather than always choosing the maximum-probability word.
[T4] What sampled sequence is shown in the lecture?
[T4] Why can sampling produce a different sequence from greedy decoding?
Sampling can select a non-maximum-probability word according to its probability.
[T4] What other generation methods are mentioned after greedy decoding and sampling?
Top-k and top-p.
[IT4 — PAGE 64] Compare the generated sequences produced by greedy decoding and sampling in the lecture.
Greedy decoding produces I want to eat lunch , while the shown sampling path produces I want to eat Chinese food .
[T4] What two practical problems are identified when applying an N-gram sentence-probability formula?
Underflow from multiplying very small numbers and unseen words.
[T4] Why can multiplying many language-model probabilities cause underflow?
The individual probabilities are small, so their product can become extremely close to zero.
[T4] How does the lecture propose dealing with numerical underflow?
Work in log space.
[T4] What logarithmic identity allows products of probabilities to be handled as sums?
p1×p2×p3×p4 = exp(log p1 + log p2 + log p3 + log p4).
[T4] What is Zipf's law as described in the lecture?
A small number of word types occur very frequently, while a large number of word types occur much less frequently.
[A4] What types of words are given as examples of highly frequent words under Zipf's law?
Function words such as “the,” “of,” and “in.”
[T4] What implication does Zipf's law have for corpus coverage?
A very large corpus is needed to observe many of the rare word types.
[T4] What distinction does the lecture make between different zero counts?
Some zeroes represent truly impossible events, while others represent low-frequency events that simply were not observed in the corpus.
[T4] Why is a zero N-gram probability a problem when computing a sequence probability?
Multiplying by one zero makes the probability of the whole sequence zero.
[T4] What solutions to unseen events are listed in the lecture?
Smoothing methods such as add-one, Good-Turing, backoff, and deleted interpolation, plus use of an
[T4] What does
An unknown or unseen word.
[T4] During inference, what probability can be used for an unseen word?
The probability of
[T4] How does the lecture suggest creating
Substitute the first occurrence of every word type with