L4 Language Models: N-Grams, Neural LMs & Evaluation

0.0(0)
Studied by 0 people
call kaiCall Kai
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/233

encourage image

There's no tags or description

Looks like no tags are added yet.

Last updated 1:50 AM on 10/5/26
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

234 Terms

1
New cards

[T4] What is the main topic of Lecture 4?

Language models.

2
New cards

[T4] What three main topics are covered in Lecture 4?

N-gram language models (count-based), neural language models, and language model evaluation.

3
New cards

[T4] What shift in focus does Lecture 4 make relative to earlier material?

From discussing words mostly in isolation to discussing sequences of words.

4
New cards

[T4] What previous topic is reviewed at the beginning of Lecture 4?

Real-word spelling errors.

5
New cards

[T4] What is the first step for correcting real-word spelling errors?

Collect common confusion sets C={C1,…,Cn}.

6
New cards

[A4] What confusion-set examples are reviewed at the beginning of Lecture 4?

{their, they’re, there}, {to, too, two}, and {weather, whether}.

7
New cards

[T4] When is the real-word spelling-correction procedure triggered?

Whenever an encountered word c' belongs to one of the confusion sets Ci.

8
New cards

[T4] After encountering c' ∈ Ci, what is computed first?

The probability of the sentence containing c'.

9
New cards

[T4] What is done with the other words c ∈ Ci where c ≠ c'?

Substitute each into the sentence and compute the probability of the resulting sentence.

10
New cards

[T4] How is the final correction selected for a real-word spelling error?

Choose the candidate yielding the highest-probability sentence.

11
New cards

[A4] In the lecture example, what two sentences are compared?

“The whether forecast for today is sunny” and “The weather forecast for today is sunny.”

12
New cards

[T4] What is a probabilistic language model used to compute?

The probability of a sentence or sequence of words.

13
New cards

[T4] What other fundamental task can a language model perform?

Predict the next word given the preceding words.

14
New cards

[T4] What notation is used for a sequence of words from w1 through wn?

w1^n.

15
New cards

[T4] What probability does a language model assign to a complete sequence?

P(w1,w2,…,wn).

16
New cards

[T4] What conditional probability represents next-word prediction?

P(wn | w1,…,wn−1).

17
New cards

[T4] How can language models help spelling correction?

Prefer a more probable word sequence over a less probable one.

18
New cards

[A4] What spelling-correction comparison is given in the lecture?

P(about fifteen minutes from) > P(about fifteen minuets from).

19
New cards

[T4] How can language models help machine translation?

Prefer a more probable translation sequence.

20
New cards

[A4] What machine-translation probability comparison is given?

P(high winds tonight) > P(large winds tonight).

21
New cards

[T4] How can language models help speech recognition?

Prefer the more probable word sequence among acoustically plausible alternatives.

22
New cards

[A4] What speech-recognition comparison is given?

P(I saw a van) >> P(eyes awe of an).

23
New cards

[T4] What additional applications of sentence probabilities are mentioned?

Summarization and question answering.

24
New cards

[T4] What applications specifically use next-word prediction according to the lecture?

Text generation and autocomplete.

25
New cards

[T4] What is the lecture's first proposed method for estimating the probability of a whole sentence?

Count the sentence and divide by the total number of sentences.

26
New cards

[T4] Write the direct whole-sentence frequency estimator.

P(w1^n) = count(w1^n) / #sentences.

27
New cards

[T4] Why is estimating complete sentence probabilities directly from sentence counts problematic?

Most possible sentences will never appear in the corpus or will appear only once.

28
New cards

[T4] What key point about statistical NLP is stated in the lecture?

Your corpus should be much larger than your sample space.

29
New cards

[A4] If there are 10^5 word types and average sentence length is 10 words, approximately how large is the space of possible 10-word sequences?

Greater than 10^50.

30
New cards

[T4] What large-corpus example is given in the lecture?

The Colossal Clean Crawled Corpus (C4), with about 1.4 trillion tokens and over 100 billion sentences.

31
New cards

[T4] Why does even a huge corpus such as C4 fail to cover the complete sentence sample space?

The number of possible sentences is vastly larger than the number of observed sentences.

32
New cards

[T4] What is Method #2 for computing sentence probabilities?

Decompose the joint probability using the chain rule.

33
New cards

[T4] Write the chain-rule decomposition for a sequence w1…wn.

P(w1…wn) = P(w1)P(w2|w1)P(w3|w1w2)…P(wn|w1…wn−1).

34
New cards

[T4] What does the chain rule do to a sentence probability?

Expresses it as a product of conditional probabilities.

35
New cards

[A4] Write the chain rule for three variables A, B, and C.

P(ABC) = P(C|AB)P(B|A)P(A).

36
New cards

[A4] Expand P(the big red dog barks) using the full chain rule.

P(the) × P(big|the) × P(red|the big) × P(dog|the big red) × P(barks|the big red dog).

37
New cards

[T4] Does applying the chain rule by itself solve the data sparsity problem?

No.

38
New cards

[T4] What assumption is introduced to simplify the chain-rule probabilities?

The Markov assumption.

39
New cards

[T4] What is the basic idea of the Markov assumption in language modeling?

The entire prefix history is not necessary; approximate the next-word probability using only a limited recent history.

40
New cards

[T4] In a unigram model, what is P(wn | history) approximated by?

P(wn).

41
New cards

[T4] In a bigram model, what is P(wn | history) approximated by?

P(wn | wn−1).

42
New cards

[T4] In a trigram model, what is P(wn | history) approximated by?

P(wn | wn−2, wn−1).

43
New cards

[T4] How many previous words does a unigram model condition on?

Zero.

44
New cards

[T4] How many previous words does a bigram model condition on?

One.

45
New cards

[T4] How many previous words does a trigram model condition on?

Two.

46
New cards

[T4] In general, how many previous words does an N-gram model use to predict the next word?

N−1.

47
New cards

[A4] Under a bigram model, how is P(the big red dog barks) decomposed?

P(the|) × P(big|the) × P(red|big) × P(dog|red) × P(barks|dog).

48
New cards

[T4] What does represent?

The beginning of a sentence.

49
New cards

[T4] Why is P(the|) not the same as the unigram probability P(the)?

P(the|) specifically measures the probability that “the” begins a sentence.

50
New cards

[T4] What special symbols are shown for a trigram model at the beginning of a sentence?

and .
51
New cards

[A4] Under a trigram model, how is P(the big red dog barks) decomposed?

P(the|) × P(big| the) × P(red|the big) × P(dog|big red) × P(barks|red dog).

52
New cards

[IT4 — PAGE 31] Compare the sentence-probability formulas for unigram, bigram, and trigram language models.

Unigram multiplies P(wk); bigram uses P(w1|) followed by P(wk|wk−1); trigram uses two start symbols and conditions each later word on the previous two words.

53
New cards

[T4] Write the unigram approximation to sentence probability.

P(w1…wn) = ∏_{k=1}^n P(wk).

54
New cards

[T4] Write the bigram approximation to sentence probability.

P(w1…wn) = P(w1|) ∏_{k=2}^n P(wk|wk−1).

55
New cards

[T4] Write the trigram approximation to sentence probability.

P(w1…wn) = P(w1|) P(w2|w1) ∏_{k=3}^n P(wk|wk−2,wk−1).

56
New cards

[T4] How is a bigram conditional probability estimated from corpus counts?

P(wn|wn−1) = C(wn−1,wn) / C(wn−1).

57
New cards

[T4] What is the numerator when estimating P(wn|wn−1)?

The count of the bigram wn−1 wn.

58
New cards

[T4] What is the denominator when estimating P(wn|wn−1)?

The count of the previous word wn−1.

59
New cards

[T4] How can C(wn−1) be expressed in terms of bigram counts?

C(wn−1) = Σw C(wn−1,w).

60
New cards

[A4] In the lecture corpus " The big red dog barks at the big pink dog", what is C(big,red)?

1.

61
New cards

[A4] In that same corpus, what is C(big)?

2.

62
New cards

[A4] In that corpus, what is P(red|big)?

C(big,red)/C(big) = 1/2.

63
New cards

[T4] How many word tokens are reported for the example corpus?

10.

64
New cards

[T4] How many word types are reported for the example corpus?

7.

65
New cards

[T4] What is the general maximum-likelihood N-gram estimate shown in the lecture?

P(wk | wk−N+1…wk−1) = C(wk−N+1…wk) / C(wk−N+1…wk−1).

66
New cards

[T4] What is ?

The end-of-sentence token.

67
New cards

[A4] What three sentences are used in the second bigram-counting example?

I am Sam ; Sam I am ; I do not like green eggs and ham .
68
New cards

[T4] What information does an N-gram frequency table contain?

Counts of one word following another word.

69
New cards

[IT4 — PAGE 44] In the lecture's frequency table, what do rows and columns jointly represent?

A cell gives the count C(wn−1,wn) for a particular previous-word/next-word pair.

70
New cards

[T4] How is a frequency table converted into a bigram probability table?

Normalize each next-word count by the total count of the corresponding previous word to obtain P(wn|wn−1).

71
New cards

[IT4 — PAGE 47] Read the highlighted probabilities from the bigram table.

P(want|I)=.32; P(to|want)=.65; P(food|Chinese)=.56; P(lunch|eat)=.055.

72
New cards

[A4] What is P(want|I) in the lecture's probability table?

.32.

73
New cards

[A4] What is P(to|want) in the lecture's probability table?

.65.

74
New cards

[A4] What is P(food|Chinese) in the lecture's probability table?

.56.

75
New cards

[A4] What is P(lunch|eat) in the lecture's probability table?

.055.

76
New cards

[A4] Is P(I|want) necessarily equal to P(want|I)?

No.

77
New cards

[A4] What does P(I|want) ask?

Given that the previous word is “want,” what is the probability that the next word is “I”?

78
New cards

[A4] What does P(want|I) ask?

Given that the previous word is “I,” what is the probability that the next word is “want”?

79
New cards

[T4] Why don't the displayed cells in each row of the shown probability matrix necessarily sum to 1?

The slide is not showing the entire matrix.

80
New cards

[T4] What is the basic idea of language generation using an N-gram model?

Choose N-grams according to their probabilities and string them together.

81
New cards

[T4] What is greedy decoding?

At each step, choose the highest-probability next word.

82
New cards

[T4] What sequence does the lecture's greedy-decoding example generate?

I want to eat lunch .
83
New cards

[T4] What is sampling in language generation?

Choose the next word according to the model's probability distribution rather than always choosing the maximum-probability word.

84
New cards

[T4] What sampled sequence is shown in the lecture?

I want to eat Chinese food .
85
New cards

[T4] Why can sampling produce a different sequence from greedy decoding?

Sampling can select a non-maximum-probability word according to its probability.

86
New cards

[T4] What other generation methods are mentioned after greedy decoding and sampling?

Top-k and top-p.

87
New cards

[IT4 — PAGE 64] Compare the generated sequences produced by greedy decoding and sampling in the lecture.

Greedy decoding produces I want to eat lunch , while the shown sampling path produces I want to eat Chinese food .

88
New cards

[T4] What two practical problems are identified when applying an N-gram sentence-probability formula?

Underflow from multiplying very small numbers and unseen words.

89
New cards

[T4] Why can multiplying many language-model probabilities cause underflow?

The individual probabilities are small, so their product can become extremely close to zero.

90
New cards

[T4] How does the lecture propose dealing with numerical underflow?

Work in log space.

91
New cards

[T4] What logarithmic identity allows products of probabilities to be handled as sums?

p1×p2×p3×p4 = exp(log p1 + log p2 + log p3 + log p4).

92
New cards

[T4] What is Zipf's law as described in the lecture?

A small number of word types occur very frequently, while a large number of word types occur much less frequently.

93
New cards

[A4] What types of words are given as examples of highly frequent words under Zipf's law?

Function words such as “the,” “of,” and “in.”

94
New cards

[T4] What implication does Zipf's law have for corpus coverage?

A very large corpus is needed to observe many of the rare word types.

95
New cards

[T4] What distinction does the lecture make between different zero counts?

Some zeroes represent truly impossible events, while others represent low-frequency events that simply were not observed in the corpus.

96
New cards

[T4] Why is a zero N-gram probability a problem when computing a sequence probability?

Multiplying by one zero makes the probability of the whole sequence zero.

97
New cards

[T4] What solutions to unseen events are listed in the lecture?

Smoothing methods such as add-one, Good-Turing, backoff, and deleted interpolation, plus use of an token.

98
New cards

[T4] What does represent?

An unknown or unseen word.

99
New cards

[T4] During inference, what probability can be used for an unseen word?

The probability of .

100
New cards

[T4] How does the lecture suggest creating examples in the training corpus?

Substitute the first occurrence of every word type with .