Send a link to your students to track their progress
13 Terms
1
New cards
Describe the primary architectural distinction between BERT-like models and GPT-like models in terms of their use of the original Transformer's Encoder and Decoder submodules.
BERT uses only the Transformer encoder (bidirectional, for understanding/classification); GPT uses only the decoder (unidirectional, for next-token generation).
2
New cards
Why is it beneficial, in terms of context understanding and model training, to add both token embeddings and positional embeddings to the raw input sequence, rather than just relying on the token embeddings alone?
Self-attention has no built-in sense of order, so token embeddings alone would treat the input as a bag of tokens. Positional embeddings inject order so the model can learn word position and context.
3
New cards
Explain the importance of using <|endoftext|> tokens when processing multiple independent documents (e.g., books or articles) into a single large input corpus for LLM training.
4
New cards
What is the fundamental difference between the core self-attention mechanism and the attention mechanisms used historically in recurrent neural networks (RNNs) for tasks like translation?
RNN attention was cross-attention between a decoder's hidden state and the encoder's states, computed step by step with recurrence. Self-attention relates every token to every other token within the same sequence, in parallel, with no recurrence.
5
New cards
Why is the causal attention mask necessary for generative LLMs like GPT, and specifically, how does the causal attention matrix structure prevent information leakage from future tokens?
GPT predicts the next token, so it must not see later tokens. The causal mask sets attention scores above the diagonal to −∞ before softmax, giving future positions zero weight so each token only attends to itself and earlier tokens.
6
New cards
Discuss the specific computational advantage achieved by implementing Multi-Head Attention using weight splitting and batched matrix multiplications, compared to stacking multiple independent single-head attention layers via a wrapper class.
A single large Q/K/V linear projection is split (reshaped) into heads and processed with one batched matrix multiplication, instead of looping over separate heads sequentially. Same math, but fewer operations and far better GPU parallelism.
7
New cards
Explain how Layer Normalization works on a batch of token embeddings with shape [batch_size, num_tokens, embedding_size], and specify which dimension it normalizes across.
LayerNorm operates on each token's embedding vector independently along the last dimension (dim=-1, embedding size): subtract the mean, divide by √(var + ε), then apply learnable scale and shift. Batch and token dimensions are untouched.
8
New cards
Contrast the Layer Normalization position used in modern models like GPT-2 (Pre-LayerNorm) with the position in the original transformer model (Post-LayerNorm), and state the benefit attributed to the modern approach.
Original Transformer: Post-LN (normalize after attention/FFN and the residual add). GPT-2: Pre-LN (normalize before attention/FFN). Pre-LN gives more stable gradients and training, with less dependence on learning-rate warmup.
9
New cards
Explain the practical implication of applying Temperature Scaling with a value greater than 1.0 (e.g., T = 5) during text generation. What effect does this have on the resulting probability distribution and the diversity of the output text?
Logits are divided by T; T > 1 flattens the distribution toward uniform, so sampling is more diverse and creative but also more likely to produce low-probability, nonsensical tokens.
10
New cards
What is the fundamental mechanism of Top-K sampling, and how does combining it with Temperature Scaling help mitigate the risk of generating nonsensical output tokens compared to using high temperatures alone?
Top-K keeps only the K highest logits, masks the rest to −∞, then applies softmax and samples. Combined with temperature, the flattening only redistributes probability among plausible tokens, so high T can't surface garbage tokens.
11
New cards
Why is calculating perplexity often considered a more interpretable metric for evaluating generative text models compared to using the raw Cross-Entropy Loss value?
Perplexity = exp(loss). It can be read as the effective number of tokens the model is "choosing between" at each step, which is more intuitive than a raw log-based loss.
12
New cards
In the process of loading pretrained model weights from OpenAI's GPT-2, explain the concept of "weight tying" and how it differs from using separate weight matrices for the token embedding layer and the final output layer.
Weight tying reuses the token embedding matrix (transposed) as the output head weights, so one matrix serves both roles. Separate matrices add ~38.6M parameters; the original GPT-2 ties them, so loading OpenAI weights requires assigning the embedding weights to the output head.
13
New cards
When calculating the variance for Layer Normalization using the implemented LayerNorm class, the unbiased=False setting is used. Explain the technical difference this setting introduces (related to n versus n − 1), and why this "biased" approach was adopted for compatibility with GPT-2.
unbiased=False divides the variance by n; the unbiased version divides by n − 1 (Bessel's correction). The biased form matches GPT-2's original TensorFlow implementation so pretrained weights behave identically; with large embedding dims the difference is negligible.