1/48
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Difference of AI, ML, and DL
Artificial Intelligence:
Broad field of building systems that perform tasks associated with intelligence
Machine Learning:
Systems learn patterns from data rather than relying only on manually specified rules
Deep Learning
A class of ML methods based on multi-layer neural networks

What is ML perspective?
Traditional programming: Input + Rules → Output
ML : Input + Desired Output→ Learned Model

Why DL is learning through multiple computational layers?
Neural networks performs a sequence of transformations
Each layer transforms 1 representation into another
x → h1 → h2 → … → y
Earlier layers operate closer to raw input
Later layers operate on increasingly transformed representations
So, ‘Deep’ generally refers to composing multiple layers of computation
What are the basic 4 components of DL training system?
Data → what the model observes
Model → how predictions are computed
Loss function→ measures prediction error
Optimizer → determines how parameters are updated
=> tgt they form a learning system
What does training data consist of, and what role does it play?
Training data consists of pairs: (x_i , y_i)
x_i : input
y_i : target/ label
Examples:
Image → object category
Test → sentiment
Patient measurements→ diagnosis
House characteristics→ price
Role: dataset provides evidence from which the model learns its parameters

What is a model in DL?
Model: parameterized function that maps an input x to a prediction y(hat)
→ (*formula in pict)
→ where theta represents the trainable parameters, such as weights and biases
Diff parameter values → Diff predictions
So, training searches for parameter values that perform well on the task

What is the loss function and what does the size of the loss tell us?
Loss function: measures disagreement btw the model’s prediction, y(hat) and the target,y
→ *formula in pict
Small loss → prediction is good
Large loss → prediction is bad/poor
Training tries to find parameters(theta*) that minimise the loss

What is an optimizer and how is it related to the gradient?
Optimizer: determines how the model parameters should change
Gradient (*formula in pict) : tells how loss changes when parameters change
Optimizer use gradient information to update parameters
Common optimizers:
Gradient Descent
SGD
Momentum
Adam

What happens during one complete learning/training iteration? Or how does 4 components interact?
The sequence is:
Select data → Compute predictions → Compare predictions with targets → Compute loss → Compute gradients→ Update parameters → Repeat
Therefore, training is an iterative optimization process
What are scalars, vectors, matrices, and tensors in DL?
Scalar → single #
e.g : loss value
Vector → 1D collection of #
e.g : feature vector
Matrix → 2D array
e.g : weight matrix
Tensor → general multidimensional array
=> Tensors are core data structure used in DL frameworks
What are the common tensor structures for different kinds of data?
Scalar loss → [ ]
Feature vector → [ features]
Batch of feature vectors → [batch, features]
Grayscale images → [batch, channels, height, width]
=> each dimension depends on problem
Why are tensor shapes important in neural networks?
Operations require compatible dimensions
*Example
Correct shapes are essential bcs many implementations errors are actually shape mismatches

What common mathematical notation is used for neural networks?
Important to conceptually distinguish inputs, parameters, predictions, and targets

What is a linear transformation in a neural network?
A basic layer can be represented as:
z = Wx + b
x = input
W = weights
b = bias
z = output
Both W and b can be learned from data
A neural network is built from repeated parameterized transformations

What happens during a basic feedforward computation?
For: *forms in pict
The network:
Receives input x → Perform a linear transformation → Applies non-linear function g → Compute output → Produces prediction y(hat)
=> This input-to-output computation is called the forward pass

Why can a neural networks be viewed as composition of functions?
Deep networks repeatedly compose functions
The ‘forward pass’ evaluate these functions input→output, while backpropagation later moves through the same structure in the opposite direction.

What is computational graph?
Computational graph represents a computation using:
Variables: nodes/values
Mathematical operations: transformations
Dependencies btw computations
Each intermediate result bcms part of the graph

Why are computational graphs important for gradient computation?
Final loss depends on many intermediate computations, and each parameter may influence the loss through multiple operations
Computational graph tells:
what depends on what
The order computation occurs
The paths along which gradients must propagate
So, modern DL frameworks construct and use these graphs automatically
What does derivative represent?
For: y = f(x) → the derivative: dy/dx
=> measures how sensitive y is to a small change in x
+ve derivative → increasing x tends to increase y
-ve derivative → increasing x tends to decrease y
Large magnitude → high sensitivity
Small magnitude → low sensitivity

What is the chain rule, and why is it important for neural networks?
For composed functions:
z = f(y) , y = g(x)
influence of x on z is found by multiplying local derivatives along the path
Conceptually: Local effect x Local effect = Total effect
Neural networks contain many composed functions, so the chain rule is the mathematical foundation of backpropagation.


Explain the chain rule using a=2x, y=a², and x=3
Forward:
a = 2(3) = 6
y = 6² = 36
Local derivatives:
da/dx = 2
dy/da = 2a = 12
Therefore, total sensitivity combines the local sensitivities:
dy/dx = (dy/da)(da/dx) = 12(2) =24
This principle scales to much larger computational graphs

What is backpropagation?
backpropagation: Efficiently compute the gradient of the loss with respect to model parameters
It works backward through the computational graph :
Start from loss → Compute local derivatives → Apply chain rule → Propagate gradients backward → Obtain a gradient for every trainable parameter
What is the difference between the forward pass and backward pass?
Forward pass:
Input → Prediction
Compute intermediate activations
Produces predictions
Compute the loss
Backward pass:
Loss → Parameters
Compute gradients
Determines how sensitive the loss is to the parameters

What does the gradient with respect to model parameters tell us?
If: theta = (w1,w2,…,wn)
Then: (*form in pict) is a collection of partial derivatives
Each component essentially ask:
“ If this parameter changes slightly, how will the loss change?”
The sign indicates direction of sensitivity while the magnitude indicates how sensitive the loss is locally

Is backpropagation the same as gradient descent?
NO
Backpropagation → Compute gradients
Optimisation/ gradient descent→ Use gradients to change parameters
So:
Backpropagation → Gradients → Optimiser → Parameter update
Therefore: Backpropagation IS NOT Gradient Descent
This distinction is especially important in PyTorch
What is the basic idea of gradient descent?
Once gradients are available :
Calculate the current slope of the loss
Move parameters in a direction that reduces loss
Recalculate
Repeat
=> Learning rate controls the size of each parameter update
How can optimization be understood as searching a loss landscape?
Imagine the loss as a landscape:
Horizontal axes → model parameters
Vertical axes → loss
Training starts at an initial parameter location
Gradients provide local directional information, and optimization repeatedly moves toward lower-loss regions
Real neural networks may contain millions or billions of parameter dimensions, so their loss landscapes normally cannot be directly visualised

What is Stochastic Gradient Descent (SGD) and why is it used?
Computing gradients using the entire dataset can be expensive
SGD uses subset of the training data to estimate the gradient
Advantages:
Lower computation per update
More frequent parameter updates
Suitable for large datasets
Trade-off(Disadvantage):
Gradient estimates bcm noisier
What is mini-batch training, and why is it commonly used?
Instead of processing:
One example at a time, OR
The entire dataset at once
DL commonly processes a mini-batch or a subset of examples together
It provides a practical balance between:
Computational efficiency
Memory usage
Gradient stability
If a dataset has 50,000 examples and batch size = 100, approximately how many parameter updates occur per epoch?
50,000/100 =500
So approximately 500 parameter updates per epoch
What is the difference between a sample, batch, iteration, and epoch?
Sample: One training example
Batch/ Mini-batch: Subset of examples processed together
Iteration: One forward-backward-update cycle
Epoch: One complete pass through the training dataset
Example: 10,000 samples with batch size 100
10,000/100=100 iterations
So approximately 100 iterations = 1 epoch
What is the learning rate?
The learning rate controls the magnitude of parameter updates
Too small:
Training progresses slowly
Many iterations may be required
Too large:
Updates may overshoot useful regions
Loss may oscillate or diverge
Appropriate:
Loss decreases at useful and reasonably stable rate

What is momentum and why is it useful?
SGD updates can vary substantially between mini-batches
Momentum incorporate information from previous updates to:
Build speed in directions with consistent gradients
Reduce sensitivity to short-term fluctuations
Smooth the optimisation trajectory
Analogy: A ball rolling downhill
Gradient = slope
Momentum = accumulated velocity

What is Adam and what is its main intuition?
Adam combines:
Momentum-like accumulation
Adaptive scaling of parameter updates
Therefore:
Different parameters can receive differently scaled updates
Recent gradient behaviour influences the update
Adam is widely used as a general-purpose optimizer
=> Importantly, Adam changes how gradient are used; It does not replace backpropagation
How do Gradient Descent, SGD, Mini-batch SGD, Momentum, and Adam differ?

What are the three key PyTorch operations in a training interation?
forward()
Compute model predictions and constructs the forward computation
backward()
Compute gradients using backpropagation
step()
Use gradients to update parameters
=> forward → loss → backward → step
What happens during forward() in PyTorch?
During forward pass:
Input tensors enter the model
Layers perform transformations
Intermediate results are computed
The model produces predictions
PyTorch records operations needed for later gradient computation
Example: pred = model(x)

What happens when loss.backward() is called?
The backward pass:
Starts from the loss
Traveses the computational graph backward
Applies the chain rule
Computes gradients for trainable parameters
=> Afterward, the parameters have associated gradient information

What happens when optimizer.step() is called?
The optimizer:
Reads parameter gradients
Applies its update rule
Modifies the model parameters
=> exact update depends on the optimizer, such as SGD or Adam

Why must gradients be reset before the next training iteration in PyTorch?
PyTorch accumulates gradients unless they are cleared
Therefore: optimizer.zero_grad() → clears the previous gradients bfr calculating new ones
Standard iteration:
Clear gradients → Predictions → Loss → New gradients → Parameter update
This ensures each mini-batch contributes the intended gradient information
What is the difference between CPU and GPU for DL?
CPU
General-purpose processor
Small number of powerful cores
Good for flexible comtrol and sequential computation
GPU
Large number of parallel computational units
Effective for large tensor operations
Well suited for matrix-heavy neural networks computation
=> DL benefits heavily from parallel numerical computation which is why GPUs are commonly used
Why must the model and data be on the same device in PyTorch?
The following must be on the same device :
Model parameters
Input tensors
Target tensors
=> A device mismatch prevents computation from proceeding correctly


What is the minimal PyTorch training loop, and what does each line mean?
Meaning:
for x, y in dataloader → get training data
zero_grad() → clear previous gradients
model(x) → forward computation
loss_fn(pred,y) → compute prediction error
loss.backward() → backpropagation / compute gradients
optimizer.step() → update parameters
Conceptually:
DATA → MODEL → LOSS → OPTIMIZER

Why can neural network training be described as a repeated feedback process?
At iteration t :
Current parameters → Predictions → Loss → Gradients → Parameter update → New parameters → New predictions
So: theta0 → theta1 → theta2 → …
Repeated many times, traing bcms a feedback process driven by predictions error

Explain how a neural network learns from data from beginning to end.
The model receives input and performs a forward pass to produce a prediction.
The loss function compares the prediction with the target.
Backpropagation then uses the chain rule to compute how each parameter contributed to the loss.
The optimizer uses these gradients to update the parameters toward lower loss.
This process repeats over many iterations.
Why do we need both backpropagation and an optimizer? Why can’t one replace the other?
They perform different jobs.
Backpropagation calculates the gradients, telling us how the loss is sensitive to each parameter.
The optimizer uses those gradients to actually change the parameters.
Therefore:
Backpropagation = find how to change
Optimizer = perform the change
Why is the chain rule essential to training a deep neural network?
A deep neural network is a composition of many functions/layers.
A parameter in an early layer affects later computations and eventually the loss.
The chain rule allows us to combine local derivatives across these operations to determine how an earlier parameter affects the final loss.
This is what makes backpropagation possible.
Why don’t we normally compute gradients using the entire dataset for every update?
Using the entire dataset can be computationally expensive.
Mini-batches allow gradients to be estimated using only part of the data, giving:
Lower computation per update
More frequent updates
Better practicality for large datasets
The trade-off is that the gradient estimate becomes noisier.
How are forward pass, loss, backward pass, gradient, and optimizer connected?
Think of the entire training process as:
Input → Forward → Prediction → Loss → Backward → Gradients → Optimizer → Updated Parameters
=> Then the updated parameters are used for the next forward pass, and the cycle repeats