[DL] 1. Foundation DL

0.0(0)
Studied by 0 people
call kaiCall Kai
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/48

encourage image

There's no tags or description

Looks like no tags are added yet.

Last updated 9:39 AM on 10/7/26
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

49 Terms

1
New cards

Difference of AI, ML, and DL

  1. Artificial Intelligence:

    Broad field of building systems that perform tasks associated with intelligence

  2. Machine Learning:

    Systems learn patterns from data rather than relying only on manually specified rules

  3. Deep Learning

    A class of ML methods based on multi-layer neural networks



<ol><li><p>Artificial Intelligence:</p><p>Broad field of building systems that perform tasks associated with intelligence</p></li><li><p>Machine Learning:</p><p>Systems learn patterns from data rather than relying only on manually specified rules</p></li><li><p>Deep Learning</p><p>A class of ML methods based on multi-layer neural networks</p></li></ol><p></p><p></p>
2
New cards

What is ML perspective?

  • Traditional programming: Input + Rules → Output

  • ML : Input + Desired Output→ Learned Model



<ul><li><p>Traditional programming: Input + Rules → Output</p></li><li><p>ML : Input + Desired Output→ Learned Model</p><p></p></li></ul><p></p>
3
New cards

Why DL is learning through multiple computational layers?

  • Neural networks performs a sequence of transformations

  • Each layer transforms 1 representation into another

    x → h1 → h2 → … → y

  • Earlier layers operate closer to raw input

  • Later layers operate on increasingly transformed representations

  • So, ‘Deep’ generally refers to composing multiple layers of computation


4
New cards

What are the basic 4 components of DL training system?

  1. Data → what the model observes

  2. Model → how predictions are computed

  3. Loss function→ measures prediction error

  4. Optimizer → determines how parameters are updated

    => tgt they form a learning system


5
New cards

What does training data consist of, and what role does it play?

  • Training data consists of pairs: (x_i , y_i)

    • x_i : input

    • y_i : target/ label


  • Examples:

    • Image → object category

    • Test → sentiment

    • Patient measurements→ diagnosis

    • House characteristics→ price


  • Role: dataset provides evidence from which the model learns its parameters


<ul><li><p>Training data consists of pairs: (x_i , y_i)</p><ul><li><p>x_i : input</p></li><li><p>y_i : target/ label</p><p></p></li></ul></li><li><p>Examples:</p><ul><li><p>Image → object category</p></li><li><p>Test → sentiment</p></li><li><p>Patient measurements→ diagnosis</p></li><li><p>House characteristics→ price</p><p></p></li></ul></li><li><p>Role: dataset provides evidence from which the model learns its parameters</p></li></ul><p></p>
6
New cards

What is a model in DL?

  • Model: parameterized function that maps an input x to a prediction y(hat)

→ (*formula in pict)

→ where theta represents the trainable parameters, such as weights and biases

  • Diff parameter values → Diff predictions

  • So, training searches for parameter values that perform well on the task


<ul><li><p>Model: parameterized function that maps an input x to a prediction y(hat)</p></li></ul><p>→ (*formula in pict)</p><p>→ where theta represents the trainable parameters, such as weights and biases</p><ul><li><p>Diff parameter values → Diff predictions</p></li><li><p>So, training searches for parameter values that perform well on the task</p></li></ul><p></p>
7
New cards

What is the loss function and what does the size of the loss tell us?

  • Loss function: measures disagreement btw the model’s prediction, y(hat) and the target,y

    → *formula in pict

    • Small loss → prediction is good

    • Large loss → prediction is bad/poor

  • Training tries to find parameters(theta*) that minimise the loss


<ul><li><p>Loss function: measures disagreement btw the model’s prediction, y(hat) and the target,y</p><p>→ *formula in pict </p><ul><li><p>Small loss → prediction is good</p></li><li><p>Large loss → prediction is bad/poor</p></li></ul></li><li><p>Training tries to find parameters(theta*) that minimise the loss</p></li></ul><p></p>
8
New cards

What is an optimizer and how is it related to the gradient?

  • Optimizer: determines how the model parameters should change

  • Gradient (*formula in pict) : tells how loss changes when parameters change

  • Optimizer use gradient information to update parameters

  • Common optimizers:

    • Gradient Descent

    • SGD

    • Momentum

    • Adam


<ul><li><p>Optimizer: determines how the model parameters should change</p></li><li><p>Gradient (*formula in pict) : tells how loss changes when parameters change </p></li><li><p>Optimizer use gradient information to update parameters </p></li><li><p>Common optimizers:</p><ul><li><p>Gradient Descent</p></li><li><p>SGD</p></li><li><p>Momentum</p></li><li><p>Adam</p></li></ul></li></ul><p></p>
9
New cards

What happens during one complete learning/training iteration? Or how does 4 components interact?

The sequence is:

Select data → Compute predictions → Compare predictions with targets → Compute loss → Compute gradients→ Update parameters → Repeat

  • Therefore, training is an iterative optimization process


10
New cards

What are scalars, vectors, matrices, and tensors in DL?

  • Scalar → single #

    e.g : loss value

  • Vector → 1D collection of #

    e.g : feature vector

  • Matrix → 2D array

    e.g : weight matrix

  • Tensor → general multidimensional array

    => Tensors are core data structure used in DL frameworks


11
New cards

What are the common tensor structures for different kinds of data?

  • Scalar loss → [ ]

  • Feature vector → [ features]

  • Batch of feature vectors → [batch, features]

  • Grayscale images → [batch, channels, height, width]

    => each dimension depends on problem


12
New cards

Why are tensor shapes important in neural networks?

  • Operations require compatible dimensions

  • *Example

  • Correct shapes are essential bcs many implementations errors are actually shape mismatches


<ul><li><p>Operations require compatible dimensions </p></li><li><p>*Example</p></li><li><p>Correct shapes are essential bcs many implementations errors are actually shape mismatches </p></li></ul><p></p>
13
New cards

What common mathematical notation is used for neural networks?

  • Important to conceptually distinguish inputs, parameters, predictions, and targets


<ul><li><p>Important to conceptually distinguish inputs, parameters, predictions, and targets </p></li></ul><p></p>
14
New cards

What is a linear transformation in a neural network?

  • A basic layer can be represented as:

    z = Wx + b

    • x = input

    • W = weights

    • b = bias

    • z = output

  • Both W and b can be learned from data

  • A neural network is built from repeated parameterized transformations


<ul><li><p>A basic layer can be represented as: </p><p>z = Wx + b </p><ul><li><p>x = input</p></li><li><p>W = weights</p></li><li><p>b = bias</p></li><li><p>z = output </p></li></ul></li><li><p>Both W and b can be learned from data</p></li><li><p>A neural network is built from repeated parameterized transformations </p></li></ul><p></p>
15
New cards

What happens during a basic feedforward computation?

For: *forms in pict

  • The network:

Receives input x → Perform a linear transformation → Applies non-linear function g → Compute output → Produces prediction y(hat)

=> This input-to-output computation is called the forward pass

<p>For: *forms in pict</p><ul><li><p>The network:</p></li></ul><p>Receives input x → Perform a linear transformation → Applies non-linear function g → Compute output → Produces prediction y(hat)</p><p>=&gt; This input-to-output computation is called the <strong>forward pass </strong></p>
16
New cards

Why can a neural networks be viewed as composition of functions?

  • Deep networks repeatedly compose functions

  • The ‘forward pass’ evaluate these functions input→output, while backpropagation later moves through the same structure in the opposite direction.


<ul><li><p>Deep networks repeatedly compose functions </p></li><li><p>The ‘forward pass’ evaluate these functions input→output, while backpropagation later moves through the same structure in the opposite direction. </p></li></ul><p></p>
17
New cards

What is computational graph?

  • Computational graph represents a computation using:

    • Variables: nodes/values

    • Mathematical operations: transformations

    • Dependencies btw computations

  • Each intermediate result bcms part of the graph


<ul><li><p>Computational graph represents a computation using:</p><ul><li><p>Variables: nodes/values</p></li><li><p>Mathematical operations: transformations</p></li><li><p>Dependencies btw computations </p></li></ul></li><li><p>Each intermediate result bcms part of the graph </p></li></ul><p></p>
18
New cards

Why are computational graphs important for gradient computation?

  • Final loss depends on many intermediate computations, and each parameter may influence the loss through multiple operations

  • Computational graph tells:

    • what depends on what

    • The order computation occurs

    • The paths along which gradients must propagate

  • So, modern DL frameworks construct and use these graphs automatically


19
New cards

What does derivative represent?

For: y = f(x) → the derivative: dy/dx

=> measures how sensitive y is to a small change in x

  • +ve derivative → increasing x tends to increase y

  • -ve derivative → increasing x tends to decrease y

  • Large magnitude → high sensitivity

  • Small magnitude → low sensitivity


<p>For: y = f(x) → the derivative: dy/dx</p><p>=&gt; measures how sensitive y is to a small change in x</p><ul><li><p>+ve derivative → increasing x tends to increase y</p></li><li><p>-ve derivative → increasing x tends to decrease y</p></li><li><p>Large magnitude → high sensitivity </p></li><li><p>Small magnitude → low sensitivity </p></li></ul><p></p>
20
New cards

What is the chain rule, and why is it important for neural networks?

  • For composed functions:

    z = f(y) , y = g(x)

    influence of x on z is found by multiplying local derivatives along the path

  • Conceptually: Local effect x Local effect = Total effect

  • Neural networks contain many composed functions, so the chain rule is the mathematical foundation of backpropagation.


<ul><li><p>For composed functions:</p><p>z = f(y)  ,  y = g(x)</p><p>influence of x on z is found by multiplying local derivatives along the path </p></li><li><p>Conceptually: Local effect x Local effect = Total effect </p></li><li><p><span>Neural networks contain many composed functions, so the chain rule is the <strong>mathematical foundation of backpropagation</strong>.</span></p></li></ul><p></p>
21
New cards
<p>Explain the chain rule using a=2x, y=a², and x=3</p>

Explain the chain rule using a=2x, y=a², and x=3

  • Forward:

    • a = 2(3) = 6

    • y = 6² = 36

  • Local derivatives:

    • da/dx = 2

    • dy/da = 2a = 12

  • Therefore, total sensitivity combines the local sensitivities:

    dy/dx = (dy/da)(da/dx) = 12(2) =24

  • This principle scales to much larger computational graphs


<ul><li><p>Forward: </p><ul><li><p>a = 2(3) = 6 </p></li><li><p>y = 6² = 36</p></li></ul></li><li><p>Local derivatives:</p><ul><li><p>da/dx = 2</p></li><li><p>dy/da = 2a = 12</p></li></ul></li><li><p>Therefore, total sensitivity combines the local sensitivities:</p><p>dy/dx = (dy/da)(da/dx) = 12(2) =24 </p></li><li><p>This principle scales to much larger computational graphs </p></li></ul><p></p>
22
New cards

What is backpropagation?

  • backpropagation: Efficiently compute the gradient of the loss with respect to model parameters

  • It works backward through the computational graph :

    Start from loss → Compute local derivatives → Apply chain rule → Propagate gradients backward → Obtain a gradient for every trainable parameter


23
New cards

What is the difference between the forward pass and backward pass?

Forward pass:

  • Input → Prediction

  • Compute intermediate activations

  • Produces predictions

  • Compute the loss


Backward pass:

  • Loss → Parameters

  • Compute gradients

  • Determines how sensitive the loss is to the parameters


<p>Forward pass:</p><ul><li><p>Input → Prediction</p></li><li><p>Compute intermediate activations</p></li><li><p>Produces predictions </p></li><li><p>Compute the loss </p></li></ul><p></p><p>Backward pass: </p><ul><li><p>Loss → Parameters</p></li><li><p>Compute gradients</p></li><li><p>Determines how sensitive the loss is to the parameters </p></li></ul><p></p>
24
New cards

What does the gradient with respect to model parameters tell us?

If: theta = (w1,w2,…,wn)

Then: (*form in pict) is a collection of partial derivatives

  • Each component essentially ask:

    “ If this parameter changes slightly, how will the loss change?”

  • The sign indicates direction of sensitivity while the magnitude indicates how sensitive the loss is locally


<p>If: theta = (w1,w2,…,wn)</p><p>Then: (*form in pict) is a collection of partial derivatives </p><ul><li><p>Each component essentially ask:</p><p>“ If this parameter changes slightly, how will the loss change?”</p></li><li><p>The sign indicates direction of sensitivity while the magnitude indicates how sensitive the loss is locally </p></li></ul><p></p>
25
New cards

Is backpropagation the same as gradient descent?

  • NO

  • Backpropagation → Compute gradients

  • Optimisation/ gradient descent→ Use gradients to change parameters

  • So:

    Backpropagation → Gradients → Optimiser → Parameter update

  • Therefore: Backpropagation IS NOT Gradient Descent

  • This distinction is especially important in PyTorch


26
New cards

What is the basic idea of gradient descent?

Once gradients are available :

  1. Calculate the current slope of the loss

  2. Move parameters in a direction that reduces loss

  3. Recalculate

  4. Repeat

    => Learning rate controls the size of each parameter update


27
New cards

How can optimization be understood as searching a loss landscape?

  • Imagine the loss as a landscape:

    • Horizontal axes → model parameters

    • Vertical axes → loss

  • Training starts at an initial parameter location

  • Gradients provide local directional information, and optimization repeatedly moves toward lower-loss regions

  • Real neural networks may contain millions or billions of parameter dimensions, so their loss landscapes normally cannot be directly visualised


<ul><li><p>Imagine the loss as a landscape:</p><ul><li><p>Horizontal axes → model parameters  </p></li><li><p>Vertical axes → loss</p></li></ul></li><li><p>Training starts at an initial parameter location</p></li><li><p>Gradients provide local directional information, and optimization repeatedly moves toward lower-loss regions</p></li><li><p>Real neural networks may contain millions or billions of parameter dimensions, so their loss landscapes normally cannot be directly visualised </p></li></ul><p></p>
28
New cards

What is Stochastic Gradient Descent (SGD) and why is it used?

  • Computing gradients using the entire dataset can be expensive

  • SGD uses subset of the training data to estimate the gradient


Advantages:

  • Lower computation per update

  • More frequent parameter updates

  • Suitable for large datasets


Trade-off(Disadvantage):

  • Gradient estimates bcm noisier


29
New cards

What is mini-batch training, and why is it commonly used?

  • Instead of processing:

    • One example at a time, OR

    • The entire dataset at once

  • DL commonly processes a mini-batch or a subset of examples together

  • It provides a practical balance between:

    • Computational efficiency

    • Memory usage

    • Gradient stability


30
New cards

If a dataset has 50,000 examples and batch size = 100, approximately how many parameter updates occur per epoch?

50,000/100 =500

So approximately 500 parameter updates per epoch


31
New cards

What is the difference between a sample, batch, iteration, and epoch?

  • Sample: One training example

  • Batch/ Mini-batch: Subset of examples processed together

  • Iteration: One forward-backward-update cycle

  • Epoch: One complete pass through the training dataset


Example: 10,000 samples with batch size 100

  • 10,000/100=100 iterations

  • So approximately 100 iterations = 1 epoch


32
New cards

What is the learning rate?

  • The learning rate controls the magnitude of parameter updates

  • Too small:

    • Training progresses slowly

    • Many iterations may be required

  • Too large:

    • Updates may overshoot useful regions

    • Loss may oscillate or diverge

  • Appropriate:

    • Loss decreases at useful and reasonably stable rate


<ul><li><p>The learning rate <strong>controls the magnitude of parameter updates</strong></p></li><li><p>Too small:</p><ul><li><p>Training progresses slowly</p></li><li><p>Many iterations may be required</p></li></ul></li><li><p>Too large:</p><ul><li><p>Updates may overshoot useful regions</p></li><li><p>Loss may oscillate or diverge</p></li></ul></li><li><p>Appropriate:</p><ul><li><p>Loss decreases at useful and reasonably stable rate </p></li></ul></li></ul><p></p>
33
New cards

What is momentum and why is it useful?

  • SGD updates can vary substantially between mini-batches

  • Momentum incorporate information from previous updates to:

    • Build speed in directions with consistent gradients

    • Reduce sensitivity to short-term fluctuations

    • Smooth the optimisation trajectory

  • Analogy: A ball rolling downhill

    • Gradient = slope

    • Momentum = accumulated velocity


<ul><li><p>SGD updates can vary substantially between mini-batches</p></li><li><p>Momentum incorporate information from previous updates to:</p><ul><li><p>Build speed in directions with consistent gradients</p></li><li><p>Reduce sensitivity to short-term fluctuations</p></li><li><p>Smooth the optimisation trajectory</p></li></ul></li><li><p>Analogy: A ball rolling downhill</p><ul><li><p>Gradient = slope</p></li><li><p>Momentum = accumulated velocity</p></li></ul></li></ul><p></p>
34
New cards

What is Adam and what is its main intuition?

  • Adam combines:

    • Momentum-like accumulation

    • Adaptive scaling of parameter updates

  • Therefore:

    • Different parameters can receive differently scaled updates

    • Recent gradient behaviour influences the update

  • Adam is widely used as a general-purpose optimizer

=> Importantly, Adam changes how gradient are used; It does not replace backpropagation



35
New cards

How do Gradient Descent, SGD, Mini-batch SGD, Momentum, and Adam differ?

knowt flashcard image
36
New cards

What are the three key PyTorch operations in a training interation?

  1. forward()

    Compute model predictions and constructs the forward computation

  2. backward()

    Compute gradients using backpropagation

  3. step()

    Use gradients to update parameters


=> forward → loss → backward → step


37
New cards

What happens during forward() in PyTorch?

During forward pass:

  1. Input tensors enter the model

  2. Layers perform transformations

  3. Intermediate results are computed

  4. The model produces predictions

  5. PyTorch records operations needed for later gradient computation


Example: pred = model(x)


<p>During forward pass:</p><ol><li><p>Input tensors enter the model</p></li><li><p>Layers perform transformations</p></li><li><p>Intermediate results are computed</p></li><li><p>The model produces predictions </p></li><li><p>PyTorch records operations needed for later gradient computation</p></li></ol><p></p><p>Example:  pred = model(x)</p><p></p>
38
New cards

What happens when loss.backward() is called?

The backward pass:

  1. Starts from the loss

  2. Traveses the computational graph backward

  3. Applies the chain rule

  4. Computes gradients for trainable parameters

=> Afterward, the parameters have associated gradient information



<p>The backward pass:</p><ol><li><p>Starts from the loss</p></li><li><p>Traveses the computational graph backward</p></li><li><p>Applies the chain rule</p></li><li><p>Computes gradients for trainable parameters </p></li></ol><p>=&gt; Afterward, the parameters have associated gradient information</p><p></p><p></p>
39
New cards

What happens when optimizer.step() is called?

The optimizer:

  1. Reads parameter gradients

  2. Applies its update rule

  3. Modifies the model parameters


=> exact update depends on the optimizer, such as SGD or Adam


<p>The optimizer:</p><ol><li><p>Reads parameter gradients</p></li><li><p>Applies its update rule</p></li><li><p>Modifies the model parameters</p></li></ol><p></p><p>=&gt; exact update depends on the optimizer, such as SGD or Adam</p><p></p>
40
New cards

Why must gradients be reset before the next training iteration in PyTorch?

  • PyTorch accumulates gradients unless they are cleared

  • Therefore: optimizer.zero_grad() → clears the previous gradients bfr calculating new ones

  • Standard iteration:

    Clear gradients → Predictions → Loss → New gradients → Parameter update

  • This ensures each mini-batch contributes the intended gradient information


41
New cards

What is the difference between CPU and GPU for DL?

  1. CPU

  • General-purpose processor

  • Small number of powerful cores

  • Good for flexible comtrol and sequential computation


  1. GPU

  • Large number of parallel computational units

  • Effective for large tensor operations

  • Well suited for matrix-heavy neural networks computation


=> DL benefits heavily from parallel numerical computation which is why GPUs are commonly used


42
New cards

Why must the model and data be on the same device in PyTorch?

The following must be on the same device :

  • Model parameters

  • Input tensors

  • Target tensors


=> A device mismatch prevents computation from proceeding correctly


<p>The following must be on the same device :</p><ul><li><p>Model parameters </p></li><li><p>Input tensors</p></li><li><p>Target tensors </p></li></ul><p></p><p>=&gt; A device mismatch prevents computation from proceeding correctly </p><p></p>
43
New cards
<p>What is the minimal PyTorch training loop, and what does each line mean?</p>

What is the minimal PyTorch training loop, and what does each line mean?

Meaning:

  • for x, y in dataloader → get training data

  • zero_grad() → clear previous gradients

  • model(x) → forward computation

  • loss_fn(pred,y) → compute prediction error

  • loss.backward() → backpropagation / compute gradients

  • optimizer.step() → update parameters

Conceptually:

DATA → MODEL → LOSS → OPTIMIZER

<p><span>Meaning:</span></p><ul><li><p><span>for x, y in dataloader → get training data</span></p></li><li><p><span>zero_grad() → clear previous gradients</span></p></li><li><p><span>model(x) → forward computation</span></p></li><li><p><span>loss_fn(pred,y) → compute prediction error</span></p></li><li><p><span>loss.backward() → backpropagation / compute gradients</span></p></li><li><p><span>optimizer.step() → update parameters</span></p></li></ul><p><span>Conceptually:</span></p><p><span><strong>DATA → MODEL → LOSS → OPTIMIZER</strong></span></p>
44
New cards

Why can neural network training be described as a repeated feedback process?

At iteration t :

Current parameters → Predictions → Loss → Gradients → Parameter update → New parameters → New predictions


So: theta0 → theta1 → theta2 → …

Repeated many times, traing bcms a feedback process driven by predictions error

<p>At iteration <em>t </em>:</p><p>Current parameters → Predictions → Loss → Gradients → Parameter update → New parameters → New predictions </p><p></p><p>So:  theta0 → theta1 → theta2 → …</p><p>Repeated many times, traing bcms a feedback process driven by predictions error</p>
45
New cards

Explain how a neural network learns from data from beginning to end.

  • The model receives input and performs a forward pass to produce a prediction.

  • The loss function compares the prediction with the target.

  • Backpropagation then uses the chain rule to compute how each parameter contributed to the loss.

  • The optimizer uses these gradients to update the parameters toward lower loss.

  • This process repeats over many iterations.


46
New cards

Why do we need both backpropagation and an optimizer? Why can’t one replace the other?

They perform different jobs.

  1. Backpropagation calculates the gradients, telling us how the loss is sensitive to each parameter.

  2. The optimizer uses those gradients to actually change the parameters.

Therefore:

  • Backpropagation = find how to change

  • Optimizer = perform the change


47
New cards

Why is the chain rule essential to training a deep neural network?


  • A deep neural network is a composition of many functions/layers.

  • A parameter in an early layer affects later computations and eventually the loss.

  • The chain rule allows us to combine local derivatives across these operations to determine how an earlier parameter affects the final loss.

  • This is what makes backpropagation possible.


48
New cards

Why don’t we normally compute gradients using the entire dataset for every update?


  • Using the entire dataset can be computationally expensive.

  • Mini-batches allow gradients to be estimated using only part of the data, giving:

    • Lower computation per update

    • More frequent updates

    • Better practicality for large datasets

  • The trade-off is that the gradient estimate becomes noisier.


49
New cards

How are forward pass, loss, backward pass, gradient, and optimizer connected?

Think of the entire training process as:

Input → Forward → Prediction → Loss → Backward → Gradients → Optimizer → Updated Parameters


=> Then the updated parameters are used for the next forward pass, and the cycle repeats