1/25
Vocabulary flashcards covering topics from Machine Learning and Deep Neural Networks lectures, including Decision Trees, Random Forests, Perceptrons, Backpropagation, CNNs, ResNet, Autoencoders, Embeddings, and Reinforcement Learning.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress

Decision Tree Structure
The structural anatomy of a decision tree model, consisting of a root node at the top, internal nodes representing questions/tests, branches representing answers, and leaf nodes representing final output decisions or classifications.
Gini Impurity
A measure of node impurity used in decision tree construction, calculated for binary outcomes as 1−(P(Yes))2−(P(No))2, where a pure node equals 0 and maximum impurity equals 0.5.
Random Forest
An ensemble learning method that improves decision tree robustness by building many small, independent trees and determining predictions through majority voting.
Bagging (Bootstrap Aggregating)
An ensemble technique used in random forests where each decision tree is trained on a random subset of features and data points to reduce variance and prevent overfitting.
Boosting
An iterative ensemble method that builds decision trees sequentially by reweighting misclassified training samples to focus on data points that are initially hard to classify.

Interpretability vs. Accuracy Trade-off
The trade-off in machine learning where simpler models like linear regression and decision trees offer higher interpretability but lower accuracy, whereas complex models like random forests and neural networks achieve higher accuracy at the cost of lower interpretability.
Perceptron
An artificial neuron model, described by Warren McCulloch and Walter Pitts (1943) and implemented by Frank Rosenblatt (1957), that functions as a simple binary linear classifier by computing a weighted sum of inputs and applying a threshold or activation function.
Sigmoid Function
A smooth, differentiable activation function defined as σ(x)=1+e−x1, mapping input values to outputs in the range (0,1).
ReLU (Rectified Linear Unit)
A computationally efficient activation function defined as f(x)=0 if x≤0 and f(x)=x if x>0.
XOR Problem (Perceptron Limit)
The fundamental limitation that a single perceptron cannot implement the XOR Boolean function because the outputs for XOR are not linearly separable.
Hornik's Universal Approximation Theorem
A mathematical theorem establishing that feedforward neural networks with even a single hidden layer can approximate any bounded continuous mathematical function.
Backpropagation
An algorithm for training neural networks that applies the calculus chain rule to compute partial derivatives of a loss function with respect to weights and biases during a backward pass.
Dropout
A regularization technique applied only during neural network training where a random subset of neurons is deactivated by setting their weights to zero with probability p.
Shallow vs. Deep Neural Networks
The distinction between networks with a single hidden layer (such as word2vec) and deep neural networks with tens or hundreds of hidden layers (such as CNNs and transformers).
MNIST Database
A classic machine learning benchmark dataset containing 70,000 grayscale images (60,000 training / 10,000 testing) of handwritten digits sized 28×28=784 pixels.
Convolutional Neural Network (CNN)
A deep neural network architecture designed for grid-structured data like images, which uses sliding learned kernels to extract local spatial features.
Pooling Layer
A operation in a convolutional neural network that reduces the spatial dimensions of feature maps by aggregating information within local neighborhoods.
AlexNet
A landmark 8-layer deep convolutional neural network with 60 million parameters and 650,000 neurons that utilized GPUs for parallel training and won the 2012 ILSVRC competition.
Vanishing / Exploding Gradient Problem
A difficulty in training deep neural networks where partial derivatives amplify or shrink exponentially as they propagate backward through many layers.

Residual Architecture (ResNet)
A neural network design that uses skip connections (residual connections) bypassing layers to mitigate vanishing gradients, enabling the successful training of models with up to 152 layers.

Autoencoder Architecture
An unsupervised neural network consisting of an encoder and a decoder separated by a narrow bottleneck layer, designed to learn compact, low-dimensional representations of data.
Denoising Autoencoder
An autoencoder variant trained by adding synthetic noise (such as Gaussian or salt-and-pepper noise) to the input while forcing the network to reconstruct the original clean data.
Word Embeddings
Dense vector representations of words or tokens in a continuous high-dimensional space where semantically similar words map to nearby coordinates.
Reinforcement Learning
A decision-making framework modeled as a Markov Decision Process (MDP) where an agent interacts with an environment to learn an optimal policy that maximizes long-term cumulative reward.
Q-Learning
A model-free reinforcement learning algorithm that estimates a value function Qπ(s,a), representing the expected cumulative future reward of taking action a in state s.
Deep Q-Learning (DQN)
A reinforcement learning method that combines Q-learning with deep neural networks (such as CNNs) to approximate Q-values in environments with high-dimensional state spaces.