1/65
Comprehensive vocabulary flashcards covering foundational machine learning, deep learning, optimization, evaluation metrics, and intelligent agent principles from the lecture review questions.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Machine Learning
The practice of learning patterns from data to make predictions or decisions rather than manually writing every decision rule.
Feature
An input variable in a dataset used by a model to make predictions.
Target Variable
The output variable that a model is intended to predict.
Classification
A supervised learning problem where the model predicts categorical targets or class labels, such as detecting spam versus not spam.
Regression
A supervised learning problem where the target to be predicted is a continuous numeric value, such as house prices.
Model Generalization
The ability of a trained machine learning model to perform well on new, unseen data.
Overfitting
A condition occurring when a model learns noise from the training data and performs poorly on new data.
Underfitting
A condition occurring when a model is too simple to learn the relevant underlying patterns in the data.
Bias-Variance Trade-off
The principle stating that increasing model complexity can reduce bias but increase variance.
Baseline Model
A simple model that provides a reference level of performance to compare against more complex models.
Supervised Learning
A learning approach that requires input data paired with known target values or labels.
Binary Classification
A classification task where the target has two possible classes.
Multiclass Classification
A classification task where there are more than two possible classes, typically assigning one label per example.
Label Noise
The presence of incorrect, inconsistent, or unreliable target labels in a dataset.
Confusion Matrix
A table used primarily to display the counts of true and predicted classes for a classification model.
False Negative
An error in screening or evaluation where a diseased or positive instance is incorrectly predicted as healthy or negative.
Training Set
The subset of data used specifically for learning model parameters.
Validation Set
The subset of data used for selecting hyperparameters and comparing candidate models.
Test Set
A held-out dataset reserved strictly for final, unbiased evaluation after all modeling choices are complete.
Data Leakage
A flaw occurring when information from validation or test data improperly influences model training.
AI Agent
A system that observes its environment and takes actions.
k-Fold Cross-Validation
A validation technique where the dataset is divided into k subsets (folds); the model trains on k−1 folds and validates on the remaining fold across k iterations, averaging the scores.
Stratified k-Fold Cross-Validation
A cross-validation strategy preferred in classification with class imbalance to ensure each fold preserves the original class proportions.
Nested Cross-Validation
A validation method especially useful when hyperparameter tuning and unbiased performance estimation are both needed simultaneously.
Gradient Descent
An optimization algorithm whose goal is to adjust model parameters to reduce the loss function.
Learning Rate (η)
A step-size parameter in gradient descent; if set too large, training may overshoot or diverge, and if set too small, training progresses very slowly.
Batch Gradient Descent
A variant of gradient descent that computes parameter gradients using the entire training dataset per update.
Stochastic Gradient Descent (SGD)
A variant of gradient descent that updates model parameters using one training sample at a time.
Mini-Batch Gradient Descent
A gradient descent variant that uses a small subset of training samples per update, balancing computational efficiency and gradient noise.
Local Minimum
A parameter region on a loss surface that is lower than nearby points but not necessarily the globally lowest point.
Backpropagation
An algorithm relying on the calculus chain rule to compute gradients of the loss function with respect to neural network parameters.
Forward Pass
The execution phase of a neural network where input features are transformed into model predictions.
Backward Pass
The execution phase of a neural network where the model computes gradients of the loss with respect to parameters.
Vanishing Gradients
A problem in deep networks where products of many small derivatives approach zero, preventing weights from updating effectively.
Exploding Gradients
A problem in deep networks where products of repeated large derivatives or weight factors grow excessively large.
Gradient Clipping
A training technique used to limit excessively large gradients to prevent instability.
Epoch
One complete pass through the entire training dataset during model training.
Batch
A group of training samples processed together in a single optimization update step.
Activation Function
A function applied to neural network layers to introduce nonlinearity into the model.
ReLU (Rectified Linear Unit)
An activation function defined mathematically as max(0,x), commonly used in hidden layers of modern deep networks.
Sigmoid Function
An activation function with an output range between 0 and 1, typically appropriate for binary classification outputs.
Hyperbolic Tangent (Tanh)
An activation function with an output range between −1 and 1.
Dying ReLU
A failure mode where a neuron permanently outputs 0 and remains inactive for many inputs, often mitigated by using Leaky ReLU.
Softmax
An activation function used in multiclass classification to convert raw scores into a probability distribution over classes.
Linear Regression
A regression model with the general form y=β0+β1x1+⋯+βpxp, predicting a continuous target variable.
Residual
The difference between the observed target value and the predicted target value in regression.
Multicollinearity
A condition in regression where predictor variables are highly correlated with one another.
Regularization
A technique used to penalize overly complex or excessively large coefficient values to reduce overfitting.
PEAS
An acronym describing the task environment of an AI agent: Performance measure, Environment, Actuators, and Sensors.
Intelligent Agent Cycle
The core operational loop of an agent: perceive, decide (evaluate/choose action), and act.
Rational Agent
An agent that selects the action expected to maximize its performance measure given its percept history, prior knowledge, and available actions.
Logistic Regression
A classification model that models a linear relationship in log-odds and maps continuous scores to probabilities using the sigmoid function.
Odds Ratio
A statistic describing the change in odds of an event associated with a one-unit change in a predictor variable.
Decision Tree
A model that makes predictions by recursively splitting data according to feature-based rules.
Root Node
The very first split or starting node at the top of a decision tree.
Leaf Node
A terminal node in a decision tree that outputs the final prediction.
Gini Impurity
A metric measuring the degree of class mixture or impurity within a node of a classification tree.
Entropy
A metric used in decision trees to quantify the level of class uncertainty within a node.
Random Forest
An ensemble model consisting of a collection of decision trees whose predictions are aggregated to reduce variance and improve generalization.
Bagging (Bootstrap Aggregating)
An ensemble technique that trains multiple models independently on bootstrap samples of the data and aggregates their predictions.
Hard Voting
An ensemble aggregation technique in classification that selects the class receiving the majority of individual model votes.
Support Vector Machine (SVM)
A classification algorithm whose primary objective is to maximize the margin between classes.
Support Vectors
The training points located closest to the decision boundary that directly define and influence the boundary.
SVM Parameter C
A hyperparameter controlling the trade-off between maximizing the margin width and minimizing classification violations.
Kernel Trick
A mathematical technique in SVMs that computes inner products in high-dimensional feature spaces implicitly using a kernel function.
Exploration vs. Exploitation
The trade-off in an agent between trying new actions to gather information for better future decisions (exploration) and choosing the best currently known option (exploitation).