AI and Machine Learning Foundations

0.0(0)
Studied by 0 people
call kaiCall Kai
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/65

flashcard set

Earn XP

Description and Tags

Comprehensive vocabulary flashcards covering foundational machine learning, deep learning, optimization, evaluation metrics, and intelligent agent principles from the lecture review questions.

Last updated 2:22 PM on 10/7/26
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

66 Terms

1
New cards

Machine Learning

The practice of learning patterns from data to make predictions or decisions rather than manually writing every decision rule.

2
New cards

Feature

An input variable in a dataset used by a model to make predictions.

3
New cards

Target Variable

The output variable that a model is intended to predict.

4
New cards

Classification

A supervised learning problem where the model predicts categorical targets or class labels, such as detecting spam versus not spam.

5
New cards

Regression

A supervised learning problem where the target to be predicted is a continuous numeric value, such as house prices.

6
New cards

Model Generalization

The ability of a trained machine learning model to perform well on new, unseen data.

7
New cards

Overfitting

A condition occurring when a model learns noise from the training data and performs poorly on new data.

8
New cards

Underfitting

A condition occurring when a model is too simple to learn the relevant underlying patterns in the data.

9
New cards

Bias-Variance Trade-off

The principle stating that increasing model complexity can reduce bias but increase variance.

10
New cards

Baseline Model

A simple model that provides a reference level of performance to compare against more complex models.

11
New cards

Supervised Learning

A learning approach that requires input data paired with known target values or labels.

12
New cards

Binary Classification

A classification task where the target has two possible classes.

13
New cards

Multiclass Classification

A classification task where there are more than two possible classes, typically assigning one label per example.

14
New cards

Label Noise

The presence of incorrect, inconsistent, or unreliable target labels in a dataset.

15
New cards

Confusion Matrix

A table used primarily to display the counts of true and predicted classes for a classification model.

16
New cards

False Negative

An error in screening or evaluation where a diseased or positive instance is incorrectly predicted as healthy or negative.

17
New cards

Training Set

The subset of data used specifically for learning model parameters.

18
New cards

Validation Set

The subset of data used for selecting hyperparameters and comparing candidate models.

19
New cards

Test Set

A held-out dataset reserved strictly for final, unbiased evaluation after all modeling choices are complete.

20
New cards

Data Leakage

A flaw occurring when information from validation or test data improperly influences model training.

21
New cards

AI Agent

A system that observes its environment and takes actions.

22
New cards

kk-Fold Cross-Validation

A validation technique where the dataset is divided into kk subsets (folds); the model trains on k−1k - 1 folds and validates on the remaining fold across kk iterations, averaging the scores.

23
New cards

Stratified kk-Fold Cross-Validation

A cross-validation strategy preferred in classification with class imbalance to ensure each fold preserves the original class proportions.

24
New cards

Nested Cross-Validation

A validation method especially useful when hyperparameter tuning and unbiased performance estimation are both needed simultaneously.

25
New cards

Gradient Descent

An optimization algorithm whose goal is to adjust model parameters to reduce the loss function.

26
New cards

Learning Rate (η\eta)

A step-size parameter in gradient descent; if set too large, training may overshoot or diverge, and if set too small, training progresses very slowly.

27
New cards

Batch Gradient Descent

A variant of gradient descent that computes parameter gradients using the entire training dataset per update.

28
New cards

Stochastic Gradient Descent (SGD)

A variant of gradient descent that updates model parameters using one training sample at a time.

29
New cards

Mini-Batch Gradient Descent

A gradient descent variant that uses a small subset of training samples per update, balancing computational efficiency and gradient noise.

30
New cards

Local Minimum

A parameter region on a loss surface that is lower than nearby points but not necessarily the globally lowest point.

31
New cards

Backpropagation

An algorithm relying on the calculus chain rule to compute gradients of the loss function with respect to neural network parameters.

32
New cards

Forward Pass

The execution phase of a neural network where input features are transformed into model predictions.

33
New cards

Backward Pass

The execution phase of a neural network where the model computes gradients of the loss with respect to parameters.

34
New cards

Vanishing Gradients

A problem in deep networks where products of many small derivatives approach zero, preventing weights from updating effectively.

35
New cards

Exploding Gradients

A problem in deep networks where products of repeated large derivatives or weight factors grow excessively large.

36
New cards

Gradient Clipping

A training technique used to limit excessively large gradients to prevent instability.

37
New cards

Epoch

One complete pass through the entire training dataset during model training.

38
New cards

Batch

A group of training samples processed together in a single optimization update step.

39
New cards

Activation Function

A function applied to neural network layers to introduce nonlinearity into the model.

40
New cards

ReLU (Rectified Linear Unit)

An activation function defined mathematically as max⁡(0,x)\max(0, x), commonly used in hidden layers of modern deep networks.

41
New cards

Sigmoid Function

An activation function with an output range between 00 and 11, typically appropriate for binary classification outputs.

42
New cards

Hyperbolic Tangent (Tanh)

An activation function with an output range between −1-1 and 11.

43
New cards

Dying ReLU

A failure mode where a neuron permanently outputs 00 and remains inactive for many inputs, often mitigated by using Leaky ReLU.

44
New cards

Softmax

An activation function used in multiclass classification to convert raw scores into a probability distribution over classes.

45
New cards

Linear Regression

A regression model with the general form y=β0+β1x1+⋯+βpxpy = \beta_0 + \beta_1 x_1 + \dots + \beta_p x_p, predicting a continuous target variable.

46
New cards

Residual

The difference between the observed target value and the predicted target value in regression.

47
New cards

Multicollinearity

A condition in regression where predictor variables are highly correlated with one another.

48
New cards

Regularization

A technique used to penalize overly complex or excessively large coefficient values to reduce overfitting.

49
New cards

PEAS

An acronym describing the task environment of an AI agent: Performance measure, Environment, Actuators, and Sensors.

50
New cards

Intelligent Agent Cycle

The core operational loop of an agent: perceive, decide (evaluate/choose action), and act.

51
New cards

Rational Agent

An agent that selects the action expected to maximize its performance measure given its percept history, prior knowledge, and available actions.

52
New cards

Logistic Regression

A classification model that models a linear relationship in log-odds and maps continuous scores to probabilities using the sigmoid function.

53
New cards

Odds Ratio

A statistic describing the change in odds of an event associated with a one-unit change in a predictor variable.

54
New cards

Decision Tree

A model that makes predictions by recursively splitting data according to feature-based rules.

55
New cards

Root Node

The very first split or starting node at the top of a decision tree.

56
New cards

Leaf Node

A terminal node in a decision tree that outputs the final prediction.

57
New cards

Gini Impurity

A metric measuring the degree of class mixture or impurity within a node of a classification tree.

58
New cards

Entropy

A metric used in decision trees to quantify the level of class uncertainty within a node.

59
New cards

Random Forest

An ensemble model consisting of a collection of decision trees whose predictions are aggregated to reduce variance and improve generalization.

60
New cards

Bagging (Bootstrap Aggregating)

An ensemble technique that trains multiple models independently on bootstrap samples of the data and aggregates their predictions.

61
New cards

Hard Voting

An ensemble aggregation technique in classification that selects the class receiving the majority of individual model votes.

62
New cards

Support Vector Machine (SVM)

A classification algorithm whose primary objective is to maximize the margin between classes.

63
New cards

Support Vectors

The training points located closest to the decision boundary that directly define and influence the boundary.

64
New cards

SVM Parameter CC

A hyperparameter controlling the trade-off between maximizing the margin width and minimizing classification violations.

65
New cards

Kernel Trick

A mathematical technique in SVMs that computes inner products in high-dimensional feature spaces implicitly using a kernel function.

66
New cards

Exploration vs. Exploitation

The trade-off in an agent between trying new actions to gather information for better future decisions (exploration) and choosing the best currently known option (exploitation).