1/73
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Machine Learning
The field of study that gives computers the ability to learn without being explicitly programmed.
Training Set
The examples used to teach/learn the model.
Training Instance
One example in the training set.
Label
The desired output for a training instance.
Feature
An input variable/attribute used by the model to make a prediction.
Inference
Using a trained model to make predictions on new data.
Supervised Learning
Learning from labeled training data.
Unsupervised Learning
Learning from data without labeled outputs.
Reinforcement Learning
An agent observes an environment, takes actions, and receives rewards or penalties to learn a policy that maximizes cumulative rewards
Classification
Predicting a discrete/categorical class.
Regression
Predicting a numerical/continuous value.
Instance-Based Learning
A method that learns examples and uses similarity to make predictions on new instances.
Model-Based Learning
A method that builds a model of the data and uses that model to make predictions.
Generalization
A model's ability to perform well on new, unseen data.
Sampling Bias
When the training data is not representative of the population/data the model will encounter.
Overfitting
When a model fits the training data too closely and performs poorly on new data
Underfitting
When a model is too simple to capture the patterns in the data.
Regularization
Constraining a model to make it simpler and reduce overfitting.
Model Parameter
A value learned by the model from the training data.
Hyperparameter
A value set by the practitioner rather than learned directly from the training data.
Training Set
Data used to train the model.
Validation Set
Held-out data used to evaluate/tune models and hyperparameters during development.
Test Set
Data held aside for the final evaluation of the model.
Data Snooping Bias
Bias that occurs when information from the test set influences model development.
Why never tune on the test set?
It can make the model appear better than it actually is on unseen data.
RMSE (Root Mean Square Error)
Measures prediction error by taking the square root of the average squared errors.

MAE (Mean Absolute Error)
Measures the average absolute prediction error.

Pearson Correlation Coefficient
Measures the linear relationship between two variables; ranges from −1 to +1.
Feature Engineering
Creating or transforming features to make them more useful for a model.
Missing Values
Data entries that are absent; can be handled by dropping rows/attributes or imputing values.
Imputation
Replacing missing values with an estimated value, such as the median.
One-Hot Encoding
Converts a categorical attribute with \(k\) categories into \(k\) binary attributes, with only one being 1 at a time.
Min-Max Scaling
Rescales values to a specified range, commonly [0,1].
![<p><span>Rescales values to a specified range, commonly [0,1].</span></p>](https://assets.knowt.com/user-attachments/b2f817e5-1ebb-40e1-b3d8-348407fa5997.png)
Standardization
Rescales data to have approximately mean 0 and standard deviation 1.

Transformer
An object that learns how to transform data; uses fit() and transform().
Estimator
A model/object that learns from data using fit().
Predictor
An estimator that can make predictions using predict().
Pipeline
A sequence of data-processing steps chained together.
K-Fold Cross-Validation
Splits training data into \(k\) folds and repeatedly trains/evaluates using different folds as validation data.
Why fit transformations only on training data?
To prevent information from the test set from leaking into the model.
Binary Classification
Classification between exactly two classes
Multiclass Classification
Classification involving more than two classes
One-vs-Rest (OvR)
Creates one classifier for each class, where that class is positive and all other classes are negative.

One-vs-One (OvO)
Creates a classifier for every pair of classes.

Confusion Matrix
A table showing actual vs. predicted classifications

True Positive (TP)
Actual positive, predicted positive.
True Negative (TN)
Actual negative, predicted negative.
False Positive (FP)
Actual negative, predicted positive.
False Negative (FN)
Actual positive, predicted negative.
Accuracy
Fraction of all predictions that are correct.

Precision
Of the examples predicted positive, how many were actually positive?

Recall
Of the actual positives, how many did the model correctly identify?

F1 Score
Harmonic mean of precision and recall.

Decision Score
A numerical score produced by a classifier that can be compared with a threshold.
Threshold
The cutoff used to determine whether a prediction is classified as positive or negative.
Precision/Recall Trade-Off
Increasing the threshold generally increases precision but decreases recall; lowering it generally increases recall but decreases precision.
ROC Curve
A plot of True Positive Rate (recall) against False Positive Rate at different thresholds.
False Positive Rate (FPR)
Fraction of actual negatives incorrectly classified as positive.

AUC
Area Under the ROC Curve; measures overall classifier performance across thresholds.
AUC = 1 → perfect
AUC ≈ 0.5 → random
PR Curve
Precision-Recall curve; often more informative than ROC when positive examples are rare.
Linear Regression
A model that predicts a numerical output as a linear combination of input features.
Parameter Vector θ
The vector containing the coefficients/parameters of the linear regression model.
Vector Form of Linear Regression
y^ = θTx
Matrix Form
y^ = xθ
MSE (Mean Squared Error)
Average of the squared prediction errors.

Cost Function
A function measuring how poorly a model performs; training seeks to minimize it.
Normal Equation
A closed-form method for finding the parameters that minimize linear regression MSE.

Gradient Descent
An optimization method that repeatedly moves parameters in the direction that decreases the cost function.
Gradient
The direction and rate of steepest increase of the cost function.
Learning Rate η
Controls the size of each Gradient Descent step.
Gradient Descent Update Rule

Learning Rate Too Small
Gradient Descent converges very slowly.
Learning Rate Too Large
The algorithm can overshoot the minimum or diverge
Stopping Criterion
A condition used to decide when Gradient Descent should stop, such as when the gradient becomes sufficiently small.