1/81
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Large language model (LLM)
An AI model trained on text that predicts the next token; its predictions can generate language.
Token
A unit of text processed by a language model, such as a word, part of a word, or punctuation.
Pre-training
Initial training on a large text collection to learn patterns and next-token prediction.
Fine-tuning
Further training a pretrained model on examples to adjust its behavior or output.
Prompt
The input sent to a language model.
System/developer role
Instructions that set the model's overall rules or behavior.
User role
The user's message to the model; the full input can include material beyond the visible text box.
Assistant role
The model's response in a conversation.
Machine learning algorithm
A procedure that learns a model from data, such as linear regression.
Model
The learned rules or mathematical relationship produced by a machine learning algorithm.
Training
Providing data to an algorithm so it can build a model.
Testing (verification)
Evaluating a trained model on independent data to estimate its performance.
Deployment (inference)
Using a trained model to make predictions on new inputs.
Supervised learning
Learning from examples whose true outputs are known.
Unsupervised learning
Learning patterns from data without known target outputs.
Bias
Error introduced by a model's simplifying assumptions; how far its predictions are from the truth on average.
Underfitting
When a model is too simple to capture the underlying relationship in the data.
Variance
How much a model's predictions change when it is trained on a different sample of data.
Overfitting
When a model fits training data too closely and performs poorly on new data.
Data leakage
When information that should be unavailable during training or selection influences the model, making evaluation overly optimistic.
Decision tree regression
A nonparametric supervised regression method that uses branching rules to make piecewise constant predictions.
Independent variable (feature)
An input or predictor used to predict the target variable.
Dependent variable (target)
The output a model aims to predict; in logistic regression it is categorical.
Sigmoid (logistic) function
A function that maps any linear score to a value between 0 and 1, used as a probability in logistic regression.
Odds
The probability of an event divided by the probability that it does not occur: p/(1-p).
Log-odds (logit)
The natural logarithm of the odds; logistic regression models it as a linear combination of the input features.
Coefficient
An estimated model parameter describing how an input feature changes the modeled log-odds, holding other features fixed.
Intercept
The constant term in logistic regression: the modeled log-odds when all input features equal zero.
Maximum likelihood estimation (MLE)
A method for estimating model parameters by maximizing the likelihood of the observed data.
Decision threshold
The probability cutoff used to turn a predicted probability into a class label.
Confusion matrix
A table comparing predicted and actual classes, showing true positives, true negatives, false positives, and false negatives for binary classification.
Precision
Among predicted positives, the fraction actually positive: TP/(TP+FP).
Recall (sensitivity)
Among actual positives, the fraction correctly identified: TP/(TP+FN).
F1 score
The harmonic mean of precision and recall: 2PR/(P+R).
Accuracy
The fraction of all examples classified correctly.
Multiclass classification
Classification into more than two categories.
Hyperplane
A decision boundary in feature space; for a linear classifier it can be written w·x+b=0.
Support vector
A training point closest to an SVM decision boundary that helps determine the boundary and margin.
Margin
The distance from an SVM decision boundary to its nearest training points; SVM seeks a large one.
Kernel
A function that lets an SVM represent a nonlinear decision boundary by computing similarity in a transformed feature space.
Hard margin
An SVM boundary that maximizes the margin while requiring perfect separation of the training data.
Soft margin
An SVM boundary that permits margin violations or misclassification in exchange for a better tradeoff between margin size and errors.
C (SVM)
The regularization parameter controlling the penalty for SVM margin violations; higher ___ penalizes them more strongly.
Hinge loss
A loss function that penalizes misclassified examples and points inside an SVM margin.
Cross-validation
A model evaluation method that repeatedly trains on part of the available data and validates on another part.
K-fold cross-validation
Divide data into k folds; train on k-1 folds and validate on the remaining fold, repeating until each fold has been used for validation.
Holdout
A single division of data into training and evaluation sets.
Grid search
Testing specified hyperparameter combinations, typically using cross-validation, to select a model configuration.
Agglomerative clustering
Bottom-up hierarchical clustering that begins with one cluster per point and repeatedly merges nearby clusters.
Hierarchical clustering
Clustering that builds a nested tree of groups at multiple levels; it may merge smaller groups or split larger ones.
Linkage
The rule used to measure distance between clusters and choose which clusters to merge.
Distance matrix
A symmetric table of pairwise distances between data points.
Dendrogram
A tree diagram showing hierarchical cluster merges and the distances at which they occur.
Ward linkage
A linkage method that chooses merges minimizing the increase in within-cluster sum of squares; it uses Euclidean distance.
Complete linkage (maximum linkage)
Defines distance between two clusters as the greatest distance between any point in one and any point in the other.
Average linkage (UPGMA)
Defines distance between two clusters as the average distance across all pairs of points in the two clusters.
Single linkage (minimum linkage)
Defines distance between two clusters as the shortest distance between any point in one and any point in the other.
Silhouette score
The average silhouette coefficient across points; it summarizes how well clusters are separated, from -1 to 1.
Silhouette coefficient
For one point, a measure comparing its average distance to points in its own cluster with its distance to the nearest other cluster.
Within-cluster sum of squares (WCSS)
The sum of squared distances from points to their cluster centers; lower values indicate more compact clusters.
Distance threshold
A cutoff on dendrogram merge distance used to form clusters by cutting the hierarchy at that height.
Fixed number of clusters (maxclust)
A stopping choice that cuts the hierarchy to produce a specified number of clusters.
Inconsistency criterion
A way to cut a hierarchy based on how unusual each merge height is relative to nearby merges.
Elbow method
A method for looking for a sharp change in merge distances or another clustering measure to choose the number of clusters.
Euclidean distance
Straight-line distance between points, equal to the square root of the sum of squared coordinate differences.
Linkage matrix
A table recording each hierarchical merge, its distance, and the size of the new cluster.
Centroid
The center of a cluster, usually the mean of its members' coordinates.
Chaining effect
In single linkage, a tendency to connect points through short links into long, thin clusters.
Inconsistency coefficient
A measure comparing a merge height with nearby merge heights in a hierarchy.
Uniform data
Data distributed without a clear natural cluster structure.
Non-uniform data
Data with cluster structures that differ in shape, size, or density.
Cluster compactness
How closely points within a cluster are grouped.
Cluster separation
How far apart distinct clusters are, or how clearly their boundaries are separated.
data
raw content or the fundamental piece of information
metadata
the information about the data, providing context,
structure, and administrative details
private data
data that is hidden: Some Farmer Data, Proprietary Data, Human Subjects
Research Data
protected data
data that is still hidden but can be released in certain ways
public data
Publicly funded research ___, "Generic" and accessible
categorical data
data in groups that have no inherent numerical order and aren’t able to be assigned with numbers alone
ordinal data
data that is words with meanings that are not based on their length
autoregressive generation
the output is generated step by step by tacking (appending) output to the end
diffusion generation
generation by successive refinement