1/237
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Steps of KKN
- Choose K
- Compute distance from test point to all training points
- Sort distances
- Pick k closest
- Majority vote
What is K in KKN?
The number of neighbors used in voting
In KNN what happens if one feature has much larger values than others?
The large value dominates the distance
PCA
The technique in KNN is used to reduce dimensionality
Euclidean Distance Formula
d = √((x2 - x1)^2 + (y2 - y1)^2)
Manhattan Distance Formula
| x1-x2 | + | y1-y2 |
Hamming Distance
Counting the different binary positions
In the Minkowski Distance formula what value give Manhattan Distance?
1
In the Minkowski Distance formula what value give Euclidean Distance?
2
What property is not required for Minkowski Distance metric?
Linearity
Effect of a Small K in KKN
Overfitting
Effect of a large K in KKN
Underfitting
KKN's Main Flaw
Curse of Dimensionality
Curse of Dimensionality
High-dimensional data requires more samples
Naive Bayes
Probabilistic classifier that calculates likelihood of each class and chooses the most likely
Naive Bayes Steps
- Compute prior P(c)
- Compute likelihoods
- Multiply
- Choose max
What does Naive Bayes produce?
Multiple probabilities of classes
Naive Bayes Assumption
Independent Features
What happens if Naive Bayes' assumption is violated?
Accuracy may decrease
What happens if Naive Bayes encounters a feature value not seen in training?
Probability becomes zero
What happens if a probability becomes zero in Naive Bayes?
Entire product becomes zero
NB Normalized Term
P(X)
NB Prior Class
P(Y)
NB Likelihood
P(X|Y)
NB Posterior
P(Y|X)
Laplace Smoothing
Technique to handle zero probabilities in classification
Support Vector Machines (SVM)'s Purpose
Finding the best separating hyperplane / margin between classification classes in a model
SVM Decision Boundary Formula
(w^T)x + b = 0
Support Vectors
Data points closest to the decision boundary that determine the margin
What are support vectors ins SVM?
Closest points to boundary that determine the hyperplane
SVM Margin
The distance between the boundary and closest points
Small C in SVM
- Simpler Decision Boundary
- Larger Margin
- High bias, low variance
- Risks underfitting
Large C in SVM
- Penalizes error heavily
- Smaller Margin
- Lower bias, high variance
- Risks overfitting
Margin Formula
2/||w||
SVM Optimization Formulas
- Hard Margin
- Soft Margin
- Hinge Loss
Hinge Loss
Penalizes points that are inside margin or have been misclassified
SVM C Variable
Margin vs misclassification tradeoff
SVM Kernel Trick Purpose
Transforms data to higher dimensions
What happens when features are not scaled in SVM?
Margin becomes skewed
What is the expected outcome of a large Gamma(Y) In Guassian RBF?
Overfitting & Wiggly Boundary
What is the expected outcome of a small Gamma(Y) In Guassian RBF?
Underfitting & Smooth boundaries
What are the assumptions when using Gaussian Naive Bayes?
- Independent Features
- Continuous Features
- Normal Distribution
What are the required parameters of GNB
- mean
- variance
What type of algorithm is SOFTMAX Regression?
Multi-Class Classification
What algorithm is SOFTMAX Regression a version of?
Logistic Regression
What does Softmax Regression Produce?
One probability value per class
Range of Softmax Regression
Sum to 1
What does SOFTMAX produce?
- Outputs from range sum to 1
- Probabilities
What is the assumption when using SOFTMAX?
Classes are mutually exclusive
How to calculate total parameters in SOFTMAX Regression?
(features * classes) + classes
In softmax, what happens if one logit is much larger?
That class gets probability ≈ 1
What models are most sensitive to unscaled features?
- KNN
- SVM
Decision Tree Model
Supervised learning model that makes predictions by splitting data into branches based on feature values, forming a tree of decisions that leads to a final output
What is the purpose of splitting node in a Decision Tree?
To reduce impurity of child nodes
What is impurity in a Decision Tree's node?
Different classes within the same node
What are the impurity measures in a Decision Tree?
- Gini Impurity
- Entropy
What is the faster impurity measure to compute in Decision Trees?
Gini
What is the relationship between Decision Trees and Scaling?
Scaling is not required
Information Gain
Reduction in impurity after a split in a Decision Tree
Information Gain Formula
(Root Entropy) - (Weighted Entropy After Split)
What does increasing max_depth within reason do?
It helps the decision tree memorize training data and reduced underfitting
C4.5
Decision tree algorithm that extends ID3 and can handle continous attributes
CART
Decision tree algorithm that constructs binary trees via binary splits
ID3
Decision tree algorithm used for classification
What happens when max_depth is increased too much?
The decision tree memorizes too much training data/overtrains causing overfitting
What does it mean to prune a Decision Tree?
To remove unnecessary branches to improve generalization and reduce overfitting
min_samples_split
Parameter that controls splitting threshold and controls minimum samples
Entropy Increases, Information Gain __________
Decreases
Entropy Decrease, Information Gain __________
Increases
Where does max entropy occur?
p = 0.5
Max Entropy
1
What condition leads to 0 entropy?
A single-class / pure node
What expression represents Gini impurity for classes with probabilities p_i?
1 - Σ p_i^2
What do we look for to determine which entropy value is better?
The lowest value, the lower the value the higher the information gain
Steps of a Neural Network
- Take inputs
- Apply functions
- Produce an output
What Neural Network libraries are used in this course?
- Scikit-learn (MLPClassifier)
- TensorFlow
- Keras
Ensemble Learning
The combination of multiple models to improve overall performance by reducing error
Decision Tree
A single model trained on a full dataset prone to high variance/overfitting
Random Forest
A multi-model/ensemble of decision trees
Voting Classifier
Combines the class predictions from multiple models
Types of Voting Classifier
- Hard Voting
- Soft Voting
Hard Voting
Based on majority class
Soft Voting
Based on the nearest averaged predicted probabilities and selected the highest
What does Bagging/Pasting do?
They create and train multiple models independently and in parallel using subsets of data
Bagging
- Samples with replacement- Duplicates exist within the model- Uses Out-of-Bag Evaluation
Pasting
- Samples without replacement- No duplicates within the model
Out-of-Bag Evaluation
Takes the samples left out during the process of bagging and uses them as a validation set
Boosting
Models trained sequentially each focused on previous errors to reduce bias
Centroid
The average of all point in a cluster
Cluster's Job
Groups similar data points of unlabeled data & discovers structure/patterns
Dendrogram
A tree of clusters
Hierarchical Clustering
Buildings dendrogram in which each node represents a cluster of clusters
Agglomerative Clustering
Uses a proximity matrix and a distance formula to merge points together
DBSCAN
A clustering alg that groups based on density instead of distance
Pros of DBSCAN
- Finds arbitrary shapes
- Handles noise and outliers well
What are the challenges of Natural Language Processing?
- Human language is ambiguous, context-dependent, & unstructured
- Text is messy & doesn't fit tables
- Requires understanding meaning not just keywords
Bag of Words
Creates an occurrence matrix that tracks frequency of words, but ignores grammar and order
Limits of Bag of Words
- Does not preserve context
- Dominated by common words
- Understands semantics poorly
TF-IDF
Creates an occurrence matrix that tracks frequency of words but penalizes common words
Tokenization
- Splits text into words/tokens
- Removes punctuation