Data Mining (DATS 2103): Classification

0.0(0)
Studied by 0 people
call kaiCall Kai
Locked
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/48

encourage image

There's no tags or description

Looks like no tags are added yet.

Last updated 3:24 AM on 9/9/26
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

49 Terms

1
New cards

Classification

Predicts a qualitative response

2
New cards

Logistic Regression

Models the probability that Y belongs to a particular category

3
New cards

Why not Linear Regression?

Linear regression may produce values outside the [0

4
New cards

Logistic Function

Function used as the basis for logistic regression

5
New cards

Prediction Threshold

Often predict Y=1 when p(X) = Pr(Y=1|X) > 0.5

6
New cards

Odds

p/(1-p)

7
New cards

Log Odds (Logit)

The logarithm of the odds; has a linear relationship with the predictors

8
New cards

Logistic Regression Coefficients

Estimated using Maximum Likelihood Estimation (MLE)

9
New cards

Maximum Likelihood Estimation (MLE)

Selects coefficients so the predicted probability is as close as possible to the individual's observed outcome

10
New cards

Multiple Logistic Regression

Logistic regression using multiple predictors

11
New cards

Multinomial Logistic Regression

Logistic regression that allows more than two response classes (K > 2)

12
New cards

Multinomial Logistic Regression Methods

Run K−1 independent binary models or use softmax coding to run one model

13
New cards

Multinomial Logistic Regression Assumption

Odds of preferring Class A vs Class B are not dependent on the presence of other classes

14
New cards

Predictor Independence in Multinomial Logistic Regression

Predictors do not necessarily need to be independent

15
New cards

Baseline Class

A class selected as the reference when running K−1 independent binary models

16
New cards

Multinomial Coefficient Interpretation

Depends on the two classes being compared

17
New cards

Softmax Coding

Runs one multinomial model with no baseline; all K classes are treated as symmetric

18
New cards

Generative Models for Classification

Models the distribution of the predictors X separately within each response class using Bayes' Theorem

19
New cards

Why Use Generative Models?

They can be more accurate when classes are far apart

20
New cards

Prior Probability (πk)

Probability that a randomly chosen observation comes from the kth class

21
New cards

fk(x)

Density function of X for an observation that comes from the kth class

22
New cards

Bayes' Theorem

Used to calculate posterior probabilities for classification

23
New cards

Linear Discriminant Analysis (LDA)

Generative classification method that assumes predictors follow a Gaussian distribution and produces a linear decision boundary

24
New cards

LDA for p=1

Assumes X is normally or Gaussian distributed

25
New cards

LDA for Multiple Predictors

Assumes X follows a multivariate Gaussian distribution

26
New cards

LDA Parameters

Estimates prior probability

27
New cards

LDA Classification

Calculate the discriminant for each class and assign the observation to the class with the highest discriminant

28
New cards

LDA Decision Boundary

The point where the discriminants for two classes are equivalent

29
New cards

LDA Boundary

Linear

30
New cards

LDA Sample Size

Performs best with fewer observations because reducing variance is important

31
New cards

Quadratic Discriminant Analysis (QDA)

Similar to LDA but uses class-specific variance instead of a common variance

32
New cards

QDA Boundary

Quadratic

33
New cards

QDA Parameters

Estimates prior probability

34
New cards

QDA Sample Size

Performs best with large amounts of observations because reducing variance is less important

35
New cards

LDA vs QDA

LDA uses a common variance and linear boundary; QDA uses class-specific variance and quadratic boundary

36
New cards

Naive Bayes Classifier

Classification method that assumes predictors are independent within classes

37
New cards

Naive Bayes Assumption

Predictors are independent within classes

38
New cards

Quantitative Predictors in Naive Bayes

Can assume each predictor comes from a normal distribution within each class or use a non-parametric estimate

39
New cards

Histogram Estimation

Estimates the density as the fraction of training observations in a class that fall into the same histogram bin as the predictor value

40
New cards

Kernel Density Estimation

A smoothed version of histogram-based density estimation

41
New cards

Qualitative Predictors in Naive Bayes

Use the proportion of training observations corresponding to each predictor value within each class

42
New cards

K-Nearest Neighbors (KNN)

Non-parametric classification method that assigns a prediction based on the K nearest training observations

43
New cards

KNN Process

Choose K

44
New cards

KNN Sample Size

Requires many more observations than the other methods

45
New cards

KNN Coefficients

Does not report coefficients

46
New cards

Linear Decision Boundary

LDA and Logistic Regression perform well

47
New cards

Moderately Nonlinear Decision Boundary

QDA or Naive Bayes may perform better

48
New cards

Highly Complicated Decision Boundary

A non-parametric method such as KNN may be superior

49
New cards

KNN Smoothness

The level of smoothness for a non-parametric approach must be chosen carefully