the model based perspective

0.0(0)
Studied by 0 people
call kaiCall Kai
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/19

encourage image

There's no tags or description

Looks like no tags are added yet.

Last updated 12:34 AM on 10/1/26
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

20 Terms

1
New cards

model

simplified representation of a complex object

2
New cards

statistical model

a simplified representation of a data generating process (DGP)

data = model + error

the model is the way to represent predictable part of data, the goal is to minimize error (max predictability of data via model)

3
New cards

loss function

how much information is lost between model and data

4
New cards

three common uses of statistical models

  1. describe

  2. predict

  3. decide


5
New cards

describe

statistical models can help us understand data by summarizing variables or mapping relationships between different variables

6
New cards

predict

a statistical model can help us anticipate how we expect new observations, which were not in our original sample that we used to estimate the model, to look

7
New cards

decide

a statistical model can guide how we choose to behave by helping us understand the likelihood of different consequences of our choices

8
New cards

squared error

most common loss function

minimized by the mean

might try to minimize the sum of squared errors: SSE = Σ(x-x̄)2

9
New cards

mean squared error (MSE)

insensitive to sample size (SSE is sensitive to sample size, as N grows, the SSE does too)

another term for variance (both N and N-1 are just constants used to scale for sample size)

root mean squared error- like SD (original units)

10
New cards

absolute error

less senstive to extreme values, lacks some desirable statistical properties

minimized by median

11
New cards

modeling separate means

score = β0 + β1Cond + Єi

β0 is the mean of the control group

β1 is the difference in means between control and experimental group

Cond is group assignment (0 = control, 1 = experimental)

Єi = residual error

12
New cards

general linear model

yi = β0 + Σβpxip + Єi

Σβpxip for all (P) predictor variables in the model

β0 = intercept, βp is/are slope(s)

parameters are β0 …. βp AND σ2Є (variance of the residuals)

13
New cards

linear algebra form

y = Xβ + Є

y is a vector of length N representing the outcome variable

X is an N x P + 1 matrix of predictors (first column is a column of ones for the intercept (control))

β is a vector of length P + 1 of coefficients, including the intercept

Є is a vector of length N of residuals

<p>y = Xβ + Є</p><p>y is a vector of length N representing the outcome variable</p><p>X is an N x P + 1 matrix of predictors (first column is a column of ones for the intercept (control))</p><p>β is a vector of length P + 1 of coefficients, including the intercept</p><p>Є is a vector of length N of residuals</p>
14
New cards

overfit

any time you add a predictor to a linear model you estimate with a sample of data, you will reduce error in the predictions within the same sample

bias in which a model captures random variation and treats it as informative signal

but we want to know how the model will generalize/fit to new data

<p>any time you add a predictor to a linear model you estimate with a sample of data, you will reduce error in the predictions within the same sample</p><p>bias in which a model captures random variation and treats it as informative signal</p><p>but we want to know how the model will generalize/fit to new data</p>
15
New cards

cross-validation

allows us to test how well the model generalizes but it reduces our sample size so we can hold out some data for validation

16
New cards

cross-validation procedure

  1. specify a training dataset

  2. specify a testing dataset

  3. estimate model on training data

  4. make predictions about testing data using model estimated on training data

  5. calculate MSE/RMSE

  6. select model with lowest MSE/RMSE


17
New cards

k-fold cross validation

split the data into k different folds or subsets of equal size

for each fold, train the model on the remaining k-1 folds, then test it on the held out kth fold

repeat

calculate average RMSE

18
New cards

leave one out (LOO) cross validation

for each of the N datapoints, train the model on the remaining N-1 datapoints

make prediction about the single held out data point

calculate average RMSE for each held out datapoint

can be very slow

19
New cards

inference vs behavior

Fisher: we build statistical models to test theories about the world, meant for learning/describing, inference

Neyman: we do inference to decide how to behave, inductive behavior: based on what we’ve learned from data, should I behave a certain way? interventions?

20
New cards

statistical decision theory

subfield of statistics focused on how to make choices based on what we learn from statistical model

interdisciplinary