3 - Correlation & Regression

0.0(0)
Studied by 0 people
call kaiCall Kai
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/36

encourage image

There's no tags or description

Looks like no tags are added yet.

Last updated 10:11 PM on 9/27/26
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

37 Terms

1
New cards

Scatterplots

most effective way of presenting relationship data

2
New cards

Linear Relationships


a type of correlation where the relationship between two variables can be represented by a straight line.

3
New cards

Curvilinear Relationships

best described w/ curves

  • analyzing them requires special stats


4
New cards

Correlations

descriptor of how reliably a change in 1 variable predicts change in another variable

5
New cards

Positive Relationships

ones where an increase in 1 variable predicts a decrease in the other

6
New cards

Negative Relationships

ones where an increase in one variable predicts a decrease in the other

7
New cards

Correlation & Causation

correlations only describe relationship b/w 2 variables

  • alone X be used to make a definitive statement about causation

  • can always design an experiment & then could make casual judgements


8
New cards

Pearson Product Moment Correlation

quantifying a correlation, it measures the strength and direction of the linear relationship between two continuous variables, descriptive statistic

  • varies from -1 to 1, sign = direction

  • weakest/no correlation = 0

  • absolute value of 1 (-1 or 1) = strongest correlation

  • + sign = positive correlation

  • - sign = negative correlation


9
New cards

Assumptions of the Pearson Correlation

uses 2 variables, both quantitative (usually scalar, sometimes ordinal)

  • variable relationships are linear, X curved

  • minimal skew/X large outliers

  • must observe whole range for each variable → Xtrim outcurved parts of curved data


10
New cards

When can an outlier be removed?

± 3 standard deviations from mean= safe to remove

11
New cards

Correlation Setup

make sure to set variables to scalar measure in SPSS

  • comparing 2 diff variables for the same set of cases

  • each row w/ min 2 values


12
New cards

Significance of Pearson Correlation

if p < .001 or p < .05, the correlation is statistically significant, indicating a strong relationship between the variables.

  • if p >.05 nothing interesting happening


13
New cards

Direction of Correlation

correlation is nondirectional

  • + or - tells directionof the relationship between two variables, indicating whether they increase or decrease together.


14
New cards

Writing out the Pearson Correlation

r value, (pearson value given), n & significance → r(N) = r value, p value

  • ex → r(787) = .140, p < .0001

  • ex → r(117) = -.054, p = .36


15
New cards

Parametric Analysis

Statistical methods that assume data follows a specific distribution, typically normality, allowing for the application of hypotheses testing.

  • Pearson → both variables are ration/interval & normal

  • conforms to expectations

  • best to use if parametric, scalar, normal


16
New cards

Nonparametric Analysis

  • spearman’s rank

  • kendall’s tau-b

  • ETA


17
New cards

Spearman’s Rank

A nonparametric measure of correlation that assesses the relationship between two variables by ranking their values, suitable for ordinal data and skewed distributions.

18
New cards

Kendalls’ tau-b

A nonparametric statistic that measures the strength and direction of association between two variables, particularly suited for ordinal data and smaller sample sizes.

  • gen superior to spearman esp for smaller groups, less affected by error

  • T(N) = correl, p >/< sig


19
New cards

ETA

special coefficent used for curvilinear relationships

  • ate-a

  • good for nominal by interval analyses


20
New cards

Jacob Cohen’s 1988 Paper

analysis of existing data showed avg correlation in literature of 0.30

  • made a scale

  • consider correlation as variance accounted for (r² x 100)

  • can be visualized as a venn, x & y overlap w/ r²


21
New cards

Jacob Cohen’s Scale

  • 0.00-0.10 = very weak/trivial (significant but weak)

  • 0.10-0.30 = small/minor/weak

  • 0.30-0.50 = moderate

  • 0.50-0.70 = large/strong

  • 0.70-0.90 = very strong

  • 0.90-1.00 = nearly perfect


22
New cards

What is regression?

model the relationship b/w a dependent variable & 1+ independent variables

  • allowing for predictions and analysis of trends


23
New cards

Correlation vs Regression

  • correlation → measures the strength and direction of a linear relationship between two variables

  • regression → analyzes the dependency of one variable on another, providing a predictive model


24
New cards

Regression Line

A graphical representation of the relationship between the dependent and independent variables in regression analysis, indicating the predicted values.

  • a line that minimizes the avg deviation of every point

  • least squares

  • includes a slope & y intercept

  • r² value = measure of how good regression is


25
New cards

How is regression calculated?

by analyzing error, look at predictive error for the y-axis variable

  • sum of the squares of the error = calc regression line

  • goal of the algorithm is to find the line w/ least error


26
New cards

Goal of Regression

find a best fit line that minimizes the squares of error

  • bog-standard equation for a line → Y’ = bX + a


27
New cards

How does regression differ from correlation?

  • both → relationships b/w variables, work best w/ quantitative variables

  • regression → explicit predictor variables used to estimate value of some target variable

    • design needs stronger evidence of causality

    • run an experiment w/ indep variables


28
New cards

Assumptions of Linear Regresion

  • requires scalar variables

  • 1 dep variable & 1+ indep variables

    • models can get big w/ dozen of indep variables

    • AI = big regression models

  • linear relationship b/w indep & dep variables

  • homoscedastic data


29
New cards

Homoskedasticity

property of a dataset having variability that’s similar across the whole range

  • opposite = heteroskedastic


30
New cards

R

correlation b/w observed values & ones model predicts

31
New cards

R²

amount of variability in dependent variable accounted for by changes in all indep variables

32
New cards

Adjusted R²

trying to correct relationships due to chance → result from lots of indep variables

  • better to use than regular r²


33
New cards

Std Err of the Estimate

measure of how accurately model predicts dep variable

34
New cards

ANOVA

tells us whether indep variables overall predict the dep

35
New cards

Unstandardized B

predicted change of dep from a +1 change in that predictor variable

36
New cards

Beta

tells how strongly this variable predicts dep, standardized for comparison

37
New cards

t & sig

tells if variable was a significant predictor of the dep