1/36
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Scatterplots
most effective way of presenting relationship data
Linear Relationships
a type of correlation where the relationship between two variables can be represented by a straight line.
Curvilinear Relationships
best described w/ curves
analyzing them requires special stats
Correlations
descriptor of how reliably a change in 1 variable predicts change in another variable
Positive Relationships
ones where an increase in 1 variable predicts a decrease in the other
Negative Relationships
ones where an increase in one variable predicts a decrease in the other
Correlation & Causation
correlations only describe relationship b/w 2 variables
alone X be used to make a definitive statement about causation
can always design an experiment & then could make casual judgements
Pearson Product Moment Correlation
quantifying a correlation, it measures the strength and direction of the linear relationship between two continuous variables, descriptive statistic
varies from -1 to 1, sign = direction
weakest/no correlation = 0
absolute value of 1 (-1 or 1) = strongest correlation
+ sign = positive correlation
- sign = negative correlation
Assumptions of the Pearson Correlation
uses 2 variables, both quantitative (usually scalar, sometimes ordinal)
variable relationships are linear, X curved
minimal skew/X large outliers
must observe whole range for each variable → Xtrim outcurved parts of curved data
When can an outlier be removed?
± 3 standard deviations from mean= safe to remove
Correlation Setup
make sure to set variables to scalar measure in SPSS
comparing 2 diff variables for the same set of cases
each row w/ min 2 values
Significance of Pearson Correlation
if p < .001 or p < .05, the correlation is statistically significant, indicating a strong relationship between the variables.
if p >.05 nothing interesting happening
Direction of Correlation
correlation is nondirectional
+ or - tells directionof the relationship between two variables, indicating whether they increase or decrease together.
Writing out the Pearson Correlation
r value, (pearson value given), n & significance → r(N) = r value, p value
ex → r(787) = .140, p < .0001
ex → r(117) = -.054, p = .36
Parametric Analysis
Statistical methods that assume data follows a specific distribution, typically normality, allowing for the application of hypotheses testing.
Pearson → both variables are ration/interval & normal
conforms to expectations
best to use if parametric, scalar, normal
Nonparametric Analysis
spearman’s rank
kendall’s tau-b
ETA
Spearman’s Rank
A nonparametric measure of correlation that assesses the relationship between two variables by ranking their values, suitable for ordinal data and skewed distributions.
Kendalls’ tau-b
A nonparametric statistic that measures the strength and direction of association between two variables, particularly suited for ordinal data and smaller sample sizes.
gen superior to spearman esp for smaller groups, less affected by error
T(N) = correl, p >/< sig
ETA
special coefficent used for curvilinear relationships
ate-a
good for nominal by interval analyses
Jacob Cohen’s 1988 Paper
analysis of existing data showed avg correlation in literature of 0.30
made a scale
consider correlation as variance accounted for (r² x 100)
can be visualized as a venn, x & y overlap w/ r²
Jacob Cohen’s Scale
0.00-0.10 = very weak/trivial (significant but weak)
0.10-0.30 = small/minor/weak
0.30-0.50 = moderate
0.50-0.70 = large/strong
0.70-0.90 = very strong
0.90-1.00 = nearly perfect
What is regression?
model the relationship b/w a dependent variable & 1+ independent variables
allowing for predictions and analysis of trends
Correlation vs Regression
correlation → measures the strength and direction of a linear relationship between two variables
regression → analyzes the dependency of one variable on another, providing a predictive model
Regression Line
A graphical representation of the relationship between the dependent and independent variables in regression analysis, indicating the predicted values.
a line that minimizes the avg deviation of every point
least squares
includes a slope & y intercept
r² value = measure of how good regression is
How is regression calculated?
by analyzing error, look at predictive error for the y-axis variable
sum of the squares of the error = calc regression line
goal of the algorithm is to find the line w/ least error
Goal of Regression
find a best fit line that minimizes the squares of error
bog-standard equation for a line → Y’ = bX + a
How does regression differ from correlation?
both → relationships b/w variables, work best w/ quantitative variables
regression → explicit predictor variables used to estimate value of some target variable
design needs stronger evidence of causality
run an experiment w/ indep variables
Assumptions of Linear Regresion
requires scalar variables
1 dep variable & 1+ indep variables
models can get big w/ dozen of indep variables
AI = big regression models
linear relationship b/w indep & dep variables
homoscedastic data
Homoskedasticity
property of a dataset having variability that’s similar across the whole range
opposite = heteroskedastic
R
correlation b/w observed values & ones model predicts
R²
amount of variability in dependent variable accounted for by changes in all indep variables
Adjusted R²
trying to correct relationships due to chance → result from lots of indep variables
better to use than regular r²
Std Err of the Estimate
measure of how accurately model predicts dep variable
ANOVA
tells us whether indep variables overall predict the dep
Unstandardized B
predicted change of dep from a +1 change in that predictor variable
Beta
tells how strongly this variable predicts dep, standardized for comparison
t & sig
tells if variable was a significant predictor of the dep