1/98
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
correlational studies
measure correlations between predictor and criterion variables
subjects come with their own set of variables
experimental studies
studies in which the independent variables are directly manipulated and the effects on the dependent variable are examined
random assignment of subjects to experimental condition
descriptive statistics
numerical data used to measure and describe characteristics of groups.
Includes measures of central tendency (mean) and measures of variation (variance)
inferential statistics
using samples to make claims about the populations
population
all events, people, scores of interest
sample
a subset of the population
random sample
each member of the population has an equal chance of being selected
convenience sample
only members of the population who are easily accessible are selected
cons of correlational study
cannot establish causation, hard to know if the value of criterio is caused be a predictor, a lot of other variables may come into play
random assignment
assigning participants to experimental and control conditions by chance, thus minimizing preexisting differences between those assigned to the different groups
Causality
when we change the value of x, the probability of y occuring also changes
IVs
variables that are manipulated by the experimenter
DVs
key variables of interest that we measure and analyze
between-subjects design
A research design in which different groups of participants are randomly assigned to experimental conditions or to control conditions.
within-subjects design
participants are exposed to all levels of the independent variable
pros of within-subjects design
better control on the individual differences
no differences between two people
cons of within-subject design
practice effect, and boredom
2 group counter balance within subject design
divides your participants into two distinct groups to experience all experimental conditions in opposing orders, which controls for order and practice effects
categorical/ qualitative variables
lack numeical properites
can be nominal (no order)
can be ordinal (meaningful order)
numerical (quantitative) variables
have values that represent a counted or measured quantity
differences between levels are consistent and meaningful
zero point may be arbitrary or meaninful
line graph
used to show relation between quantitative measures
each value on x-axis has one data point
often, variable on x-axis is continuous
scatter plot
useful for visualising the association between two quantitative variables
each value on x-axis can have multiple data points
each dot represents one particpants
pie charts
bad because:
no common reference point for each slice
hard to trach change over time
requires clunky labels or legends
hated by colour blind people
skewness
Measure of asymmetry in data distribution.
kurtosis
the frequency/propability of scores that are far from the centre of distribution
high kurtosis
more outlier scores -> fatter tails
low kurtosis
light tails, lack of outliers
mode
the most frequently occurring score(s) in a distribution
if two adjacent scores occur with equal frequency- average of those two scores
if two non-adjacent scores occur with equal frequency distribution is bi-modal
cons of mode
- only gives info about a single score(s)
- sensitive to frequent extreme scores
- doesn't account for variability
- changing just one observation can change which value is the mode, even though the overall data barely changed.
pros of mode
- value will appear in the data set
- easy to understand
- robust to extreme scores
median
the middle score in a distribution; half the scores are above it and half are below it
- tells you how a score ranks within a distribution of scores
percentiles
the proportion of values in a sample that fall below a given value
median > mode
positive skew
mode < median
negative skew
pros of median
Not affected by extreme scores
stable even when mode is undefined
easy to define as the middle score
cons of median
difficult to use in statistical theorems and calculations
no simple formula
mean
average
pros of mean
-summarizes data in a way that is easy to understand
-uses all the data and is the most--used measure of central tendency
minimises the distances (deviations) from ach score (balance point)
every observation contributes to the mean
-best at minimizing squared errors
cons of mean
-affected by outliers
-values may not actually exist in data
trimmed means
means calculated on data for which we have discarded a certain percentage of the data at each end of the distribution
range
the difference between the highest and lowest scores in a distribution
cons of range
distorted by outliers
pros of the interquartile range
not sensitive to extreme scores
pearson r
A method of computing correlation when both variables are linearly related and continuous
- index of goodness of fit
-how close points are to line
linear relationship
A relationship that has a straight line graph
curvilinear
best fit line is characterized by curved lines
monotonic trend
as x increases y increases, might not be a straight line
non monotonic trend
As X increases, Y changes direction at least once
covariance
A measure of linear association between two variables. Positive values indicate a positive relationship; negative values indicate a negative relationship
covariance depends on
the sum on the products of deviation scores
cons of covariance
sensitive to the spread of x and y
depends on our choice of units
weak r
0.1 - 0.3 (-0.1 to -0.3)
moderate r
0.3-0.5 (-0.3 to -0.5)
strong r
0.5 - 1 (-0.5 t -1)
factors that affect correlation (r)
non-linearity
extreme scores
restricted range
heterogenous sub groups
simpson's paradox
when averages are taken across different groups, they can appear to contradict the overall averages
a trend that exists in several groups can disappear or reverse when groups are combined
pearson product moment correlation coefficient (r)
for two continuous variables
spearman's correlation coefficient for ranked data (rho, p, or ra)
two ranked/ ordinal variables
sort into ranks then you calculate r
point-biserial correlation (rpb)
1 dichotomous and 1 continuous variable
- correct/incorrect on a single MCQ vs total exam score)
calculate r
Phi Correlation (rφ)
two dichotomous variables
yes/no vs yes/no
calculate r
rs
measures monotonicity of X,Y association
rs is sensitive to
extreme scores affecting monotonicity
when to use rs
if your data is ranked and intervals are meaningless
point estimate
a summary statistic from a sample that is just one number used as an estimate of the population parameter
bootstrapping
using one sample to estimate how much your result might vary across repeated samples
r doesn't tell you anything about
slope or best fit line
Confidence Interval
the range of values within which a population parameter is estimated to lie
--% of the calculated intervals would be expected to contain the true parameter value from our population
p value
how unusual your observed result would be if the null hypothesis were true
describes how unusual your sample r is, assuming H0 is true.
type 1 error
Rejecting null hypothesis when it is true
Concluding there's evidence of a correlation when the population correlation is actually zero.
linear regression
a type of regression that models that relationship with a straight line when there's one predictor
R^2
the proportion (percent) of the variation in the values of y that can be accounted for by the least squares regression line
residual standard error
the standard deviation of residuals.
coefficients
the parameters of the regression line
multiple r-squared
Proportion of variance explained by predictors.
Multiple Regression
a statistical technique that includes two or more predictor variables in a prediction equation
normal distribution is defined by
standard deviation and mean
population standard deviation
how spread out individual values are around the population mean.
t-test
a statistical test used to evaluate the size and significance of the difference between two means
sampling distribution of the mean
the distribution of sample means over repeated sampling from one population
standard error of mean
the standard deviation of the sampling distribution of sample means
sampling distribution
probability/ frequency distribution of a sample statistic
sampling error
the difference between a sample statistic and the true population value, caused by which individuals happen to be sampled.
Central Limit Theorem
The theory that, as sample size increases, the distribution of sample means of size n, randomly selected, approaches a normal distribution.
type 2 error
we fail to reject the H0 when it is not true
there was actually a change
reject H0 (Type 1 error)
p = a
Do not reject Ho (H0 not true)
p = 1 - a
Reject H0 (H0 is false)
correct decision
Do no reject H0 (H0 is false)
p = b
Effect size
how large the real effect or difference is.
Larger effect size
Easier to detect; increases power and decreases β.
Statistical power
Probability of correctly rejecting a false null hypothesis.
1 - b
Alpha (α)
the probability of making a type I error
Beta (β)
probability of making a type II error
Lower alpha
Stricter rejection rule → more missed effects → higher b
Effect size can be measured
Difference between the true mean and the null mean
Larger sample size
More precise sample means → easier detection → lower b
Sample size and SEM
As n increases, SEM decreases
Smaller SEM
Sample means cluster more closely around the true population mean
factors affecting the probability of making a type 2 error
alpha level, effect size, and sample size