1/79
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
statistics
methods and procedures; a tool kit for analyzing data
correlational studies
measure associations between predictor and criterion variables, “subjects” come with their own set of variables
designed experiments
measure effects of indepedent variables on dependent variables, random assignment of “subjects” to experimental conditions
descriptive statistics
describes/summerizes important characteristics of data, uses graphs & statistics (mean or standard deviation), describes interesting features of sample
population
all events (subjects, scores, etc) of interest
sample
subset of a population
random sample
each member of the population has equal chance of being selected
convenience sample
taking a subset of a population that is easiest; may not be very representative
inferential statistics
use sample to make claims about a population; estimate parameters from sample statistics, investigate differences among populations by examining difference among group/sample
independent variable
manipulated by the experimenter, used in designed experiments
dependent variable
the things that are measured and constitute the data, or results, that will be analyzed; used in designed experiments
statistical idea of causality
when we intervene and change X, then the probability of Y also is changed (smoking causes cancer even though not all smokers get cancer and not all people with cancer are smokers)
regression
measures association between predictor and criterion variables
random assignment
designed experiments’ “subjects” come with their own characteristics, this procedure is done so that effects of subject differences should be unrelated to effects of independent variables
categorical variables
typically have a few discrete levels (s#x, marital status, brain area, profession), often lack numeric properties
numerical variables
typically have many levels though not always, (reaction time, body weight, income, age), the levels are ordered, differences between levels are meaningful
nomial scale
categorical, have no obvious numeric properties (eye colour, s#x, political party)
ordinal scale
order of the levels is meaningful, differences between levels may nit be meaningful (age, musical expertise, likert scales)
interval scale
numeric variables usually with many levels, but lack a true meaningful zero (time of day, fahrenheit and celsius), differences between levels are meaningful
ratio scale
numeric variables (reaction time, proportion correct, body weight), have a true non-arbirary zero on the scale, differences between levels are meaningful
need for graphs
highlight important features, patterns, trends, that cannot be visualized from the raw numbers
histogram
x aix is horizontal, binned variable, (bins defined by upper and lower limits, midpoints), y-axis → frequency count
bar graphs
value of interest is represented by height of ____, x-axis usually consists of levels on a qualitative/nominal variable
pie charts
used for visualizing proportions
line graphs
useful for illustrating trends, typically have a discrete variable on the x-axis
scatter plots
useful for visualizing the association between quantiative variables
unimodal distributions
single peak (INSERT PIC)
bimodal distributions
two peaks (INSERT PIC)
positive skew
INSERT PIC
negative skew
INSRET PIC
kurtosis
refers to the frequency/probability of scores that are far from the center of distribution
exploratory data analysis
goal to discover and summarize interesting aspects of data, discover interesting hypotheses to test, important with large and complex data sets
confirmatory data analysis
data gathered and analyzed to evaluate specific a priori hypotheses
(ex. clinical drug trials: “i think this will happen when i do this”) → find the real effect
between subjects design
randomly subject every person to one condition, minimizes chances of systematic differences between groups
within subjects design
every subject is tested in every condition, all conditions have the same subjects, so differences between conditions must be due to experimental treatments
mode
most common score, may be poorly defined, small changes in frequencies can produce big changes, actually appears in the data
median
middle score/50th percentile, robust to outliers
mean
msot commonly used measure of central tendency, average score, not robust to outliers (can trim both ends of data set to outset this effect → 10%)
mode advantages and disadvantages
A
robust to outliers, value actually appears in the data, the value with the highest probability of subjects having that score, can be found in nominal data
D
depends on how we bin scores, can be poorly defined/ unstable for flat or bimodal distributions
median advantages and disadvantages
a
robust to outliers, good index of typical scores, can be calculated even with flat distribution
d
no mathematical formula, difficult to use in equations
mean advantages and disadvantages
a
easy to use in statistical formulas, best estimate of typical score
d
value may not actually exist in the data, less robust to extreme values
range
difference between highest and lowest values, sensitive to extreme values
quartiles
most distributions can be divided into, sorted data into four subsets each containing 25% of data
interquartile range
difference between 3rd and 1st quartiles, the spread of the middle 50% of scores, stable against outliers
outliers
extreme values in data, sometimes real data (should be kept), sometimes experimental error (should be removed)
variability around the mean
use absolute values and squared deviations
variance
more common than MAD, calculate the sum of squared deviations, divide by the number of observations minus 1 (n-1), results in sample variance (s²)
why n-1 and not n
makes s² an unbiased estimator of the population variance
linearly related scales
when measures are perfectly correlated, most statistical analyses of those variables will yield the same results (celsius and fahrenheit)
bivariate
is data for which there are two variables for each observation (ex. the ages of husbands and wives of 10 married couples)
pearson’s product-moment correlation coefficient/correlation coefficient
is a measure of the strength of the linear relationship between two variables, represented with p for population and r for sample (range from 1 to -1)
properties of pearson’s r
it is unaffected by linear transformations (ex. the correlation of Weight and Height does not depend on whether Height is measured in inches, feet, or even miles)
point estimate
a single number used when trying to guess the parameter, limited in usefulness because it does not reveal the uncertainty associated with the estimate
parameter
A value calculated in a population. For example, the mean of the numbers in a population
confidence interval
are intervals constructed using a procedure that will contain the population mean a specified proportion of the time
null hypothesis
tested in significance testing. It is typically the hypothesis that a parameter is zero or that a difference between parameters is zero
criterion variable
The variable we are predicting and is referred to as Y
predictor variable
The variable we are basing our predictions and is referred to as X
simple regression
When there is only one predictor variable
simple linear regression
the predictions of Y when plotted as a function of X form a straight line
regression line
the best-fitting straight line through the points, the line that minimizes the sum of the squared errors of prediction
standard error of the estimate
a measure of the accuracy of predictions, is the square root of the average squared deviation
influence
how much the predicted scores for other observations would differ if the observation in question were not included (using Cook’s D)
leverage
based on how much the observation's value on the predictor variable differs from the mean of the predictor variable; The greater an observation's _______, the more potential it has to be an influential observationm
monotonic trend
as x increases, y increases, or decreases without reversal (may not be linear)
non-monotonic
as x increases, y changes direction at least once
linear relationship
the best fit line is straight
curvilinear relationship
best fit line is not straight
covariance
measures the degree to which two variables vary together, depends on the sum of products of deviation scores (positive when deviation scores have the same sign and negative when deviation scores have different signs): measure of association but is influenced by the spread of X and Y
heterogeneous subsamples
a statistical association observed in a population can be attenuated and even reversed within subgroups that make up a population; correlations calculated with these subgroups may be misleading
simpson’s paradox
correlation for population is positive but correlation within subgroups are negative
point-biserial correlation (rpb)
for one dichotomous and one continuous variable; calculate person’s r but call it rpb, not sensitive to units (ex. correct/incorrect answer on mc question and exam total score
spearman’s correlation
useful when observations have been replaced by their numerical ranks, is more robust to some types of extreme points (sensitive to scores affecting monotonicity), is an index of strength/direction of monotonic relationship
r phi
two dichotomous variables, code each binary variable value with two numbers (ex. gender (male/female) and pass/fail on exam)
percentile bootstrapping method
used for confidence interval, randomly selecting pairs from the data and resampling to use to estimate a confidence interval
Permutation distribution
Represents the differences we’d expect to see if they really were no difference between the groups
Permutation test
A way of testing a hypothesis by randomly rearranging the observed data To simulate what we’d expect if the hypothesis were true
P value
The probability of getting our data when the null hypothesis is true
Multiple regression
Relates y variable to multiple predictor variables X, Compute complex function/equation relating Y And x, Can be used to fit multiple lines to a set of XY data, Can be used to compute non-linear association between X and y
R² coefficient of determination
Component of regression line, Tells us how much of the variability in the outcome is explained by all of your predictors together