1/70
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
statistics
the science of variability
statistical investigative cycle
a lens that statiticians use to help them solve real world problems
steps of a statistical investigative cycle
problem- identify a statistical question
plan- chose a sample design, study design, and measures
data- collect and process data
analysis- look for patterns with summary tables, graphs, and statistical models
conclusion- interpret the results and generate new questions about the real world context
causal inferences
claims that some choice we make will cause some effect to happen
causal factor
the “initial”
outcome factor
the conclusion or the thing effected
claimed effect
the “arrow”
generalization inferences
claims about an entire group of people, based only on information from a sample of people
population of interest
the group of people it is trying to describe
attribute
the characteristic about them
What was compared?
“Was it apples-to-apples?”, related to CAUSAL claims, consider variability and uncertainty, internal validity
Who’s not here?
related to GENERALIZATION claims, how to question external validity
Data
a representation of someone or something
tidy data
a way of mapping the real world to a data set
observations
the things we are interested in, the rows, in an easy way it is all the first colums
attributes
pices of information, about the obserbations, that we are interested in, correspond to colums, all of the “data” on the top of the set, # of columns = # of attributes
measures
the way in which we ollect information about observations
quantitative data
data in which the values of an attribute are numbers representing a quantity of something, hint- decimals
categorical data
data in which the values of an attriute for an observation are selected from a set of different category labels
rating scale data
data in which the values of an attribute for an observation are selected from a predetermined list that has an order to it (strongly disagree, agree… OR scale from 1-10), rating scale has an OBVIOUS ORDER to it
text data
data in which the values of an attribute for an observation is text, typially from an open-ended question
time series data
data in which the values of an attribute for an observation indicate a moment in time, not how long something is- the MOMENT in time IT HAPPENED
RELIABILITY (memorize exact)
the extent to which the data you collect from a measure truly represents and reflects the real world characteristics of the observations
data validation
the act of ensuring that the values collected from each observation for each attribute are valid
data cleaning
removing invalid values from a dataset
univariate analysis
the analysis of a single attribute or variable at a time
standard deviation
the measure of spread, the typical amount by which each observation is different from the mean
five number summary
minimum, first quartile (25%), median, third quartile (75%), maximum
box plots
graphic depiction of a five number summary
mean
the average
median
the middle
dot plot
graph where each observation is displayed as one point on the graph
density plot
very similar to a dot plot with a line drawn across the top of all the stacks of dots
first quartile
25 percentile, the number at 25 means n had a “score” of x or LOWER
third quartile
75 percentile, the number at 75 means n had a “score” of x or LOWER
frequencies
total number of observations whose response is equal to a particular value, “how many”
relative frequencies
the percentage of all observations without missing values whose value is equal to a particular response, out of everyone, how many?
bar graph
based on a frequency table, has one bar for every response option, the bars height is equal to each response option’s frequency or relative frequency
word cloud
graphically depicts the words across the responses from the observations
distribution
the pattern that the responses from all the observations
normal distributions
bell curve, have a value near the average
unimodal distributions
distributions with only one peak
multimodal distributions
distributions with more than one peak
skewed distributions
normal distribution but with one of the side stretched out
right skew
pulled to the right, mode: closer to the left, median: closer to the right, mean: closest to the right
left skew
pulled to the left, mode: closer to right, median: closer to left, mean: closest to left
middle 95%
the distribution represents MOST of the observations
standard deviation
upper limit of middle 95%- lower limit of 95% / 4
two- way table
similar to a frequency table, except one attriute’s frequencies are presented as different rows in the table, and a second attribute’s frequencies are presented as columns`
column percents
relative frequencies based only on the total from a SINGLE column
absolute difference
two colons percents by SUBTRACTING one column percent from the other
absolute difference scale
0%- very similar, 5%- pretty similar, 10%- the SWITCH to different, 15%- VERY different
relative difference
two column percents by dividing one column from the other
relative difference scale
1.0- pretty similar, 0.8/1.25- slightly different, 0.6/1.5- very different
line graph
places time on the horizontal (x) axis, and the percentage on the vertical axis
grouped density plots
two density plots on the same graph
ratio of standard deviations
equal to the largest standard deviatio between the two groups divided by the smallest standard deviation, larger SD/ smaller SD
ratio of standard deviation scale
1.0- very similar, 2.0- pretty similar, 3.0- pretty different, 4.0 very different
effect size
equal to the mean difference divided by the larger standard deviation, group 1 mean - group 2 mean / larger SD
effect size scale
0.00- groups very similar, 0.25- pretty similar, 0.50- different, 0.75- very different, 1.00- extremely different
scatter plot
a plot in which each observation is placed as a point on a graph according to their value for each of the two quantative attributes
smoothed trend line
line through the average value of the vertical axis attribute across all values of the horixontal axis attribute
positive trend
trend line goes from the bottom left to the top right
negative trend
trend line goes from top left to bottom right
flat line
no trend
another word for slope is…
direction
another word for variability is…
strength
interaction plot
similifies comparing distributions between two sets of different groups by focusing just on the relative frequency or mean
interaction
difference in directions/ slope
margins of error
formal quantification of the uncertainty inherent to an inference
always prefer…
interval estimates over a single number of an estimate