1/55
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
statistics
science of variability
First principle (examining a sample’s external validity)
What was compared?
Second principle (examining representativeness of a sample & where the data came from)
Who’s not here?
Third principle (considering variability & uncertainty)
Incorporate ish-ness
causal inference
claims that some choice we make will cause some effect to happen
causal claims consist of…
a dependent variable, an independent variable, & a claimed effect
generalization inferences
claims about an entire group of people, based only on information from a sample of people
generalization claims consist of…
the population of interest, an attribute, & an estimate(#)
data
representation of someone or something
observations
things we are interested in (rows)
attributes
pieces of information about the observations (columns)
measures
the way we collect information about observations
quantitative data
refers to data in which the values of an attribute are numbers representing a quantity of something
categorical data
refers to data in which the values of an attribute for an observation are selected from a set of different category labels
rating scale data
refers to data in which the values of an attribute for an observation are selected from a predetermined list that has an order to it
ex. (strongly agree, agree, disagree, strongly disagree)
text data
refers to data in which the values of an attribute for an observation is text, typically from an open-ended question
time series data
refers to data in which the values of an attribute for an observation indicate a moment in time
reliability
refers to the extent to which the data you collect from a measure truly represents and reflects the real world characteristics of the observation
data validation
the act of ensuring that the values collected from each observation for each attribute are valid
data cleaning
removing invalid values from a dataset
univariate analysis
the analysis of a single attribute or variable at a time
standard deviation
a measure of spread
boxplots
graphic depiction of a five-number summary
dot plot
a graph where each observation is displayed as one point on the graph
density plot
a dot plot with a line drawn across the top of all of the stacks of dots
frequencies
the total number of observations whose response is equal to a particular value
bar graph
graph based on a frequency table, and has one bar for every response opinion
word cloud
graph that depicts the words across the responses from the observations, size of the word varies by the frequency of the word
distribution
the pattern that the responses from all the observations make
quantitative measures include…
count, mean, standard dev., 5-number summary, box plot, density curve
categorical & rating scale measures include…
count, frequencies, percentages, bar graphs
text data measures include…
count, word clouds
quantitative data characteristics include…
shape, center(location), spread
categorical, rating scale, and text data characteristics include…
frequencies and percentages
normal distribution
distributions with a bell-shape curve shape
right skew distributions
distributions that look like the right side of a normal distribution has been stretched out
left skew distributions
distributions that look like the left side of a normal distribution has been stretched out
middle 95%
part of the distribution that represents most of the observations
what are typical values?
the mean, median, & mode
two way table
a frequency table where the attributes are presented as different rows, and a different attribute is presented as columns
column percents
relative frequencies based only on the total from a single column
absolute difference
whether two column percents are similar or dissimilar
relative difference
what to compute instead of absolute difference when the column percents are small
line graph
what type of graph to use when dealing with time series attributes
grouped density plot
graph to compare two distributions → two density plots in one graph
utilizing shape, spread, and location
if the graph’s shapes are very different…
then it doesn’t make sense to compare the locations with the spread
standard deviation formula
(upper limit of middle 95% - lower limit of middle 95%) / 4
absolute difference formula
Group 1’s column percent - Group 2’s column percent
relative difference formula
Group 1’s column percent / Group 2’s column percent
ratio of standard deviations formula
largest standard deviation / smallest standard deviation
if the ratio of standard deviations is 1…
the standard deviations are pretty similar
if the ratio of standard deviations is 3+…
the standard deviations are pretty different
effect size formula
mean difference / larger standard deviation OR
(Group 1’s mean - Group 2’s mean) / larger standard deviation
if the effect size is .25…
then there is a pretty small distance between the means
if the effect size is .75…
then there is a pretty large distance between the means