1/21
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
proportion
x number of people who responded a certain way/total (AKA relative frequency)
symbols for proportion
sample: p-hat; population: p
two-way table purpose
to investigate relationship between 2 categorical variables
outlier
observed value that is noticeably distinct from other values in a dataset; outlier if smaller than Q1 - 1.5(IQR) or greater than Q3 + 1.5(IQR)
mean vs. median when data is: skewed to the right, skewed to the left, symmetric & bell-shaped
mean > median, mean < median, m = median
mean caluclation & symbol
sum of all data values/number of all data values, symbolized by x bar for sample (statistic) and mu for population (proportion)
median
middle entry (odd) or average of 2 middle entries (even)
resistance
related to the impact of outliers on a statistic
standard deviation definition & symbols
measures how much variability is in the data (larger values = more spread out data), s for sample and sigma for population
if distribution is symmetrical and bell-shaped, about what percent of data should fall within 2 standard deviations of the mean? (mean - 2 to mean + 2)
95%
z-score
how many standard deviations the value is from the mean: (x - mean)/s or (x - mu)/sigma
Pth percentile
value of a quantitative variable that is greater than P% of the data
Q1, Q2, Q3
first quartile = 25th percentile, second quartile = 50th percentile = median, third quartile = 75th percentile
range & interquartile range
max - min, Q3 - Q1
5 number summary
minimum, Q1, median, Q3, maximum
boxplots include:
appropriate numerical scale, box from Q1 → Q3, line at median, line from each quartile to most extreme data that isn’t an outlier, plotted outliers
side-by-side graphs
facilitate comparison of distributions, includes graph for quantitative variable for each group in the categorical variable
scatterplots
graphs of relationship between two quantitative variables (explanatory variable on x-axis, response variable on y-axis)
correlation
measure of strength and direction of linear association, sample = r and population = rho
correlation key things to note
always between -1 & 1, sign indicates direction of association (positive or negative), if closer to 1 or -1 it’s stronger correlation, no units, does not imply causation, just because it’s near 0 doesn’t mean variables aren’t associated, can be heavily influenced by outliers
histogram vs. box plot
histograms are for quantitative/numeric data where we are choosing our bins
mode
the value with the highest frequency OR the bin with the highest frequency when it’s a histogram