1/93
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Entities measured and studied
Individuals
The entire group to be studied
Population
Subset of pop. from which you collect data
Sample
Numerical summary of population
Parameter
Numerical summary of sample
Statistic
Methods for summarizing collected data via tables, graphs, numerical summaries (averages or percentages)
Descriptive statistics
Methods that take a result from a sample, extend it to the population and measure the reliability of the result. Always contains uncertainty
Inferential stats
Any characteristic of the individuals within the pop.
Variable
What values a variable takes and how often it takes those values
Distribution of variable
If observation takes on numerical values
Quantitative
If each observation belongs to a set of categories
Qualitative variable
Measurement if the values of the variable name, label, or category. Doesn’t allow for ranking
Nominal (qualitative)
Measurement if the variable has properties of nominal level but can be ranked
Ordinal (qualitative)
Measurement if the variable has the properties of the ordinal level but each difference in value has meaning. (A value of zero doesn’t mean the absence of the quantity)
Interval level (quantitative)
Measurement if the variable has the properties of the interval level but the ratios of the value has meaning. A value of zero means absence of quantity
Ratio level (quantitative)
When someone assigns individuals in a study to certain experimental conditions then observed outcomes
An experiment
When someone observes and records the behavior of the individual without imposing any conditions
Observational study
Explanatory variable that was not considered in the study but that affects the value of the response
Lurking variable
characteristics typical of those possessed by the population of interest
Representative sample
n subjects from a population of N size in which each possible sample size of n has the same chance of being selected (n = sample, N = pop.)
Simple random sample
When dividing the population into separate groups called strata then selects a simple random sample from each stratum. (Each stratum should be homogeneous in some way w respect to variable of interest)
Stratified sample
Dividing the population into large number of clusters then a simple random sample of clusters is selected and all individuals in the selected clusters are included
Cluster sample
Selecting every kth individual from a pop. first individual selected corresponds to a random number btween 1 and k
Systemic sample
When the technique used tends to favor one part of the pop. over another
Sampling bias
When individuals selected to be in the sample who do not respond have different opinions from those who do respond
Non response bias
Response bias
When the answers on a survey do not reflect the true opinions of the respondents
Subject do not know what treatment they’re receiving
Single blind experimen
Neither subject nor researcher knows what treatment they’re subject is getting
Double blind experiment
frequency distribution
listing each category of data and the number of observations in each
listing category of data together with the relative frequency and the proportion of observation in each category. found by taking the frequency for a particular category and dividing by total number of observations
relative frequency distribution
mean
sum of observation divided by number of observations (average)
μ - population
x̄ - sample
median
the middle observation when they’re listed in ascending order
M
mode
observation that happens most frequently
if median is bigger than mean
left skewed
median = mean
symmetric
median less than mean
right skewed
standard deviation
typical value for how far the data fell from the mean
s - sample
σ - population
always greater or equal to 0
not resistant
1 standard deviation of the mean
68%
2 standard deviations of the mean
95%
3 standard deviations of the mean
99.7%
z score
distance that a data value is from the mean in terms of the number of standard deviations
z score equation
(observation - mean)/standard deviation
approx. 68% of z score’s
fall between -1 and 1
approx. 95% of z scores
fall between -2 and 2
approx. 99.7% of z scores
fall between -3 and 3
pth percentile
means p% of the observations fall BELOW the pth percentile and only (100-p)% fall above
first quartile Q1
25th percentile
Second quartile, median (M)
50th percentile
third quartile Q3
the 75th percentile
finding quarterlies
arrange data in order
find the median of all of them
then find the median of the numbers above and below the median
interquartile range IQR
Q3-Q1
Lower fence for outliers
Q1-1.5(IQR)
Upper fence for outliers
Q3+1.5(IQR)
for approx. symmetric distributions
use the mean and standard dev.
for distributions that are skewed or have outliers
use the median and IQR
five number summary
minimum, Q1, M, Q3, maximum
approx. symmetric distributions the median is
near the center of the box
distributions that are skewed right the median is
slightly left of the center of the box.
right whisker will be longer than left whisker
distributions that are skewed left the median is
slightly right center of the box
left whisker will be longer than the right whisker
response variable (y)
measures outcome of study
explanatory variable (x)
explains or influences change in the response variable
a value of r close to 0 indicates
that the relationship is not linear
weak correlation
0.2-0.5
moderate correlation
0.5-0.8
strong correlation
0.8-1
correlation coefficient, r is
not resistant
unitless
if the absolute value of r is greater than critical value
linear relation exists between the two variables otherwise, no linear relation exists
least-squares regression
y(hat)=bo+b1x
yhat is predicted value of y for value of x
bo is the average value of y when x=0
b1 is the change in the average value of y when x increases by 1 unit
the way we quantify uncertainty
probability
probability, experiment is
an act or process of observation with uncertain results that can be repeated
Multiplication Rule of Counting
Total outcomes for an experiment = product of the number of outcomes for each individual task
Repeating a task with n possible outcomes, r times → total possible outcomes
use when you can repeat and when order does not matter
n^r
Selection where order IS important, without replacement (e.g., 1st/2nd place finishers)
Permutation
Formula for number of permutations of r objects from n
nPr = n! / (n-r)!
Selection where order is NOT important, without replacement (e.g., picking a group of friends)
Combination
Formula for number of combinations of r objects from n
nCr = n! / [r!(n-r)!]
Any collection of outcomes from a probability experiment
Event
An event with only one outcome
Simple event (eᵢ)
The probability of any event A must satisfy
0 ≤ P(A) ≤ 1
An event with a low probability of occurring (often < 0.05)
Unusual event
Probability formula when outcomes are equally likely
P(A) = (# outcomes in A) / (total # outcomes)
Estimating probability using real data: (# times A observed) / (# repetitions)
Empirical approach
the median is (outliers)
resistant
Standard deviation (outliers)
not resistant
IQR (outliers)
resistant
w/out replacement
order is important
permutation (n!)
pareto, pie, bar charts
qualitative
histogram, stem and leaf
quantitative
The Complement Rule states
P(Aᶜ) = 1
so P and Ac must add up to 1
if (Q2 - Q1) is greater than (Q3 - Q2)
skewed left
if (Q2 - Q1) is less than (Q3 - Q2)
skewed right
higher IQR means
more varied range
probability refers to
what is expected in the long-term, not short-term.