1/55
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
statistics
study of how best to collect, analyze, interpret, present and draw conclusions from data
data
any collection of numbers, characters, images, or other itens that provide info about something
descriptive statistics aka exploratory data analysis (EDA)
organizing data
summarizing data
presenting data in an informative way
inferential statistics
determining something about the group of interest (population) based on a sample
methods for making decisions/predictions and drawing conclusions about population
objectives of statistical analysis
how best can we collect data?
how should we explore/present data?
what can we infer from the analysis?
how to quantify and explain the variability?
population
the entire group of interest
sample
part of the population selected to draw conclusions abou tthe entire population
individual (subject)
a person or any specific object in a population
variable
any characteristic of an individual
can take different values for different individuals
what do we ask each individual?
qualitative or quantitative
qualitative (categorical) variable
are you a sport fan
favorite ice cream flavor
year of school
quantitative variable
how much do you spend daily
how many pets do you have
how many classes are you taking
proportions and percentages
summarize categorical/qualitative variables
means and medians
preferred measure of center of the distribution for quantitative
observational study
observes individuals and measures variables of interest but does not attempt to influence the values of the variables
purpose is to describe some group or situation
ex: samplings and surveys
why sample?
difficult to find entire population
limited resources
measurements that require destroying the item
census
official count or survey of a population, typically recording various details of individuals; difficult to organize
samples are often used
to make inferences about the population, how you draw it will affect accuracy
chance error/sampling error
random samples can vary from what is expected, in any direction
occurs when a sample is drawn from a population deviates from true population
ex: sample too small for representation
bias or systematic error
systematic error in one direction
ex: sample taken from members of costco when looking for a representative sample of cstat population
sampling frame
list from which the sample is drawn
representative sample
have the same characteristics as the population
convenience sample
individuals who are easily accessible are more likely to be included in the sample; will not be representative; not random
quota sample
first specify your desired breakdown of various subgroups, then reach those targets however you can
selection bias
systematically excluding/favoring particular groups
avoid: examine sampling frame and the method of sampling
response bias
people don’t always respond truthfully
avoid: examine nature of questions and method of surveying
nonresponse bias
people don’t always respond
people decide to take part in the study (self selected)
avoid: keep surveys short and be persistent
people who don’t respond aren’t like the people who do
random/probability sample
reduce bias
estimate the bias and chance error
quantify the uncertainty
for sample to be random:
must be able to provide the chance that any specified set of individuals will be in the sample
all ind in the pop do not need to have the same chance of being selected
you will still be able to measure the errors because you know all the probabilities
not all probability samples are necessarily good
sampling with replacement
once a member of the population is selected, that member is returned to the population for the selection of the next individual
sampling without replacement
member of the population may be chosen only once; not returned to population before next selection
if population is huge compared to sample
random sampling with and without replacement are pretty much the same
probabilities of sampling with replacement
much easier to compute
simple random sampling (SRS)
every member in population has equal chance of being selected in sample; sample selected from list of every unit in population (often difficult)
variation in samples: two or more samples form the same population, taken randomly, and having close to the same characteristics of the population will likely be diff from each other
systemic random sampling
list if members in population, pick random starting point, select every kth member of population
stratified sampling
split population into like groups (strata)
within each strata do random sampling
good for making sure certain members of pop are in sample
cluster sampling
split pop in different clusters (normally by proximity)
perform simple random sample to select clusters
within each cluster, sample every member in the cluster
often used because it is cheaper and easier
non sampling errors
self funded samples: performed to support a certain claim
misleading use of data
in an observational study
we observe a population without applying treatment
in a randomized experiment
we randomly assign the subjects in the study to the treatment
response/dependent variable
measures outcome of study
explanatory/independent variable
may explain or influence changes in a response variable
lurking/confounding variable
not among the explanatory or response variables and is still associated with both explanatory and response variables
why not always use observational studies
cannot conclude cause-effect relationship or causal relationship
subjects/experimental units
individuals studied in an experiment, particularly when they are people
factors
explanatory variables in an experiment
treatment
any specific experimental condition applied to the subjects; if experiment has more than one factor, a treatment is a combination of specific values of each factor
the purpose of an experiment
to investigate a causal relationship between variables
how can we prove that the explanatory variable is causing a change in the response variable?
necessary to isolate the effect of the explanatory variable → randomized experiment
to counter power of suggestion
researchers set aside one treatment group as control group
control group
meant to serve as baseline with which the experimental group is compared; placebo provided
placebo
control treatment that is fake but otherwise indistinguishable from experimental treatment group
placebo effect
improvement in health due not to any treatment but only to the patient/doctor’s belief that he/she will improve
randomized comparative experiment
experiment that uses both comparison of two or more treatments and chance (random) assignment of subjects to treatments
the logic of randomized comparative experiment depends on
ability to treat all the subjects identically in every way except the actual treatments being compared
blinding
preserves power of suggestion
single: subjects don’t know who is receiving real treatment
double: researchers and subjects don’t know (gold standard); necessary when investigator evaluates experimental outcome
blocking
arranging of experimental units in groups (blocks) that are similar to one another; if variable could influence response, should block it