1/62
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
case
an object upon which we collect information that we are interested in studying
respondents
cases that are answering in a survey
subjects/participants
cases who are being studied in an experiment
experimental units
non-human cases being studied in an experiment
variable
a characteristic of a case that differs
identifier variable
provides an identifier for each case, typically a variable that uniquely identifies each case
it’s possible to have more than one
data
collection of observed values of a variable
observation
collection of observed values from a case
data table
table containing sets of data where observations are contained in rows and each variable gets its own column
what are some sources of data?
polls
surveys
experiments
observational studies
census
brain scans
genetic measurements
what are the two different kinds of variables?
Categorical and Quantitative
categorical variable
divides cases into groups, placing each case into exactly one or more categories. It consists of groups or category names.
ex) eye color, marital status, political party
what are the two types of categorical variables?
nominal and ordinal
nominal categorical variables
the order of the categories does not matter
ordinal categorical variables
the ordering of the categories does matter
ex) clothing sizes, final grade
nominal or ordinal: year in college
ordinal
nominal or ordinal: which award would you rather win- nobel prize, gold olympic medal, academy award
nominal
quantitative variable
measures a numerical quantity for each case. It consists of numerical measures or counts
ex) height, temperature, number of children in a family
discrete quantitative variables
can only take on a set number of values
usually a whole number (unless like all of the available options are decimals- what matters is that it can’t just be ANY number)
ex) classes missed in a week (you can’t miss more than 2 because there only ARE two and you can’t miss 1.283 days or something)
continuous quantitative variables
can take on any value within some interval
ex) height, weight, speed
how can we provide context when analyzing the source of data?
we can try to find the:
who
what
where
why
when
how
of the experiment
population
any complete collection of people or objects that a statistician is interested in studying
parameter
value that describes a characteristic of a population
what we actually want to study about the population. what do we want to know?
sample
set of units selected from a population that a statistician analyzes to better understand the population
statistic
value calculated from a sample that serves as an estimate of a parameter
inference/inferential statistics
the process of using data from a sample to gain information about the population
sampling frame
list of cases from which a population is drawn; usually a large subset of the population
like the contact info of all the cases
in order to be valid, samples must be:
representative: characteristics of the sample closely resemble the characteristics of the population
selected randomly: observations are chosen by chance rather than by deliberately and intentionally selecting specific cases
large enough: need enough information to understand the population
sampling error
difference between the sample and population, can be large or small.
larger samples tend to have a smaller sampling error
simple random sample
method of sampling in which every member of the sampling frame has the same chance of being chosen
systematic sample
method of sampling where members are sampled according to some predetermined rule by skipping a certain number of people and then sampling the nth person
systematic sample
method of sampling in which the population is first divided into groups according to some characteristic (called strata) and then a random sample is taken from within each stratum
cluster sample
sampling method where the population is divided into similar groups (called clusters), a simple random sample of the clusters is taken, and then every member in each selected cluster becomes part of the sample
multistage sample
sampling scheme that combines several methods
census
an attempt to collect data on the entire population of interest
convenience sample
a sample obtained from the people who were the easiest to access
voluntary sample
a sample obtained when a large group of individuals is invited to participate, and all responses are recorded; statistician does little to no work other than offer the opportunity to participate
sampling bias
any systematic failure of a sampling method to represent its population
what are the two types of non-sampling bias?
nonresponse bias and response bias
nonresponse bias
introduced to a sample when a large fraction of those sample fails to respond
response bias
anything in the survey that influences responses
survey
method of data collection where respondents are asked questions and self-report responses on various topics
what are some sources of bias in designing survey questions?
complicated questions (more than one part)
vague questions
leading questions
central tendency bias (neutral responses)
error prone response options (answers out of a typical order)
voluntary response bias (some people don’t respond)
response/dependent variable
the variable of interest, what researchers are measuring
explanatory/independent variable
any variable that may influence or explain the response variable
observational study
a method of data collection where researchers allow subjects to live their lives naturally without interfering and record their behavior
looks for associations between the variables that allows researchers to make informed decisions
cannot prove the explanatory variable causes the response to occur
what are the two types of observational studies?
retrospective and prospective
retrospective observational studies
studies that analyze an outcome in the present by delving into historical records
require accurate records/subjects to recall history
more prone to bias
prospective observational studies
studies that identify a set of subjects in advance and collect data in the future as events unfold
researcher is involved in the collection of data
less prone to bias
lurking variable
a third variable that is associated with both the explanatory variable and the response variable
experiment
a type of study where the experimenter manipulates some aspect of one (or more) explanatory variables, assigns them to a subject/experimental unit, and observes a response in the future
factor
another name for an explanatory variable in an experiment
level
one of several possible attributes that can be assigned to a subject/experimental unit
treatment
combination of all factor levels assigned to a subject/experimental unit
control
making conditions as similar as possible for all treatment groups
randomize
each subject/experimental unit is assigned a treatment randomly in an attempt to spread out sources of variability evenly amongst all treatments; this is the most important step in controlling for confounding variables
replicate
taking more than one observation at each factor level to estimate the variability of measurements. when an experiment is repeated in entirety, it is said to be replicated
blocking
grouping subjects/experimental units together based on a factor that cannot be randomized, but may have an effect on the response variable; this can reduce within-treatment variability
completely randomized design
an experimental design where every possible treatment is assigned to at least one subject/ experimental unit
randomized block design
an experimental design where subjects/ experimental units are first divided into their respective blocks and are then assigned treatments within each block
factorial design
an experiment with more than one manipulated factor
a full factorial design that contains treatments for all possible combinations of factors at all levels