QTM 100 Lecture 2 – Variables and Study Design
A variable is any characteristic observed in a study. Variables can be classified as categorical or quantitative. A categorical variable places observations into categories. Examples include condition such as new or used, shipping method, and eye color. A quantitative variable takes numerical values that represent meaningful amounts.
Quantitative variables can be discrete or continuous. A discrete quantitative variable has a finite or countable number of possible values and often represents a count. Examples include the number of bids in an auction and the number of wheels included with a product. A continuous quantitative variable can take infinitely many possible values within a range. Examples include price, height, weight, and time.
Categorical variables can also have additional classifications. A dichotomous categorical variable has exactly two categories, such as helper versus hinderer or new versus used. An ordinal categorical variable has categories with a natural ordering, such as freshman, sophomore, junior, and senior.
In data analysis, an explanatory variable is used to explain or predict changes in a response variable. The response variable is the outcome of interest. If two variables have a relationship, they are associated. If no association exists, they are independent.
An experimental study assigns subjects to experimental conditions or treatments. An observational study observes explanatory and response variables without assigning treatments. Observational studies can establish association but generally cannot establish causation because confounding variables may exist.
Random assignment is used in experiments and helps researchers make causal conclusions. Random sampling is used to obtain a representative sample and helps researchers generalize results to a population.
Three common random sampling methods are simple random sampling, stratified sampling, and cluster sampling. In a simple random sample, individuals have an equal chance of being selected. In stratified sampling, the population is divided into similar groups called strata, and a random sample is taken from every stratum. In cluster sampling, naturally occurring groups are identified, some clusters are randomly selected, and individuals are sampled from those selected clusters.