Introduction to Statistics, Research Design, and Variables - Comprehensive Study Guide

Research Lifecycle and Workflow in Behavioral Sciences

Empirical research in the behavioral sciences operates through a structured, multi-stage lifecycle. Each stage informs subsequent decisions and ensures that findings remain scientifically robust and reproducible.

  1. Research Design: The process begins by posing specific research questions based on empirical observations, existing literature, or theoretical gaps. Researchers formulate hypotheses and underlying theories, then select the research design best suited to test these assertions.

  2. Data Collection: Researchers gather primary data directly (e.g., through structured survey administration or laboratory experiments) or acquire secondary data from third-party repositories (e.g., legal or institutional archival datasets).

  3. Data Management: Raw data must be prepared for statistical processing. Key management operations include:

    • Data coding and re-coding.

    • Identifying and managing missing data.

    • Detecting and handling statistical outliers.

    • Evaluating and addressing deviations from normality.

    • Merging separate data files and restructuring data formats.

    • Data cleaning and constructing aggregated variables or multi-item psychometric scales.

  4. Data Analysis: Data are subjected to descriptive statistics to summarize sample properties and inferential statistics to test hypotheses.

  5. Interpretation and Dissemination of Results: Researchers evaluate statistical findings in the context of their theoretical frameworks. Results are documented in formal papers or technical reports and shared with the scientific community through peer-reviewed journal publications and conference presentations.

  6. Continuing Research or Theory Revision: Based on community feedback and empirical outcomes, researchers design follow-up studies to explore new variables, alternative populations, or novel settings. This stage also involves conducting exact or conceptual replications and modifying or refining theories.

Definitions and Core Functions of Statistics

Statistics serves as both a theoretical discipline and an applied methodology. Prominent definitions from foundational literature include:

  • De Veaux et al. (2006, p. 2): "Statistics is a way of reasoning, along with a collection of tools and methods, designed to help us understand the world."

  • Freedman et al. (1978, p. xiii): "Statistics is the art of making numerical conjectures about puzzling questions."

  • Iversen & Gergen (1997, p. 4): "Statistics is a set of concepts, rules, and methods for (1) collecting data, (2) analyzing data, and (3) drawing conclusions from data."

  • Johnson & Tsui (1998, p. 2): "Statistics helps provide a systematic approach for obtaining reasoned answers together with some assessment of their reliability in situations where complete information is unobtainable or not available in a timely manner."

  • Mosteller et al. (1961, p. 2): "Statistics is the art and science of gathering, analyzing, and making inferences from data."

  • Ross (1996, p. 5): "Statistics is the art of learning from data. It is concerned with the collection of data, its subsequent description, and its analysis, which often leads to drawing conclusions."

  • Utts (1996, p. 5): "Statistics is a collection of procedures and principles for gaining and processing information in order to make decisions when faced with uncertainty."

  • Wallis & Roberts (1962, p. 11): "Statistics is a body of methods for making wise decisions in the face of uncertainty."

  • Peszka et al. (2023, pp. 5–7): Statistics in the behavioral sciences is an empirical tool defined as the practice of collecting and analyzing numerical data in sufficiently large quantities to infer proportions in a whole population from a representative sample.

Statistics is divided into two primary branches:

  • Descriptive Statistics: Numerical indexes that organize, summarize, and simplify raw data to convey specific properties of a dataset.

  • Inferential Statistics: Analytical techniques that use sample data and probability theory to draw conclusions about unmeasurable target populations.

Theoretical Frameworks of Probability: Frequentist vs. Bayesian Approaches

Inferential statistics relies on formal probability models. Two main statistical frameworks exist: the Frequentist approach and the Bayesian approach.

Frequentist Approach

  • Probabilities are defined by the set of all possible outcomes of a random experiment (i.e., the sample space) under theoretical distributions.

  • An event is a defined subset of the sample space (e.g., obtaining heads on a coin toss has a probability of 12\frac{1}{2}).

  • Events are treated as binary: any given event either occurs or does not occur.

  • The probability of an event is measured by its observed relative frequency across a high number of repeated experimental trials. Inferential logic depends on comparing observed empirical distributions against theoretical probability distributions.

Bayesian Approach

  • Originally formulated by 18th-century mathematician Thomas Bayes (1701–1761) and later expanded by mathematician Pierre-Simon Laplace (1749–1827).

  • Defines probability not as long-run relative frequencies from theoretical outcome distributions, but as a reasonable expectation or degree of belief given a specific state of prior knowledge.

  • Evaluates evidential probabilities by establishing a prior probability for a hypothesis. This value is updated as new empirical data are acquired, producing a revised posterior probability.

  • Formalized via Bayes' Theorem:

P(H∣E)=P(E∣H)P(H)P(E)P(H|E) = \frac{P(E|H)P(H)}{P(E)}

Bayes Theorem Formula

Where:

  • PP = Probability

  • HH = Hypothesis

  • EE = Evidence

Bayesian inference is applied in behavioral science, law, and public policy, particularly for:

  • Modeling changing event frequencies over time (e.g., shifts in trial conviction rates versus plea bargain resolutions).

  • Weighting criminal evidence to guide decisions regarding prosecution or guilt.

  • Assessing low-frequency, high-consequence events (e.g., police use of lethal force or aviation accidents).

Violation of Independence Assumptions: The Sally Clark Case (1999)

Incorrect probability calculations often stem from invalidly assuming that events are independent. A prominent example is the 1999 trial of Sally Clark, who was prosecuted for murder following the Sudden Infant Death Syndrome (SIDS) deaths of her two infants.

An expert witness testified that the baseline probability of a single infant dying from SIDS was approximately 18,500\frac{1}{8,500} (≈0.00011765\approx 0.00011765). The expert calculated the probability of two SIDS deaths in a single family by multiplying the two individual probabilities:

18,500×18,500=172,250,000≈0.00000014\frac{1}{8,500} \times \frac{1}{8,500} = \frac{1}{72,250,000} \approx 0.00000014

This calculation was mathematically flawed because it assumed the two infant deaths were completely independent events. In reality, shared genetic predispositions and environmental factors establish dependency between incidents. The occurrence of a first SIDS death significantly increases the probability of a subsequent SIDS death in the same family.

Populations, Samples, Parameters, and Statistics

Empirical research relies on distinguishing between target populations and collected samples.

Population, Sample, Parameter, and Statistic Diagram
  • Population: The complete set of measurements for a specified group of interest (e.g., Life Satisfaction Scores for all adult residents in a nation).

  • Sample: A measured subset selected from a specified population.

  • Parameter: A numerical or nominal characteristic of an entire population (e.g., population mean μ\mu, population standard deviation σ\sigma).

  • Statistic: A numerical or nominal characteristic calculated from a sample (e.g., sample mean Xˉ\bar{X}, sample standard deviation SS or S^\hat{S}).

Because entire target populations are usually unmeasurable, sample statistics are used to estimate unknown population parameters. Accuracy improves as sample sizes increase and sampling methods achieve greater population representativeness.

Empirical Sampling Example

  • Total Population: All residents in the United States (N=331,000,000N = 331,000,000).

  • Target Population: All adult residents in the United States (N=258,300,000N = 258,300,000).

  • Population Parameter: True population mean Life Satisfaction Score (LSS\text{LSS}, measured on a 1−71-7 scale): μ=?\mu = ?

  • Sample 1: N1=2,000N_1 = 2,000 adult U.S. residents; calculated sample mean statistic Xˉ1=6.54\bar{X}_1 = 6.54

  • Sample 2: N2=1,850N_2 = 1,850 adult U.S. residents; calculated sample mean statistic Xˉ2=6.37\bar{X}_2 = 6.37

Measurement Theory, Working Models, and Error Structure

Measurement transforms unobservable behavioral phenomena into quantifiable data.

Definition of Measurement

John & Benet-Martinez (2014, p. 474) define measurement as follows:

"The raw data…of the social and behavioral sciences…consist of infinitely minute observations of ongoing behavior and attributes of individuals, social groups, social environments, and other entities or objects that populate the social world. Measurement is the process by which these infinitely varied observations are reduced to compact descriptions or models that are presumed to represent meaningful regularities in the entities that are observed… Accordingly, measurement consists of rules that assign scale or variable values to entities to represent the constructs that are thought to be theoretically meaningful."

Measurement as Working Models

  1. Measurement operationalizes abstract constructs as measurable variables to create simplified representations of real-world phenomena.

  2. Measurement models simplify complex phenomena into working approximations. Like all theoretical models, they are revised or replaced over time as more precise models are developed.

  3. Models represent both individual constructs and the structural relationships among multiple constructs.

Structure of Measurement Error

Observation error represents the difference between an observed measured value and the true actual value of a target variable:

Measurement Error=Random Error+Systematic Error\text{Measurement Error} = \text{Random Error} + \text{Systematic Error}

  • Random Error: Unpredictable, naturally occurring variation present in all empirical measurement.

  • Systematic Error: Predictable, artificial bias introduced into measurement, such as instrument miscalibration or systematic observer bias.

In statistical terminology, datum refers to a single individual measurement, whereas data represents the plural set of measurements that form the foundation of statistical analysis.

Epistemology of Science and Scientific Norms

Empirical research is a systematic, data-driven approach governed by communal standards to achieve three core levels of adequacy:

  1. Critical Adequacy: Evaluating how phenomena or relationships are demonstrated, ensuring appropriate tools and methodologies are selected.

  2. Descriptive and Predictive Adequacy: Establishing whether phenomena exist, identifying their conditions of occurrence, and modeling relational patterns.

  3. Explanatory and Manipulable Adequacy: Identifying underlying causal mechanisms and designing interventions to modify phenomena or relationships.

Merton's Scientific Norms

Sociologist Robert K. Merton established four foundational scientific norms, supplemented by the principle of revisability:

  • Universalism: Scientific claims are evaluated strictly on their empirical merit using preestablished criteria, independent of the researcher's background, credentials, or status. Application: Anyone can conduct valid scientific work without advanced institutional standing.

  • Communality: Scientific knowledge is created collectively and belongs to the public domain. Application: Researchers must transparently share data and findings with peers and the public.

  • Disinterestedness: Researchers seek objective truth and must not let personal convictions, political ideals, or financial interests influence findings. Application: Results must be accepted as dictated by data, without personal spin or bias.

  • Organized Skepticism: All claims, theories, established concepts, and traditional beliefs must be critically questioned. Application: Claims are not accepted at face value; empirical evidence is always required.

  • Revisability: Scientific conclusions remain open to revision whenever new disconfirming evidence emerges.

In practice, researchers sometimes deviate from these norms due to institutional pressures, cognitive biases, or personal incentives.

Science and Empiricism vs. Intuition

Feature

Science / Empiricism

Intuition

Data Base

Uses systematic, principled, critical thinking with large, representative samples.

Uses small, unrepresentative sample sizes, personal anecdotes, gut reactions, hearsay, news, or social media.

Claim Type

Makes probabilistic claims.

Makes deterministic claims.

Evidence Search

Tests falsifiable hypotheses and actively searches for convergent evidence across studies and methodologies.

Avoids searching for additional evidence, particularly disconfirming evidence, resulting in bias.

Flexibility

Continuously open to revision based on data.

Can remain rigidly resolute or be swayed by non-empirical factors.

Clinical Example: A physician who relies solely on clinical intuition may fall prey to confirmation bias or ageism, misdiagnosing patient symptoms. While clinical experience builds valuable domain knowledge, diagnostic decisions must be guided by systematic empirical data (e.g., diagnostic lab tests and systematic exclusion of alternative explanations).

Objectives of Behavioral Science

Behavioral science applies empirical methods to describe, predict, explain, and control human behavior and mental processes.

Scientific Method Models

  • Inductive Model: Observations →\rightarrow Problem Definition →\rightarrow Formulate Hypothesis →\rightarrow Gather Empirical Evidence →\rightarrow Retain or Reject Hypothesis →\rightarrow Build Theory →\rightarrow Publish Results.

  • Deductive Model: Existing Theory →\rightarrow Formulate Specific Hypothesis →\rightarrow Define Problem →\rightarrow Gather Empirical Evidence →\rightarrow Retain or Reject Hypothesis →\rightarrow Support or Modify Theory →\rightarrow Publish Results.

Research Questions, Hypotheses, and Variables

Research Questions

Research questions stem from empirical observations, literature reviews, media analysis, policy assessments, or stakeholder inquiries.

  • Observation: Individuals often express support for consequentialist sentencing policies (focused on future crime reduction), yet act retributively when given the opportunity to punish offenders.

  • Research Question: What primarily drives criminal punishment behavior: retribution, deterrence, or both?

Hypotheses

A hypothesis is a formal, testable statement predicting the relationship between two or more variables.

  • Null Hypothesis (H0H_0): Predicts no effect, no relationship, or no difference between conditions (e.g., sentence lengths assigned to offenders stealing 4,9954,995 will not differ from those assigned for stealing 505505).

  • Alternative Hypothesis (H1H_1): Predicts a specific effect, relationship, or difference (e.g., sentence lengths assigned to offenders stealing 4,9954,995 will be significantly longer than those assigned for stealing 505505).

To be scientifically valid, hypotheses must be:

  • Testable: Concepts must be operationalized into measurable and manipulable concrete variables.

  • Falsifiable: Hypotheses must be framed such that empirical data could potentially disprove them.

Classification of Variables and Scales of Measurement

A variable is any measurable property, element, or outcome that varies in amount or form (e.g., height, test performance, GDP, criminal sentences).

Quantitative vs. Categorical Variables

  • Quantitative Variables: Indicate numerical variations in amount.

    • Continuous Variables: Measured along an unbroken numerical continuum containing fractional or decimal values (e.g., 61.5%61.5\%).

    • Discrete Variables: Measured in distinct whole-number units where intermediate decimal values are not meaningful (e.g., 22 children versus 2.52.5 children).

  • Categorical Variables (Qualitative): Indicate differences in kind or category rather than numerical amount. Groups are assigned numerical codes (e.g., Control Group = 00, Treatment Group = 11; Incarceration = 11, Probation = 00).

Stevens' Typology of Measurement Scales

  1. Nominal: Numbers serve solely as arbitrary labels to categorize items and convey no quantitative information (e.g., Dichotomous: Yes = 11, No = 00; Multilevel: Gender assigned as Male = 11, Female = 22, Non-Binary = 33).

  2. Ordinal: Numbers represent rank order, but intervals between ranks are unequal or unknown (e.g., preference rankings 1st, 2nd, 3rd; risk scores).

  3. Interval: Numbers represent equal intervals between units, but the scale lacks a true absolute zero point (e.g., Temperature in Celsius where 0∘C0^\circ\text{C} does not mean absence of heat; credit scores ranging 200−800200-800).

  4. Ratio: Numbers represent equal intervals and include a true absolute zero point indicating complete absence of the measured construct (e.g., height in inches, duration in seconds, monetary value stolen).

Characteristics of Four Scales of Measurement

Scale Classification Example: Satisfaction with Life Scale (SWLS)

  • Individual Items (rated on a 1−71-7 scale from 1=Strongly Disagree1 = \text{Strongly Disagree} to 7=Strongly Agree7 = \text{Strongly Agree}): Discrete quantitative variables measured on an ordinal scale.

  • Aggregated Sum Score (summing 5 items, range 5−355-35): Discrete quantitative variable measured on an ordinal scale.

  • Aggregated Mean Score (averaging 5 items, range 1−71-7 with intermediate decimals such as 6.546.54): Continuous quantitative variable treated as an interval scale.

Data Architecture and SPSS Variable Specification

In statistical software such as IBM SPSS Statistics, variables are configured within the Variable View interface across 10 parameters: Type, Width, Decimals, Label, Values, Missing, Columns, Align, Measure, and Role.

SPSS Variable Creation Mapping

Variable Type Mapping in SPSS

Variable Category

Numerical Coding

SPSS Variable Type

Measurement Scale

SPSS Scale Classification

Categorical

No numerical codes

String

Nominal

Nominal

Categorical

Coded numerically

Numeric

Nominal

Nominal

Quantitative (Discrete)

Coded numerically

Numeric

Ordinal

Ordinal

Quantitative (Continuous)

Continuous numbers

Numeric

Interval or Ratio

Scale

An exemplary SPSS dataset mapping includes:

  • Accused_ID: String variable, Nominal measurement scale.

  • Case_Num: String variable, Nominal measurement scale.

  • Age: Numeric variable, Ratio measurement scale (SPSS Scale).

  • Gender: Numeric variable with value labels (1=Male1 = \text{Male}, etc.), Nominal scale.

  • Race: Numeric variable with value labels (1=Asian American1 = \text{Asian American}, etc.), Nominal scale.

  • LSMCI_Tot: Numeric variable (0−420-42 score), Ordinal scale.

  • Arrest_Charge_Tot: Numeric variable, Ratio measurement scale (SPSS Scale).

  • Missing values are defined across variables using a designated code (e.g., −99-99).

Experimental Demonstration: Retributive vs. Deterrence Drivers

To test whether punishment is driven by retribution or deterrence, researchers operationalize independent variables to examine their effects on assigned prison sentence lengths (0−600-60 months).

  • Retributive Factor (IV 1): Magnitude of harm, operationalized as low (505505) versus high (4,9954,995) monetary value stolen.

  • Deterrence Factor (IV 2): Probability of detection via audit, operationalized as low (15%15\%) versus high (85%85\%) detection probability.

  • Dependent Variable (DV): Length of assigned prison sentence (0−600-60 months).

When harm is treated categorically (505505 vs. 4,9954,995 stolen), results show a mean sentence of 5.905.90 months for low harm compared to 22.2422.24 months for high harm. When harm is analyzed as a continuous quantitative variable (0−50000-5000 dollars stolen), a positive linear relationship emerges, supporting rejection of the Null Hypothesis.

Critical Thinking in Empirical vs. Legal Frameworks

  • Scientific / Empirical Thinking: Probabilistic, centered on uncertainty, quantified using pp-values, confidence intervals, and effect sizes, and subject to ongoing revision.

  • Legal Thinking: Deterministic, relies on deductive logic and binary outcomes (e.g., liable vs. not liable; guilty vs. not guilty), and is less amenable to empirical revision.

  • Societal Context: As noted by Wilks (1951), quoting H.G. Wells: "Statistical thinking will one day be as necessary for efficient citizenship as the ability to read and write."

Research Methodologies and Taxonomy

Research methods can be categorized along three primary dimensions:

  1. Quantitative vs. Qualitative: Quantitative research operationalizes variables using numerical scales to maximize precision and predictive adequacy. Qualitative research uses non-numerical, categorical approaches.

  2. Basic vs. Applied: Basic research seeks to build theoretical knowledge without immediate practical application. Applied research addresses specific real-world problems.

  3. Laboratory vs. Field: Laboratory research occurs in controlled environments to maximize internal validity. Field research occurs in naturalistic settings to maximize external and ecological validity.

Four Categories of Research Methods

Ordered by their strength in drawing causal inferences:

  1. Observational / Descriptive Methods: Researchers record natural phenomena to document their existence and characterize baseline population trends (e.g., tracking public support for sentencing reform on a 1−91-9 scale from 20102010 to 20202020).

  2. Correlational Methods: Researchers measure two or more variables to evaluate the strength and direction of their linear association (−1.00-1.00 to +1.00+1.00).

    • Methodological Limitation: Correlation does not equal causation. Confounding third variables (ZZ) can drive spurious relationships between XX and YY (e.g., high outdoor temperatures ZZ drive both increased ice cream consumption XX and violent crime rates YY).

    • Conditions for Inferring Causality: Establishing causality requires (1) an empirical association between XX and YY, (2) temporal precedence (XX occurs before YY), (3) elimination of all plausible alternative explanations, and (4) a logical theoretical rationale.

  3. Experimental Methods: Researchers manipulate one or more independent variables, measure dependent variables, and exert strict experimental control over extraneous confounds.

    • Factorial Experimental Assignment: Participants are randomly assigned across conditions. In a 2×22 \times 2 factorial design crossing Deterrence (15%15\% vs. 85%85\% audit probability) and Harm (505505 vs. 4,9954,995 stolen), N=180N = 180 participants would be allocated equally (n=45n = 45 per cell).

    • Advanced Variable Roles: Path models evaluate covariates (measured variables controlled for potential influence on the DV), moderators (variables that alter the strength or direction of the IV-DV relationship via interaction terms X×WX \times W), and mediators (variables explaining the indirect mechanisms connecting an IV to a DV).

  4. Quasi-Experimental Methods: Evaluate the causal impact of an intervention without random assignment.

Hierarchy of Quasi-Experimental Designs

An example of an Interrupted Time Series Design is analyzing state prison populations across regular intervals before and after a policy change (e.g., Nebraska prison population trends from 19801980 to 20192019 surrounding the enactment of legislative bill LB605).

Methodological Validity and Selection Frameworks

Types of Validity

  • Internal Validity: The degree to which observed changes in a dependent variable can be unequivocally attributed to the manipulated independent variable, free from confounding factors.

  • External Validity: The degree to which empirical results generalize to other populations, settings, and times.

  • Construct Validity: The extent to which operationalized measures and manipulations accurately represent theoretical constructs.

  • Statistical Conclusion Validity: The extent to which appropriate statistical techniques are used to draw valid inferences about variable covariation without violating probability assumptions.

Research Design Trade-Offs: Exercising tight laboratory control increases internal validity but often reduces external ecological validity. Conversely, field studies maximize external validity at the expense of internal control. No single study maximizes all validities simultaneously.

Decision Decision Frameworks for Statistical Selection

Selecting appropriate statistical techniques requires evaluating the research question, design structure, and scale properties of the data.

Decision Tree for Descriptive StatisticsDecision Tree for Inferential StatisticsSummary Table of Introductory Statistics

Statistical Procedures Reference

Descriptive Statistics

  • Frequency Distributions: Summarize raw data frequencies (ff, nn, percentages %\%%).

  • Central Tendency: Measures of typicality, including Mean, Median, and Mode.

  • Variability: Measures of dispersion, including Range, Interquartile Range (IQR), Standard Deviation (SS), Variance (S2S^2), Skewness, and Kurtosis.

  • Effect Sizes: Indexes quantifying the magnitude of an effect or proportion of variance explained, including Cohen's dd, Eta squared (η2\eta^2), Partial eta squared (ηp2\eta_p^2), Epsilon squared (ϵ2\epsilon^2), Phi (ϕ\phi), Cramér's V (ϕc\phi_c), and Rank-biserial correlation (rrbr_{rb}).

  • Single Score Evaluation: Evaluating extreme individual values using standardization (zz-scores).

Parametric Inferential Statistics (Normal Distribution Assumed)

  • Pearson's Product-Moment Correlation (rr): Measures linear association between two continuous quantitative variables.

  • Linear Regression: Models the predictive relationship between one continuous quantitative IV and one continuous quantitative DV.

  • Binary Logistic Regression: Predicts a single dichotomous categorical DV using continuous or categorical IVs.

  • Student's tt-Test: Compares means across groups for a continuous DV (One-Sample tt-test, Independent-Samples tt-test, Paired-Samples tt-test).

  • Analysis of Variance (ANOVA / FF-Test): Compares means across three or more conditions (One-Way Independent ANOVA, Repeated-Measures ANOVA, Factorial ANOVA).

  • General Linear Model (GLM / ANCOVA): Extends ANOVA to incorporate continuous control covariates alongside categorical IVs.

  • Multivariate Analysis of Variance (MANOVA / MANCOVA): Evaluates effects across two or more continuous DVs simultaneously.

Advanced Structural Models

  • Path Analysis & Mediation: Tests direct and indirect mechanisms connecting IVs to DVs through intermediate mediator variables.

  • Factor Analysis: Evaluates underlying latent variable structures using Exploratory Factor Analysis (EFA) or Confirmatory Factor Analysis (CFA).

  • Hierarchical Linear Modeling (HLM): Analyzes nested data structures (e.g., students within classrooms) and longitudinal repeated measures.

Non-Parametric Inferential Statistics (Distribution-Free)

  • Chi-Square (χ2\chi^2) Tests: Evaluates categorical frequency distributions (Goodness-of-Fit Test and Test of Independence).

  • Wilcoxon Signed-Rank Test / Sign Test: Non-parametric equivalent of paired-samples tt-tests.

  • Mann-Whitney UU Test: Non-parametric equivalent of independent-samples tt-tests.

  • Kruskal-Wallis Test: Non-parametric equivalent of one-way independent ANOVAs.

  • Mood's Median Test: Evaluates differences among group medians.

  • Friedman Test: Non-parametric equivalent of repeated-measures ANOVAs.

  • Spearman's Rank Correlation (rsr_s): Evaluates monotonic relationships between ordinal variables.