Data Processing and Statistical Analysis

Data Preparation Overview

  • Data preparation serves as the critical bridge between data collection and data analysis.
  • Raw data collected from survey questionnaires, laboratory analyses, or field assessments are rarely ready for immediate analysis, as they frequently contain errors, omissions, inconsistent formatting, or invalid responses.
  • The sequential pipeline for data preparation follows this path:
    • Raw Data Collection →\rightarrow Data Cleaning →\rightarrow Handling Missing Values →\rightarrow Variable Coding →\rightarrow Analysis-Ready Dataset.

Data Cleaning Procedures

  • Data cleaning involves inspecting the raw dataset to identify and correct errors that could skew analytical results.
  • Format Standardization:
    • Ensures consistent measurement units across the dataset (e.g., converting all weight values to kilograms).
    • Standardizes date formats across all observations.
    • Establishes uniform text entries.
  • Duplicate Detection:
    • Identifies and removes accidental repeated submissions or double entries.
  • Outlier Screening:
    • Identifies extreme or unusual values.
    • Outliers can originate from data entry errors (e.g., entering an age of 250250 instead of 2525).
    • Outliers can also represent genuine extreme biological or behavioral variation.

Managing Missing Values

  • Missing data occurs when no value is stored for a variable in an observation.
  • Correctly handling missing data requires diagnosing the underlying missingness mechanism:
    • Missing Completely at Random (MCAR):
    • The probability of missing data is completely unrelated to any observed or unobserved variable.
    • Example: A laboratory sample tube accidentally breaks.
    • Missing at Random (MAR):
    • Missingness depends on observed variables but does not depend on the missing value itself.
    • Example: Older participants are less likely to complete an online dietary survey.
    • Missing Not at Random (MNAR):
    • Missingness depends directly on the unobserved value itself.
    • Example: Individuals with high household income intentionally decline to report their income.
  • Remediation Strategies for Missing Data:
    • Listwise Deletion (Complete Case Analysis):
    • Discards any observation containing missing values.
    • Simple to perform, but reduces overall sample size and introduces bias if data is not MCAR.
    • Pairwise Deletion:
    • Utilizes all available data for each specific statistical test.
    • Preserves sample size, but can lead to mathematically inconsistent sample bases across different tests.
    • Mean/Median Imputation:
    • Replaces missing values with the variable's overall mean or median.
    • Quick to execute, but artificially shrinks variance and distorts standard errors.
    • Advanced Imputation (Multiple Imputation / Regression Imputation):
    • Uses statistical models based on observed variables to estimate missing values.
    • Preserves sample variability and represents the preferred approach for complex datasets.

Coding and Transforming Variables

  • Variable coding translates raw qualitative or quantitative data into structured numerical representations suitable for software processing.
  • Categorical Coding:
    • Assigns numerical codes to discrete non-numeric categories.
    • Examples: Assigning 0=Male0 = \text{Male} and 1=Female1 = \text{Female}; or assigning 1=Low Income1 = \text{Low Income}, 2=Middle Income2 = \text{Middle Income}, and 3=High Income3 = \text{High Income}.
  • Dummy / One-Hot Encoding:
    • Converts a categorical variable with kk categories into k−1k - 1 binary indicators (00 or 11) for inclusion in regression models.
  • Recoding and Binning:
    • Collapses continuous variables into meaningful ordinal bands.
    • Example: Converting continuous Body Mass Index (BMI) values into standard World Health Organization categories (Underweight, Normal, Overweight, Obese).
  • Variable Normalization and Scaling:
    • Applies log transformations, z-score standardization, or min-max scaling.
    • Used to resolve non-normal distributions or to standardize measurement scales across different variables.

Descriptive Statistics

  • Descriptive statistics summarize and organize the features of a specific dataset, offering a clear numerical narrative before complex modeling or inference takes place.
  • Measures of Central Tendency:
    • Central tendency metrics locate the numerical center or typical value of a distribution.
    • Mean:
    • Definition: The arithmetic average of all values in a dataset.
    • Best Used For: Continuous data with symmetric, normal distributions. It is sensitive to extreme outliers.
    • Median:
    • Definition: The exact middle score when data points are arranged in numerical order.
    • Best Used For: Skewed continuous data or ordinal data (e.g., income, house prices).
    • Mode:
    • Definition: The value that appears most frequently in the dataset.
    • Best Used For: Nominal data and categorical classifications (e.g., primary language).
  • Measures of Dispersion:
    • Dispersion metrics quantify the spread, variability, or scatter of data points around the central value.
    • Variance: The average squared distance of data points from the sample mean.
    • Standard Deviation (ss or SD\text{SD}): The square root of variance, expressed in the exact same physical units as the original observations.
    • Interquartile Range (IQR\text{IQR}): The distance between the 75th percentile (Q3Q_3) and the 25th percentile (Q1Q_1), representing the central 50%50\% of the dataset.
  • Confidence Intervals (CI):
    • Provides an estimated range within which an unknown population parameter is expected to fall, accompanied by a specified probability level (typically 95%95\%).

Inferential Statistics

  • Inferential statistics allows researchers to test hypotheses, draw generalizable conclusions, and make predictions about a target population based on data drawn from a sample.
  • Parametric and Non-Parametric Framework:
    • The choice of statistical test depends directly on the measurement scale of the variables and the distribution of the underlying data.
    • Parametric Tests:
    • Require strict distributional assumptions.
    • Normality: Data are normally distributed (evaluated via tests such as Shapiro-Wilk or Kolmogorov-Smirnov).
    • Homogeneity of Variance: Variances across comparison groups are equal (evaluated via Levene's test).
    • Interval or Ratio Scale: Variables are measured on a quantitative continuous scale.
    • Non-Parametric Tests:
    • Distribution-free techniques used when distributional assumptions are violated, sample sizes are small (n<30n < 30), or variables are measured on nominal or ordinal scales.
  • Comparative Overview of Inferential Statistical Tests:
    • Comparing 22 Independent Groups:
    • Parametric Test: Independent Samples t-test
    • Non-Parametric Equivalent: Mann-Whitney U test
    • Comparing 22 Paired/Matched Groups:
    • Parametric Test: Paired Samples t-test
    • Non-Parametric Equivalent: Wilcoxon Signed-Rank test
    • Comparing ≥3\ge 3 Independent Groups:
    • Parametric Test: One-Way ANOVA
    • Non-Parametric Equivalent: Kruskal-Wallis test
    • Comparing ≥3\ge 3 Repeated Measures:
    • Parametric Test: Repeated Measures ANOVA
    • Non-Parametric Equivalent: Friedman test
    • Evaluating Association Between 22 Categorical Variables:
    • Parametric Test: N/A
    • Non-Parametric Equivalent: Chi-Square (χ2\chi^2) Test of Independence
    • Evaluating Linear Correlation:
    • Parametric Test: Pearson Correlation Coefficient (rr)
    • Non-Parametric Equivalent: Spearman Rank Correlation (ρ\rho)
  • Regression Models:
    • Regression analysis estimates the relationship between a dependent outcome variable and one or more independent predictor variables.
    • Simple Linear Regression: Models a continuous outcome (YY) based on a single continuous or binary predictor (XX).
    • Multiple Linear Regression: Incorporates multiple predictors (X1,X2,…,XpX_1, X_2, \dots, X_p) to isolate independent predictor effects while controlling for potential confounders.
    • Binary Logistic Regression: Models the log-odds of a binary categorical outcome (Y=1Y = 1 versus Y=0Y = 0), estimating Odds Ratios (OR\text{OR}).

Software Applications in Data Analysis

  • Modern statistical analysis relies on software suites that enable data processing, execution of complex models, and reproduction of results.
  • SPSS:
    • Primary Use Case: Social Sciences, Health Studies, Introductory Research.
    • Key Strengths: Graphical user interface (GUI); low learning curve; point-and-click operations.
    • Considerations: Proprietary commercial license; limited flexibility for advanced custom algorithms.
  • R:
    • Primary Use Case: Academic Research, Data Science, Advanced Biostatistics.
    • Key Strengths: Free, open-source; vast ecosystem of statistical packages (CRAN); state-of-the-art graphics (ggplot2).
    • Considerations: Command-line driven; steep learning curve for non-programmers.
  • Stata:
    • Primary Use Case: Public Health, Epidemiology, Economics, Social Sciences.
    • Key Strengths: Powerful built-in epidemiological and longitudinal data analysis commands; excellent documentation.
    • Considerations: Proprietary commercial license; graphical output requires specialized syntax to customize.
  • Python:
    • Primary Use Case: Machine Learning, Big Data Integration, Predictive Analytics.
    • Key Strengths: Versatile programming language; libraries like pandas, scipy, statsmodels, and seaborn.
    • Considerations: Requires foundational programming proficiency.

Bioethics and Ethical Considerations

  • Core Ethical Principles:
    • Ethics in biomedical and public health research is anchored in four fundamental principles, codified in milestone frameworks such as the Belmont Report (1979) and the Declaration of Helsinki (1964, continuously updated).
    • Autonomy (Respect for Persons):
    • Individual agency must be acknowledged and protected.
    • Potential participants must enter research voluntarily, with adequate information to make an informed choice.
    • Vulnerable populations with diminished autonomy (e.g., minors, cognitively impaired individuals, incarcerated persons) require additional safeguards against coercion or undue influence.
    • Beneficence:
    • Researchers are obligated to secure the well-being of participants by maximizing potential benefits (direct individual benefits or broader scientific knowledge) while minimizing potential risks.
    • Non-Maleficence:
    • Derived from the Hippocratic oath ("do no harm").
    • Mandates that risks must never be unreasonably disproportionate to anticipated benefits.
    • Experimental procedures must be designed to avoid physical, psychological, or socio-economic injury.
  • Institutional Review Boards (IRB):
    • An Institutional Review Board (IRB)—also referred to as an Independent Ethics Committee (IEC)—is an officially designated committee responsible for approving, monitoring, and reviewing biomedical and behavioral research involving humans.
    • Clearance and Approval Workflow:
    • Protocol Preparation: Submission of a detailed dossier including the scientific protocol, data collection tools, safety monitoring plans, recruitment materials, and consent forms.
    • Risk Assessment Categories:
      • Exempt Review: Minimal risk, involving anonymized data or routine educational/survey procedures.
      • Expedited Review: Minimal risk with non-invasive collection methods (e.g., standard dietary monitoring, small blood draws in healthy adults).
      • Full Board Review: Greater than minimal risk, interventions, biological sampling in sensitive areas, or inclusion of vulnerable populations.
    • Formal Evaluation: The board reviews scientific validity, risk-benefit balance, privacy protections, and consent mechanisms.
    • Approval and Continuing Review: Protocol approval must be secured prior to initiating field activity or data collection. Studies require periodic renewal (typically annual) and immediate reporting of protocol amendments or adverse events.
  • Informed Consent:
    • Informed consent is an ongoing operational process of communication, not merely a signed document.
    • Dimensions and Requirements:
    • Information: Full disclosure of study objectives, procedures, expected duration, potential risks/discomforts, direct benefits, and data confidentiality measures.
    • Comprehension: Information must be written in plain, non-technical language tailored to local literacy levels and translated into local dialects where necessary.
    • Voluntariness: Explicit statement that participation is entirely optional, and refusal or withdrawal at any stage carries no penalty or loss of standard medical care.
    • Assent (Minors): For pediatric populations (typically ages 7–177\text{--}17), written or verbal assent must be obtained from the child in tandem with formal written consent from legal guardians.
  • Scientific Integrity and Research Misconduct:
    • Scientific progress relies on public trust and absolute fidelity to data integrity. Research misconduct undermines the credibility of the broader scientific enterprise.
    • Fabrication: Making up data or research findings and reporting or recording them as real.
    • Falsification: Manipulating research materials, equipment, or processes, or altering or omitting data such that the research record is not accurately represented.
    • Plagiarism: The appropriation of another person's ideas, text, methodologies, or data without giving appropriate credit or attribution (including self-plagiarism).
    • Conflicts of Interest (COI): Financial, institutional, or personal ties that could bias or appear to bias scientific judgment. Transparency requires explicit declaration of all funding sources and competing interests upon protocol submission and journal publication.