Basic Statistics in Research: Definitive Study Notes

Fundamentals of Statistics in Research

Statistics is defined as the science of collecting, organizing, analyzing, interpreting, and communicating data. In academic and applied research, its core purpose is to summarize empirical evidence, evaluate specific research questions, estimate population characteristics, and support reasoned, evidence-based decisions. Statistical methods are widely applied across diverse academic disciplines, including education, psychology, health sciences, business, and the social sciences.

A fundamental guiding principle in research design is that statistical methods must be selected directly based on the research question, study objectives, and research design—never chosen arbitrarily from software options or software menus. Research begins with a clear inquiry, which dictates the necessary data collection, measurement level, and statistical test required to yield valid conclusions.

Research Variables and Selection Standards

Variables are fundamental components of research that enable the quantitative and qualitative measurement and analysis of data. A variable is defined as any characteristic, property, or attribute that can take on different values across individuals, objects, or events over time. When conceptualizing a study, researchers must maintain strict variable relevance, explicitly avoiding the inclusion of demographic or background variables that do not directly pertain to the research objectives or theoretical framework.

For example, in a study investigating Factors Influencing Non-Compliance with Driver’s Licensing Rules Among College Students, relevant profile variables include Age, Sex, College or Program, Year Level, Driving Experience, Type of Vehicle Driven, Frequency of Driving, and Driver’s License Status. Conversely, background characteristics such as Favorite Food, Favorite Color, Religion, Height, and Birth Month are unnecessary and irrelevant to the research topic and should be excluded from the instrument and analysis.

Levels and Scales of Measurement

Variables in statistical research are broadly categorized into categorical variables and continuous variables. Categorical variables include nominal and ordinal scales, while continuous variables encompass interval and ratio scales. Binary, dummy-coded, or dichotomous variables represent a specialized form of nominal variable that contains exactly two distinct categories.

Nominal Measurement Scale: Represents mutually exclusive named categories without any intrinsic order or quantitative ranking. Applicable statistical analyses for nominal data are limited to modes, frequencies, counts, and percentages. Key examples include occupation, political preference, favorite animal, place of residence, sex, religion, favorite color, and school type.

Ordinal Measurement Scale: Incorporates named categories that possess a distinct, meaningful order or rank, but lacks proportionate or equal intervals between levels. Appropriate analytical tools include modes, frequencies, medians, minimum/maximum values, ranges, and percentiles. Key examples include Likert scale items, rank order items (such as 1st1\text{st}, 2nd2\text{nd}, 3rd3\text{rd} place), year level in school, and levels of satisfaction measured across ordered options (Very Dissatisfied, Dissatisfied, Neutral, Satisfied, Very Satisfied).

Interval Measurement Scale: Features named, ordered levels with equal, proportionate intervals between values, but lacks an absolute or true zero starting point. An interval score of zero does not indicate the complete absence of the measured attribute. Appropriate analytical tools include means, variances, and standard deviations. Key examples include temperature measured in Celsius or Fahrenheit, composite means of Likert scales, and specific times of day (such as 1pm1\,\text{pm} or 2pm2\,\text{pm}).

Ratio Measurement Scale: The highest level of measurement, possessing named, ordered levels, equal proportionate intervals, and a meaningful absolute zero starting point. An absolute zero indicates a total absence of the variable being measured. Appropriate statistical tools include means, variances, standard deviations, and exact ratio comparisons. Key examples include temperature in Kelvin, height, weight, time duration (such as 20mins20\,\text{mins} or 1hour1\,\text{hour}), and age.

The structural properties across the four scales of measurement can be summarized by specific benchmark features. Nominal scales possess labeled values only. Ordinal scales possess labeled values and a meaningful order. Interval scales possess labeled values, a meaningful order, and measurable differences between points. Ratio scales possess labeled values, a meaningful order, measurable differences, and a true zero starting point.

Descriptive and Inferential Statistics

Descriptive statistics focuses on summarizing, organizing, and presenting observed sample data in a meaningful way without drawing conclusions beyond the immediate data set. It utilizes statistical metrics such as frequency distributions, percentages, measures of central tendency (mean, median, mode), and measures of dispersion (variance, standard deviation), alongside visual presentations including tables and graphs.

Examples of descriptive research questions include evaluating the level of students' motivation across four specific dimensions: intrinsic motivation, extrinsic motivation, task value, attainment value, and expectancy for success. Descriptive questions also assess levels of student self-efficacy in terms of problem-solving self-efficacy and academic self-efficacy.

Inferential statistics utilizes sample data to make generalized statements, predictions, or inferences about a larger target population. It relies on confidence intervals, hypothesis testing, regression models, and other model-based analyses. Inferential methods require rigorous attention to statistical uncertainty, sampling techniques, and underlying distributional assumptions.

Examples of inferential research questions include determining whether a statistically significant relationship exists between students' motivation and their mathematics performance, assessing whether there is a significant difference in students' mathematics performance when grouped according to demographic profiles, and testing whether specific teaching strategies significantly influence students' mathematics performance.

Populations, Samples, Parameters, and Statistics

A population represents the entire group of individuals, units, or observations that possess the characteristics required to answer a research inquiry. A parameter is a numerical value that describes a specific characteristic of an entire population, such as the population mean mathematics score, represented conceptually as μ\mu.

A sample is a smaller, carefully selected subset of individuals drawn from the population. A statistic is a numerical value computed from sample data that describes a sample characteristic, such as the sample mean mathematics score, represented as Xˉ\bar{X}. Inferential statistics uses sample statistics to estimate population parameters.

Sampling Techniques in Quantitative Research

Sampling methods are divided into probability sampling and non-probability sampling design framework categories. Probability sampling procedures ensure that every unit in the population has a known, non-zero chance of selection, thereby supporting strong probability-based statistical inferences.

Simple Random Sampling: The foundational sampling technique where subjects are selected from a larger population entirely by chance, giving every individual an equal probability of inclusion in the final sample.

Stratified Random Sampling: Involves dividing a heterogeneous population into non-overlapping, homogeneous subgroups known as strata. A simple random sample is subsequently drawn from within each individual stratum. Stratified sampling can utilize equal allocation, where identical sample sizes are drawn from each stratum, or proportional allocation, where sample sizes drawn from each stratum are directly proportional to the relative size of that stratum within the total population.

Systematic Random Sampling: A method where the sample is constructed by selecting every kthk\text{th} unit from an ordered population frame, starting from a randomly selected initial unit. The variable kk represents the sampling interval, and its reciprocal 1k\frac{1}{k} represents the sampling fraction.

Cluster Random Sampling: Involves dividing the population into naturally occurring groups or clusters, typically defined along geographic or administrative boundaries (e.g., Barangays 11 through 3030). A random sample of entire clusters is selected, and measurements are collected from all individual units within the selected sample clusters.

Non-Probability Sampling: Includes convenience sampling, purposive sampling, quota sampling, and snowball sampling. These techniques rely on targeted, non-random selection methods and are applied when study designs, specific participant criteria, or extreme population access limitations necessitate targeted recruitment.

Conceptual Paradigms and Theoretical Frameworks

Research studies rely on visual paradigms to represent relationships between independent variables (IV) and dependent variables (DV). One common paradigm evaluates Independent Variables consisting of Students' Motivation (comprising Intrinsic Motivation, Extrinsic Motivation, Task Value, Attainment Value, and Expectancy for Success) predicting Dependent Variables consisting of Self-Efficacy (comprising Problem Solving Self-Efficacy, Academic Self-Efficacy, and Self-Regulatory Self-Efficacy) and Mathematics Performance measured via test scores or academic grades.

Another conceptual framework evaluates the Independent Variable Adversity Quotient—broken down into Control, Ownership, Reach, and Endurance—predicting Dependent Variables consisting of Mathematics Proficiency and English Proficiency.

A complex organizational paradigm examines the Profile of School Paper Writers (Sex, Year level, Performance in English, Seminars or conferences attended in campus journalism) alongside the Profile of Teacher-Advisers (Age, Sex, Civil Status, Educational Qualification, Seminars or conferences attended in campus journalism) to measure their direct effect on Writing Competence across six distinct domain dimensions: News writing, Editorial writing, Feature writing, Sports writing, Copyreading, and Photojournalism.

Research frameworks can also follow an Input-Process-Output (IPO) model. For example, in an evaluation of training effectiveness, Inputs include Qualified Trainers, Adequate Facilities and Equipment, Time Span, and Respondent Profile/Research Values (Age, Gender, Educational Attainment, Employment Status). The Process consists of conducting surveys using questionnaires, computing mean ratings, and evaluating percentage and frequency distributions. The Output is the measured Effectiveness of the Flight Attendant Training Module in the BS Tourism Program at Systems Plus College Foundation.

Statistical Assumptions and Outlier Diagnostics

Parametric inferential procedures depend on key underlying assumptions regarding data quality and structure. Common statistical assumptions include approximate normality of data distribution where required, independence of observations, homogeneity of variance across groups, and linearity of relationships in model-based designs.

Assumptions must be evaluated using graphical and formal diagnostic methods. Essential diagnostics include inspecting histograms, evaluating Quantile-Quantile (Q-Q) plots, examining boxplots, and conducting formal statistical tests such as the Shapiro-Wilk test. Researchers must not rely on any single test mechanically, but should combine formal tests with substantive domain and design knowledge.

An outlier is an observation that is unusually distant from the general pattern of the data. When encountered, researchers must investigate whether the outlier stems from a data-entry error, a measurement operational failure, an unrecorded procedural anomaly, or a legitimate extreme response. Outliers must never be deleted automatically simply to alter significance results. Any exclusion rule must be formally documented, and researchers should report sensitivity analyses comparing statistical outcomes both with and without influential observations included.

Measurement Reliability, Validity, and Pilot Testing

Reliability refers to the consistency and stability of a measurement scale over time and across items. Internal consistency reliability evaluates how consistently items within a multi-item scale measure a single construct, commonly quantified using Cronbach's alpha coefficient. Other primary forms include test-retest reliability and inter-rater reliability. While reliability is necessary for effective measurement, high reliability alone does not guarantee validity.

Validity refers to the extent to which an instrument accurately measures the specific construct it is intended to measure. Evidence supporting validity includes face and content validity, construct validity, and criterion-related validity evidence.

Pilot testing represents a preliminary administration of research instruments and administrative procedures on a small sample prior to main data collection. It evaluates item clarity, wording, response options, completion time, administration logistics, and operational feasibility. Pilot study participants must closely resemble the target research population. Findings from pilot testing are used to revise or remove ambiguous items.

When a researcher adopts a standardized questionnaire from a previously published study (e.g., an instrument adopted from Marcelino et al., 2021 that underwent prior validation and reliability testing), the adopted questionnaire MUST STILL undergo validity re-evaluation and pilot reliability testing with the new target population sample, because instrument reliability and validity are properties of the data gathered from a specific administration and sample, rather than fixed properties of the questionnaire itself.

Cronbach's Alpha Benchmarks and Survey Design Principles

Internal consistency measured via Cronbach's alpha is interpreted using standardized threshold categories: Cronbach's alpha values above 0.90.9 indicate Excellent reliability; values above 0.80.8 indicate Good reliability; values above 0.70.7 indicate Acceptable reliability; values above 0.60.6 indicate Questionable reliability; values above 0.50.5 indicate Poor reliability; and values below 0.50.5 indicate Unacceptable reliability.

In a comprehensive instrument validation study examining Motivation, Self-Efficacy, and Teaching Strategies, expert validators provided an overall mean validation score of 5.005.00 across all domains. For Motivation, the scale was reduced from 4040 original items to 3131 final items, yielding an overall Cronbach's alpha of 0.860.86 (Intrinsic Motivation reduced from 88 to 77 items, α=0.85\alpha = 0.85; Extrinsic Motivation reduced from 88 to 66 items, α=0.88\alpha = 0.88; Task Value reduced from 88 to 55 items, α=0.84\alpha = 0.84; Attainment Value reduced from 88 to 55 items, α=0.82\alpha = 0.82; Expectancy for Success retained all 88 items, α=0.89\alpha = 0.89). For Self-Efficacy, all 2424 original items were retained across Academic Self-Efficacy (88 items, α=0.90\alpha = 0.90), Problem-Solving Self-Efficacy (88 items, α=0.88\alpha = 0.88), and Self-Regulatory Self-Efficacy (88 items, α=0.89\alpha = 0.89), yielding an overall alpha of 0.890.89. For Teaching Strategies, the scale was reduced from 2424 original items to 2323 final items, yielding an overall alpha of 0.870.87 (Instructional Delivery retained 88 items, α=0.89\alpha = 0.89; Active Learning Strategy retained 88 items, α=0.85\alpha = 0.85; Feedback and Support reduced from 88 to 77 items, α=0.88\alpha = 0.88).

Questionnaire item construction requires avoiding the mixing of positively worded and negatively worded statements within the same survey scale. Questionnaires should maintain statements that are consistently worded in the same direction (e.g., using all positive statements) to minimize respondent confusion, reduce measurement error, and facilitate unambiguous score interpretation.

When scoring 4-point Likert scales, mean ranges map directly to specific qualitative descriptions across research dimensions. A mean score range of 3.504.003.50 - 4.00 corresponds to Strongly Agree, described qualitatively as Highly Motivated, High Level of Self-Efficacy, or Always Practiced. A mean score range of 2.503.492.50 - 3.49 corresponds to Agree, described qualitatively as Motivated, Moderate Level of Self-Efficacy, or Frequently Practiced. A mean score range of 1.502.491.50 - 2.49 corresponds to Disagree, described qualitatively as Slightly Motivated, Low Level of Self-Efficacy, or Rarely Practiced. A mean score range of 1.001.491.00 - 1.49 corresponds to Strongly Disagree, described qualitatively as Not Motivated, Very Low Level of Self-Efficacy, or Never Practiced.

Measures of Central Tendency and Parametric Tests

Measures of central tendency summarize a data distribution using a single typical value. The Mean represents the arithmetic average of all values; it incorporates every observation in its calculation but is highly sensitive to extreme outliers. The Median represents the exact middle value when data are arranged in numerical order; it is resistant to extreme values and is preferred for skewed continuous or ordinal data. The Mode represents the most frequently occurring score in a dataset, making it useful for categorical classifications.

Parametric statistical procedures require continuous dependent variables and adherence to distributional assumptions. Ten common parametric tests used in research include:

  1. One-Sample t-test: Compares the sample mean of a single group against a known or expected population mean or standard passing score (e.g., comparing class Test Score against a fixed Passing Score).

  2. Independent-Samples t-test: Compares the sample means of two separate, independent participant groups (e.g., comparing Test Score between male and female students).

  3. Paired-Samples t-test: Compares the means of two related or repeated measurements taken from the exact same participants (e.g., comparing Pretest Score to Posttest Score for a single class).

  4. One-Way ANOVA (Analysis of Variance): Compares the sample means of three or more separate, independent groups (e.g., comparing Test Score across Year Levels 1, 2, 3, and 4).

  5. Repeated-Measures ANOVA: Compares the means of three or more related measurements collected from the same participants over time or across conditions (e.g., tracking Test Score across 3 distinct time periods).

  6. Pearson Correlation Coefficient (rr): Measures the linear strength and direction of the relationship between two continuous variables, yielding a calculated coefficient ranging from 1.00-1.00 to +1.00+1.00 (e.g., correlating Study Time with Test Score).

  7. Simple Linear Regression: Evaluates whether one continuous independent predictor variable significantly predicts a continuous outcome variable (e.g., evaluating if Study Time predicts Test Score).

  8. Multiple Linear Regression: Evaluates whether two or more independent predictor variables simultaneously predict a single continuous outcome variable (e.g., testing if Study Time and Attendance Rate predict Test Score).

  9. ANCOVA (Analysis of Covariance): Compares group means on a continuous dependent variable while statistically controlling for the confounding effect of one or more continuous covariates (e.g., comparing Posttest Score across Teaching Methods while controlling for initial Pretest Score).

  10. MANOVA (Multivariate Analysis of Variance): Compares independent groups across two or more continuous dependent variables simultaneously (e.g., comparing Teaching Method groups on combined Test Score and Motivation Score outcomes).

ANOVA Computations and Post-Hoc Comparisons

ANOVA evaluates whether statistically significant differences exist among three or more group means by computing an FF statistic, which represents the ratio of between-group variance to within-group variance. A statistically significant omnibus FF test confirms that group means are not all equal, but it does not specify which exact pairs of groups differ significantly.

To identify specific pairwise differences following a significant omnibus ANOVA, researchers apply post-hoc multiple comparison tests (such as Tukey HSD or Scheffé) that control for cumulative Type I error rates across pairwise comparisons.

In an empirical investigation assessing student attitudes toward individuals living with HIV/AIDS across demographic profiles, ANOVA yielded specific variance outputs: For Age (Between-Group Sum of Squares=.925\text{Between-Group Sum of Squares} = .925, df=2\text{df} = 2, Mean Square=.463\text{Mean Square} = .463, F=2.366F = 2.366, p=.095p = .095), the difference was Not Significant, leading to the decision to Accept the Null Hypothesis. For Course (Between-Group Sum of Squares=4.362\text{Between-Group Sum of Squares} = 4.362, df=5\text{df} = 5, Mean Square=.872\text{Mean Square} = .872, F=4.655F = 4.655, p=.000p = .000), the difference was Significant at p < .001, leading to the decision to Reject the Null Hypothesis. For Religion (Between-Group Sum of Squares=.723\text{Between-Group Sum of Squares} = .723, df=3\text{df} = 3, Mean Square=.241\text{Mean Square} = .241, F=1.225F = 1.225, p=.300p = .300), the difference was Not Significant, leading to the decision to Accept the Null Hypothesis. For Address (Between-Group Sum of Squares=1.581\text{Between-Group Sum of Squares} = 1.581, df=3\text{df} = 3, Mean Square=.527\text{Mean Square} = .527, F=2.714F = 2.714, p=.045p = .045), the difference was Significant at p < .05, leading to the decision to Reject the Null Hypothesis. Across all demographic analyses, the Total Sum of Squares was 70.52070.520 with Total df=358\text{df} = 358 (Within-Group df=353\text{df} = 353 for Course and Address; df=356\text{df} = 356 for Age; df=355\text{df} = 355 for Religion).

Follow-up Scheffé post-hoc comparisons across academic courses—which included College of Arts and Sciences (CAS), College of Education (COE), College of Agriculture and Sustainable Development (CASD), College of Health and Sciences (CHS), College of Criminal Justice Education (CCJE), and College of Computing Sciences (CCS)—revealed specific pairwise differences: Comparing COE against CASD yielded a Mean Difference of 0.32055-0.32055 (p=.012p = .012), confirming a Statistically Significant difference. Comparing CASD against CCS yielded a Mean Difference of +0.29910+0.29910 (p=.008p = .008), also confirming a Statistically Significant difference.

Non-Parametric Tests, P-Values, and Effect Sizes

When continuous data violate distributional normality assumptions, or when variables are measured at the ordinal level, non-parametric statistical tests must be substituted for their parametric counterparts:

  1. Mann-Whitney U test replaces the Independent-Samples t-test for comparing two independent groups.

  2. Wilcoxon Signed-Rank test replaces the Paired-Samples t-test for comparing two related measurements.

  3. Kruskal-Wallis H test replaces One-Way ANOVA for comparing three or more independent groups.

  4. Friedman Test replaces Repeated-Measures ANOVA for comparing three or more related measurements.

  5. Spearman Rank Correlation replaces Pearson r for measuring monotonic relationships between ordinal or non-normal variables.

A p-value quantifies the probability of observing sample data as extreme as, or more extreme than, the current result under the assumption that the null hypothesis (H0H_0) is true. Small p-values provide stronger statistical evidence against H0H_0. The standard significance threshold of p < .05 is an established scientific convention, not an absolute boundary. Furthermore, a p-value does NOT represent the probability that the null hypothesis is true. Researchers must verify sample data frequency distributions prior to executing correlation or comparative tests.

Effect size quantifies the actual magnitude or practical strength of an observed difference or relationship, independent of sample size. While large sample sizes can render trivial differences statistically significant, effect sizes clarify practical importance. Cohen's d metric categorizes effect sizes into Very small (d=0.01d = 0.01), Small (d=0.20d = 0.20), Medium (d=0.50d = 0.50), Large (d=0.80d = 0.80), Very large (d=1.20d = 1.20), and Huge (d=2.00d = 2.00). The Phi Coefficient (ϕ\phi) metric categorizes association strength into Negligible (0.00 \text{ to } <0.10), Weak (0.10 \text{ to } <0.20), Moderate (0.20 \text{ to } <0.40), Relatively strong (0.40 \text{ to } <0.60), Strong (0.60 \text{ to } <0.80), and Very strong (0.80 to 1.000.80 \text{ to } 1.00).

Statistical decision-making involves two potential classification errors: A Type I Error occurs when a researcher incorrectly rejects a true null hypothesis (H0H_0), controlled by setting alpha (typically α=.05\alpha = .05). A Type II Error occurs when a researcher fails to reject a false null hypothesis (H0H_0), directly linked to statistical power.

Selecting Statistical Tests and SOP Alignment

Selecting the correct statistical test requires a systematic 6-step framework: Step 1 involves identifying the primary research goal—whether the objective is to describe, compare group means, evaluate associations, or predict outcomes. Step 2 involves identifying the variable types and their exact levels of measurement. Step 3 involves determining the number of groups or repeated measurements and establishing whether samples are independent or related. Step 4 requires formally checking statistical assumptions and data quality. Step 5 involves selecting the matching parametric or non-parametric statistical procedure. Step 6 involves interpreting results using test statistics, exact p-values, effect sizes, and confidence intervals.

The Statement of the Problem (SOP) directly determines the appropriate analytical tool:

SOP 1: Evaluating demographic profiles (such as sex or family monthly income) uses Frequency and Percentage distributions.

SOP 2: Evaluating the general level of student mathematics performance uses Frequency, Percentage, Mean, and Average metrics.

SOP 3: Evaluating student motivation levels across intrinsic and extrinsic dimensions uses Mean and Standard Deviation.

SOP 4: Evaluating the level of effectiveness of the Nueva Vizcaya Provincial Cyber Response Team (NVPCRT) across Coordination and Collaboration and Case Handling and Resolution dimensions uses Mean and Standard Deviation.

SOP 5: Testing for significant differences in handling cybercrime cases when grouped by demographic profile uses Independent t-tests or One-Way ANOVA (or non-parametric Mann-Whitney U or Kruskal-Wallis H tests).

SOP 6: Testing for a significant relationship between the level of capability of the NVPCRT and its level of effectiveness in handling cybercrime cases uses Pearson r or Spearman rank correlation (or Chi-square tests of independence for categorical variables).

SPSS Data Preparation, Coding, and Research Workflow

Statistical Package for the Social Sciences (SPSS) is a statistical software platform used for data management, descriptive analysis, hypothesis testing, regression, and ANOVA. Statistical software executes computations, but the researcher remains fully responsible for evaluating assumptions and interpreting findings correctly.

The complete research statistics workflow follows a sequential 14-step path: Research Problem / SOP \rightarrow Variables & Design \rightarrow Sampling \rightarrow Data Collection \rightarrow Coding \rightarrow Data Entry \rightarrow Data Cleaning \rightarrow Descriptive Analysis \rightarrow Assumption Checks \rightarrow Test Selection \rightarrow Inferential Analysis \rightarrow Effect Size / Confidence Intervals \rightarrow Interpretation \rightarrow Reporting.

Data entry requires establishing numerical coding schemes prior to collection. For example, in an empirical study on Cyber Response Teams: Agency Affiliation is coded as 1=Department of Justice (DOJ) - Office of Cybercrime1 = \text{Department of Justice (DOJ) - Office of Cybercrime}, 2=PNP Anti-Cybercrime Group (PNP-ACG)2 = \text{PNP Anti-Cybercrime Group (PNP-ACG)}, 3=NBI Cybercrime Division (NBI-CCD)3 = \text{NBI Cybercrime Division (NBI-CCD)}, and 4=Other4 = \text{Other}. Length of Service as Cyber Response Team is coded as 1=Less than 1 year1 = \text{Less than 1 year}, 2=13years2 = 1-3\,\text{years}, 3=46years3 = 4-6\,\text{years}, and 4=7years and above4 = 7\,\text{years and above}. Cybercrime Trainings Attended is coded as 1=None1 = \text{None}, 2=12trainings2 = 1-2\,\text{trainings}, 3=34trainings3 = 3-4\,\text{trainings}, and 4=5and above4 = 5\,\text{and above}. Number of Cybercrime Cases Handled is coded as 1=None1 = \text{None}, 2=15cases2 = 1-5\,\text{cases}, 3=610cases3 = 6-10\,\text{cases}, 4=1115cases4 = 11-15\,\text{cases}, and 5=More than 15 cases5 = \text{More than 15 cases}.

In the SPSS interface, Data View displays observed raw values arranged with individual cases in rows and variables in columns. Variable View defines metadata including Variable Name, Type, Width, Decimals, Label, Value Labels, Missing Values, Columns, Alignment, and Measurement Level (Nominal, Ordinal, Scale).

Data cleaning requires checking for completeness, verifying entries against physical questionnaires, identifying duplicate cases, checking impossible values, and examining missing data patterns. Missing values must NEVER be coded automatically as zero, as zero represents a meaningful value on continuous and ratio scales. Treatment of missing data depends on design context, variables, and missingness patterns.

Researchers must actively avoid major statistical pitfalls, including p-hacking (repeatedly running tests until a desired p-value appears), ignoring effect sizes, deleting inconvenient outliers without documented rules, treating correlational findings as causal proof, overinterpreting non-significant results as proof of no effect, and altering data to achieve desired conclusions.

The final checklist for quantitative research requires verifying that instruments strictly align with the SOP, securing expert validation, pilot testing questionnaires, verifying data frequency distributions, checking underlying assumptions, reporting exact statistics, p-values, effect sizes, and confidence intervals, and adhering strictly to the guiding sequence: Question \rightarrow Data \rightarrow Method \rightarrow Evidence \rightarrow Interpretation.