Statistical Methods for Computer Science

Experimental and Empirical Research

  • Empirical research involves observation-based investigations to establish causal relationships by manipulating independent variables and measuring dependent variables.

Key steps in experimental/empirical research

  • Pre-Experimental Designs:

    • One-Shot Experimental Case Study: A single group receives treatment followed by observation (Tx→ObsTx \rightarrow \text{Obs}).
    • One-Group Pretest-Posttest Design: A single group is observed before and after treatment (Obs→Tx→Obs\text{Obs} \rightarrow Tx \rightarrow \text{Obs}).
    • Static Group Comparison: Two groups are compared where one receives treatment and the other does not (Group 1:Tx→Obs\text{Group 1}: Tx \rightarrow \text{Obs}, Group 2:−→Obs\text{Group 2}: - \rightarrow \text{Obs}).
  • True-Experimental vs. Quasi-Experimental Designs:

Comparison between true-experimental and quasi-experimental designs

Central Measures and Data Distributions

  • Measures of Central Tendency:
    • Mean: Average value of a set of scores.
    • Median: Middlemost item in an ordered set; unaffected by extreme values.
    • Mode: Most frequently occurring value; unaffected by extreme values and can be unimodal, bimodal, or multimodal.

Mean, median, and mode alignment in skewed and symmetric distributions

  • Central Limit Theorem (CLT):

    • States that the distribution of sample means approaches a normal distribution with standard deviation σ′=σn\sigma' = \frac{\sigma}{\sqrt{n}}.
    • Probability density function equation:   p(x)=12πσ′2e−12(x−μσ′)2p(x) = \frac{1}{\sqrt{2\pi\sigma'^2}} e^{-\frac{1}{2}\left(\frac{x-\mu}{\sigma'}\right)^2}
  • Kernel Density Estimation (KDE): A non-parametric technique for estimating the probability density function of a random variable.

  • Normal vs. Student's t-Distribution:

    • Normal Distribution: Symmetric curve centered at μ=0\mu = 0 with standard deviation σ=1\sigma = 1 (68 %68\,\%- ⁣95 %\!95\,\%- ⁣99.7 %\!99.7\,\% rule).
    • t-Distribution: Uses degrees of freedom (d.f.=n−1d.f. = n - 1); has heavier tails than a standard normal distribution but approaches the standard normal distribution (ZZ-distribution) as n→∞n \rightarrow \infty.

Statistical Hypothesis Testing and T-Tests

  • Assumptions for T-Tests: Continuous or ordinal data, random sampling, normal distribution, sufficient sample size, and equal variance across groups (for independent two-sample tests).

  • Types of T-Tests:

    • One-Sample t-Test: Compares a group mean mm against a population or theoretical mean μ\mu:     t=m−μsnt = \frac{m - \mu}{\frac{s}{\sqrt{n}}}
    • Independent Two-Sample t-Test: Compares means of two independent groups (mAm_A and mBm_B) with d.f.=nA+nB−2d.f. = n_A + n_B - 2:     t=mA−mBS2nA+S2nBt = \frac{m_A - m_B}{\sqrt{\frac{S^2}{n_A} + \frac{S^2}{n_B}}}
    • Paired Sample t-Test: Compares mean differences in repeated measurements on the same subjects with d.f.=n−1d.f. = n - 1:     t=msnt = \frac{m}{\frac{s}{\sqrt{n}}}

Analysis of Variance (ANOVA)

  • Purpose: Evaluates if statistically significant differences exist across three or more group means.
  • ANOVA Variations:
    • One-Way ANOVA: Compares multiple independent groups across a single factor (H0:μ1=μ2=⋯=μkH_0: \mu_1 = \mu_2 = \dots = \mu_k).
    • Two-Way ANOVA Without Replication: Double-tests a single group across two factors.
    • Two-Way ANOVA With Replication: Evaluates multiple groups undergoing multiple treatment conditions to assess both main effects and interaction effects.

Pearson's Chi-Square Test

  • Purpose: Evaluates independence between categorical variables in a contingency table.

  • Calculations:

    • Expected Frequencies (cc):     c=row total×column totalgrand totalc = \frac{\text{row total} \times \text{column total}}{\text{grand total}}
    • Chi-Square Statistic (χ2\chi^2):     χ2=∑(o−c)2c\chi^2 = \sum \frac{(o - c)^2}{c}
    • Degrees of Freedom (d.f.d.f.):     d.f.=(rows−1)×(columns−1)d.f. = (\text{rows} - 1) \times (\text{columns} - 1)
  • Decision Rule: Reject null hypothesis H0H_0 if calculated χ2>critical χ2\chi^2 > \text{critical } \chi^2 value or if p≤αp \le \alpha (α=0.05\alpha = 0.05).