Comprehensive Study Notes on Descriptive Statistics, Probability Theory, and Hypothesis Testing

Descriptive Epidemiology and Population Concepts

  • Descriptive Epidemiology Objectives:

    • Allows the description of specific characteristics within a given population, such as health-related problems.

    • Serves to estimate characteristics in a larger general population based on observed data from a study sample.

    • Crucial Limitation: Direct conclusion/extrapolation to the overall population is not strictly possible unless specific sampling conditions are met. Study results apply directly to the study sample, and researchers must be vigilant regarding potential biases.

  • Target Population vs. Study Sample:

    • Population (PP): Large-scale total group of individuals for which specific characteristics are sought to be studied. The basic element is the individual subject.

    • Sample (EE): Restricted subgroup belonging to the larger target population upon which measurements and studies are actually conducted to simplify feasibility.

  • Variability and Defining Parameters:

    • Inter-subject Variability: Represents differing responses or measurements among individual subjects to an identical prompt or parameter (e.g., individual season preferences: summer vs. winter).

    • Essential Context Boundaries: Any population characteristic or study variable must strictly be defined according to three mandatory dimensions:

      • Time (e.g., year 2025)

      • Place (e.g., Pau)

      • Person (e.g., P1 health science students)

  • Conditions for General Population Estimation:

    • To estimate a characteristic in a broader population from a sample, the sample must be sufficiently representative of that population.

    • Representativeness requires:

      • A sufficiently large sample size (NN).

      • Random sampling selection (drawing subjects at random).

Variable Types and Data Presentation

  • Classification of Variables:

    • Quantitative Variables (Numerical):

      • Discrete Quantitative Variables: Possess a finite, countable number of distinct values. Values are strictly whole/round numbers (e.g., number of P1 students = 10001000, never 1000.41000.4; number of pin's won).

      • Continuous Quantitative Variables: Possess an infinite number of possible real values within a given range (e.g., age, height, weight, biological parameters, blood cholesterol, active ingredient quantity in medicine).

    • Qualitative Variables (Categorical):

      • Represent non-numerical categories or modalities (e.g., sex, hair color, profession, diabetes status [yes/no, Type 1/Type 2]).

      • For analytical convenience, numerical codes can be assigned to qualitative modalities (e.g., coding disease status as 11 for diseased and 00 for non-diseased).

  • Data Summarization Techniques:

    • Raw Data Series: Unstructured lists of values (e.g., 69,75,60,71,72,71,11,73,77,88,66,73,7569, 75, 60, 71, 72, 71, 11, 73, 77, 88, 66, 73, 75) provide no synthetic overview and prevent meaningful conclusions.

    • Summarized Information: Databases aggregate data into synthetic numerical parameters (central tendency and dispersion) or graphical representations to expose variable distributions.

  • Graphical Representation Methods:

    • Bar Chart (Diagramme en barres): Standard for qualitative/categorical variables, as well as discrete quantitative variables.

    • Histogram: Designed for continuous quantitative variables, plotting frequency or percentage against continuous class intervals.

Descriptive Statistics for Quantitative Variables

  • Measures of Central Tendency (Position):

    • Must always be expressed with their proper measurement units.

    • Mean (mm or μ\mu): Arithmetic average of all observed values, divided by total observation count NN:

m=i=1NXiNm = \frac{\sum_{i=1}^N X_i}{N}

*   **Median (MM)**: Value that divides an ordered sequence of numbers into two equal halves (50% of values are greater, 50% are lower).
    *   If dataset size NN is **odd**: Median position is N+12\frac{N + 1}{2}.
    *   If dataset size NN is **even**: Median is the mean of the values at positions N2\frac{N}{2} and N2+1\frac{N}{2} + 1.
*   **Mode**: Most frequently occurring value in the dataset. A distribution may have no mode, a single unique mode, or multiple modes.
  • Distribution Shape and Central Tendency Relationships:

    • Symmetrical Distribution (Gaussian / Normal): Mode = Median = Mean.

    • Asymmetrical Distribution (Skewed): Mode, Median, and Mean take distinct separate values.

  • Comparative Evaluation of Central Tendency Indicators:

    • Mean (mm):

      • Advantages: Easy to calculate; widely recognized and universally understood.

      • Disadvantages: Highly sensitive to extreme values/outliers; poorly represents heterogeneous or bimodal populations.

    • Median (MM):

      • Advantages: Unaffected by extreme values; robust indicator for asymmetrical distributions.

      • Disadvantages: Ignores total data distribution details (extreme ranges); manually tedious to determine when NN is large.

    • Mode:

      • Advantages: Effectively represents heterogeneous or multimodal populations.

      • Disadvantages: Manually tedious to identify for large NN; fluctuates based on grouping class widths.

  • Measures of Dispersion (Variation):

    • Informs on the spread or variation of values around the central tendency. Must always retain measurement units.

    • Range (EE): Simple difference between maximum and minimum observed values:

E=ValuemaxValueminE = \text{Value}_{\max} - \text{Value}_{\min}

*   **Variance (σ2\sigma^2 or s2s^2)**: Mean of squared deviations from the arithmetic mean. Quantifies observation spread around the mean. Unit is **squared** (unit2\text{unit}^2).
    *   Population Variance:

σ2=i=1N(xiμ)2N\sigma^2 = \frac{\sum_{i=1}^N (x_i - \mu)^2}{N}

    *   Sample Variance:

s2=i=1N(xim)2Ns^2 = \frac{\sum_{i=1}^N (x_i - m)^2}{N}

*   **Standard Deviation (σ\sigma or ss)**: Square root of the variance. Restores the original measurement unit.
    *   Population Standard Deviation: σ=σ2\sigma = \sqrt{\sigma^2}
    *   Sample Standard Deviation: s=s2s = \sqrt{s^2}
*   **Percentiles (Centiles)**: The NN\-th percentile is the value below which N%N\% of observations lie.
    *   **Quartiles**: Split data into 4 equal parts (Q1=25%Q_1 = 25\%, Q2=50%Q_2 = 50\% [Median], Q3=75%Q_3 = 75\%).
    *   **Terciles**: Split data into 3 equal parts (T1=33.3%T_1 = 33.3\%, T2=66.6%T_2 = 66.6\%).
    *   **Deciles**: Split data into 10 equal parts (D1=10%D_1 = 10\%, D2=20%D_2 = 20\%, ..., D9=90%D_9 = 90\%).
  • Summary Recommendation for Parameter Reporting:

    • For Symmetrical Distributions: Report Mean + Standard Deviation.

    • For Asymmetrical Distributions: Report Median + Percentiles (e.g., Interquartile Range).

  • Qualitative Variable Description:

    • Frequency (ff): Proportion of individuals in a specific category relative to total sample size NN:

f=nNf = \frac{n}{N}

*   Where nn is the category effectif (count) and NN is the total population count. The sum of all category frequencies strictly equals 11.

Detailed Practical Examples and QCM Applications

  • Example 1: Cookie Consumption Analysis:

    • Scenario: 1010 P1 students discuss daily cookie consumption. 22 eat 33 cookies/day, 33 eat 11 cookie/day, and 55 eat 00 cookies/day.

    • Ordered Dataset: 0,0,0,0,0,1,1,1,3,30, 0, 0, 0, 0, 1, 1, 1, 3, 3 (N=10N = 10).

    • Mean Calculation:

m=(2×3)+(3×1)+(5×0)10=6+3+010=0.9cookies/daym = \frac{(2 \times 3) + (3 \times 1) + (5 \times 0)}{10} = \frac{6 + 3 + 0}{10} = 0.9\,\text{cookies/day}

*   *Median Calculation*: N=10N = 10 (even). Midpoint between 5th value (00) and 6th value (11):

M=0+12=0.5cookies/dayM = \frac{0 + 1}{2} = 0.5\,\text{cookies/day}

*   *Mode*: 0cookies/day0\,\text{cookies/day} (occurs 55 times).
*   *Distribution Shape*: Asymmetrical (Mode = 00, Median = 0.50.5, Mean = 0.90.9).
*   *Range*:

E=30=3cookies/dayE = 3 - 0 = 3\,\text{cookies/day}

*   *Variance Calculation*:

σ2=2(30.9)2+3(10.9)2+5(00.9)210\sigma^2 = \frac{2(3 - 0.9)^2 + 3(1 - 0.9)^2 + 5(0 - 0.9)^2}{10}

σ2=2(2.1)2+3(0.1)2+5(0.9)210=2(4.41)+3(0.01)+5(0.81)10=8.82+0.03+4.0510=1.4333(cookies/day)2\sigma^2 = \frac{2(2.1)^2 + 3(0.1)^2 + 5(-0.9)^2}{10} = \frac{2(4.41) + 3(0.01) + 5(0.81)}{10} = \frac{8.82 + 0.03 + 4.05}{10} = 1.4333\,(\text{cookies/day})^2

*   *Standard Deviation Calculation*:

σ=1.4333=1.1972cookies/day\sigma = \sqrt{1.4333} = 1.1972\,\text{cookies/day}

*   *Quartiles*: Q1=0cookies/dayQ_1 = 0\,\text{cookies/day}, Q2=0.5cookies/dayQ_2 = 0.5\,\text{cookies/day}, Q3=1cookie/dayQ_3 = 1\,\text{cookie/day}.
  • Example 2: Kahoot Pin's Competition QCM:

    • Scenario: Kahoot competition where 1st prize is a pin's. A random selection of 1212 P1 students won respectively: 0,0,0,0,1,1,1,2,2,3,4,50, 0, 0, 0, 1, 1, 1, 2, 2, 3, 4, 5 pin's.

    • Item A: "The variable 'number of pin's won' is a qualitative variable." -> FALSE. It is a discrete quantitative variable because it represents countable whole numbers.

    • Item B: "The mean is 1.61.6." -> TRUE.

m=(0×4)+(1×3)+(2×2)+3+4+512=0+3+4+3+4+512=19121.58331.6pin’sm = \frac{(0 \times 4) + (1 \times 3) + (2 \times 2) + 3 + 4 + 5}{12} = \frac{0 + 3 + 4 + 3 + 4 + 5}{12} = \frac{19}{12} \approx 1.5833 \approx 1.6\,\text{pin's}

*   *Item C*: "The median is 11\,€." -> **FALSE**. The unit is incorrect (1pin’s1\,\text{pin's}, not euros). Numerical calculation: midpoint between 6th (11) and 7th (11) value is 1+12=1pin’s\frac{1+1}{2} = 1\,\text{pin's}.
*   *Item D*: "The mode is 11" -> **FALSE**. Value 00 occurs 44 times, whereas 11 occurs 33 times. The mode is 0pin’s0\,\text{pin's}.
*   *Item E*: "The variance is 2.582.58" -> **FALSE**. The variance magnitude is 2.582.58, but the required unit is pin’s2\text{pin's}^2. Detailed formula:

σ2=4(01.5833)2+3(11.5833)2+2(21.5833)2+(31.5833)2+(41.5833)2+(51.5833)212=2.58pin’s2\sigma^2 = \frac{4(0 - 1.5833)^2 + 3(1 - 1.5833)^2 + 2(2 - 1.5833)^2 + (3 - 1.5833)^2 + (4 - 1.5833)^2 + (5 - 1.5833)^2}{12} = 2.58\,\text{pin's}^2

Fundamentals of Probability Theory

  • Basic Terminology:

    • Trial / Random Experiment (Épreuve): An experiment with an uncertain, random outcome (e.g., rolling a die, drawing a random patient).

    • Elementary Event: A single outcome of a trial (e.g., rolling a specific number like 33).

    • Sample Space (EE): Set of all possible elementary events.

    • Event: A subset of the sample space EE corresponding to a defined condition (e.g., rolling an even number {2,4,6}\left\{2, 4, 6\right\}).

  • Fundamental Probability Rules:

    • The probability of any event AA satisfies 0P(A)10 \le P(A) \le 1

    • Complementary Event (Aˉ\bar{A}): Event occurring when AA does not occur. Rule:

P(Aˉ)=1P(A)P(\bar{A}) = 1 - P(A)

*   **Event Inclusion (ABA \subset B)**: If event AA is completely contained in event BB, then:

P(A)P(B)P(A) \le P(B)

  • Union and Intersection Operations:

    • Union (ABA \cup B): Event "AA or BB or both occur". General additive rule:

P(AB)=P(A)+P(B)P(AB)P(A \cup B) = P(A) + P(B) - P(A \cap B)

*   **Mutually Exclusive / Incompatible Events**: If AA and BB cannot occur simultaneously (AB=A \cap B = \emptyset), then P(AB)=0P(A \cap B) = 0, simplifying to:

P(AB)=P(A)+P(B)P(A \cup B) = P(A) + P(B)

*   **Intersection (ABA \cap B)**: Event "both AA and BB occur simultaneously". General multiplicative rule:

P(AB)=P(AB)×P(B)=P(BA)×P(A)P(A \cap B) = P(A|B) \times P(B) = P(B|A) \times P(A)

*   **Independent Events**: If the occurrence of BB does not alter the likelihood of AA, then P(AB)=P(A)P(A|B) = P(A), simplifying to:

P(AB)=P(A)×P(B)P(A \cap B) = P(A) \times P(B)

Conditional Probability and Bayes' Theorem

  • Conditional Probability Definition:

    • Denotes probability of event AA given that event BB has occurred (P(AB)P(A|B) or PB(A)P_B(A)). Expressed in natural language by phrases such as "given that" or "sachant que".

P(AB)=P(AB)P(B)P(A|B) = \frac{P(A \cap B)}{P(B)}

  • Symmetry of Intersection:

P(AB)=P(BA)    P(AB)×P(B)=P(BA)×P(A)P(A \cap B) = P(B \cap A) \implies P(A|B) \times P(B) = P(B|A) \times P(A)

  • Law of Total Probability:

    • For a partition formed by event AA and its complement Aˉ\bar{A}, any event BB can be decomposed into B=(BA)(BAˉ)B = (B \cap A) \cup (B \cap \bar{A}):

P(B)=P(BA)+P(BAˉ)=P(BA)×P(A)+P(BAˉ)×P(Aˉ)P(B) = P(B \cap A) + P(B \cap \bar{A}) = P(B|A) \times P(A) + P(B|\bar{A}) \times P(\bar{A})

  • Bayes' Theorem:

    • Allows updating conditional probabilities (inverting conditions):

P(AB)=P(BA)×P(A)P(B)=P(BA)×P(A)P(BA)×P(A)+P(BAˉ)×P(Aˉ)P(A|B) = \frac{P(B|A) \times P(A)}{P(B)} = \frac{P(B|A) \times P(A)}{P(B|A) \times P(A) + P(B|\bar{A}) \times P(\bar{A})}

  • Worked Example: Probability Tree Analysis:

    • Variables: MM = eating fruits & vegetables at least 1×/week1\times/\text{week}, FF = being in shape.

    • Given Tree Values:

      • P(M)=0.55    P(Mˉ)=0.45P(M) = 0.55 \implies P(\bar{M}) = 0.45

      • P(FM)=0.85    P(FˉM)=0.15P(F|M) = 0.85 \implies P(\bar{F}|M) = 0.15

      • P(FMˉ)=0.60    P(FˉMˉ)=0.40P(F|\bar{M}) = 0.60 \implies P(\bar{F}|\bar{M}) = 0.40

    • Question Calculations:

      • Probability of eating fruits 1×/week1\times/\text{week}: P(M)=0.55P(M) = 0.55 (55%55\%).

      • Probability of being in shape given not eating fruits 1×/week1\times/\text{week}: P(FMˉ)=0.60P(F|\bar{M}) = 0.60 (60%60\%).

      • Probability of being in shape AND eating fruits 1×/week1\times/\text{week}:

P(MF)=P(FM)×P(M)=0.85×0.55=0.46750.47(47%P(M \cap F) = P(F|M) \times P(M) = 0.85 \times 0.55 = 0.4675 \approx 0.47\,(47\%

    *   Total Probability of being in shape P(F)P(F): 

P(F)=P(MF)+P(MˉF)=(0.55×0.85)+(0.45×0.60)=0.4675+0.27=0.73750.74(74%P(F) = P(M \cap F) + P(\bar{M} \cap F) = (0.55 \times 0.85) + (0.45 \times 0.60) = 0.4675 + 0.27 = 0.7375 \approx 0.74\,(74\%

    *   Probability of eating fruits 1×/week1\times/\text{week} given being in shape P(MF)P(M|F): 

P(MF)=P(MF)P(F)=0.46750.73750.63390.63(63%P(M|F) = \frac{P(M \cap F)}{P(F)} = \frac{0.4675}{0.7375} \approx 0.6339 \approx 0.63\,(63\%

Random Variables and Standard Distributions

  • Discrete Random Variables:

    • Quantitative: Take discrete integer values.

    • Qualitative: Modalities assigned numerical markers (e.g., diseased = 11, healthy = 00).

    • Characterized by a discrete probability distribution P(X=xi)P(X = x_i), expectation/mean E[X]E[X], and variance Var(X)Var(X).

    • Common Discrete Laws: Bernoulli, Binomial, Poisson distributions.

  • Continuous Random Variables:

    • Take real number continuous values within an interval (e.g., biological parameters, drug dosage).

    • Evaluated via probability density functions (area under the curve represents probability) and cumulative distribution functions.

  • Standard Normal Distribution (Gauss Distribution, N(0,1)\mathcal{N}(0,1)):

    • Continuous bell-shaped curve, symmetrical around the mean μ=0\mu = 0, with standard deviation σ=1\sigma = 1.

    • For any standard normal variable, Mean = Median = Mode.

    • Total area under the density curve equals 11.

Standard Normal Distribution Symmetry and Table Usage

  • Notation and General Normal Variable Transformation:

    • If a continuous variable XX follows a normal distribution with mean μ\mu and standard deviation σ\sigma, it is written as XN(μ,σ)X \sim \mathcal{N}(\mu, \sigma).

  • Key Probability Area Identities (UαU_\alpha Tail Rules):

    • Complementary two-tailed area definition: The value UαU_\alpha defines central confidence area 1α1 - \alpha and two outer tail areas total α\alpha (α2\frac{\alpha}{2} in each tail).

    • P(U < -U_\alpha) = P(U > +U_\alpha) = \frac{\alpha}{2}

    • P(U > -U_\alpha) = P(U \le +U_\alpha) = 1 - \frac{\alpha}{2}

    • P(-U_\alpha < U < 0) = \frac{1 - \alpha}{2}

    • P(-U_\alpha < U < +U_\alpha) = 1 - \alpha

    • P(-U_{\alpha 1} < U < +U_{\alpha 2}) = 1 - \frac{\alpha_1}{2} - \frac{\alpha_2}{2}

    • P(+U_{\alpha 1} < U < +U_{\alpha 2}) = \frac{\alpha_1 - \alpha_2}{2}

  • Reduced Normal Table (UαU_\alpha vs α\alpha mapping):

α=0.01    Uα=2.576\alpha = 0.01 \implies U_\alpha = 2.576

α=0.02    Uα=2.326\alpha = 0.02 \implies U_\alpha = 2.326

α=0.05    Uα=1.960\alpha = 0.05 \implies U_\alpha = 1.960

α=0.10    Uα=1.645\alpha = 0.10 \implies U_\alpha = 1.645

α=0.20    Uα=1.282\alpha = 0.20 \implies U_\alpha = 1.282

α=0.30    Uα=1.036\alpha = 0.30 \implies U_\alpha = 1.036

α=0.40    Uα=0.842\alpha = 0.40 \implies U_\alpha = 0.842

α=0.50    Uα=0.674\alpha = 0.50 \implies U_\alpha = 0.674

α=0.60    Uα=0.524\alpha = 0.60 \implies U_\alpha = 0.524

α=0.70    Uα=0.385\alpha = 0.70 \implies U_\alpha = 0.385

α=0.80    Uα=0.253\alpha = 0.80 \implies U_\alpha = 0.253

α=0.90    Uα=0.126\alpha = 0.90 \implies U_\alpha = 0.126

  • Worked Numerical Table Example:

    • Find P(U < 0.45). Given Uα=0.45U_\alpha = 0.45, looking up table yields corresponding two-tailed risk α=0.65\alpha = 0.65

    • Applying formula P(U < 0.45) = 1 - \frac{\alpha}{2} = 1 - \frac{0.65}{2} = 1 - 0.325 = 0.675

  • Other Related Statistical Distributions:

    • Chi-squared (χ2\chi^2) distribution.

    • Student's t-distribution.

    • Fisher's F-distribution.

Sampling Fluctuation, Estimation, and Confidence Intervals

  • Notation Summary (Population vs. Sample):

    • Population Parameters (Unknown Truth): Mean μ\mu, Variance σ2\sigma^2, Standard Deviation σ\sigma, Proportion PP.

    • Sample Statistics (Observed Estimates): Mean mm, Variance s2s^2, Standard Deviation ss, Proportion ff.

  • Sampling Fluctuation Principles:

    • Drawing different random samples from the same population leads to varying estimates (E1,E2,,EiE_1, E_2, \dots, E_i) due to random chance.

    • Representativeness relies on sufficient sample size and random drawing.

  • Confidence Intervals (IC):

    • Constructs a value range expected to contain the true unknown population parameter with probability 1α1 - \alpha (typically 95%95\% when α=0.05\alpha = 0.05).

    • For a normal distribution, 95%95\% of individual observations fall within [μ1.96σ,μ+1.96σ][\mu - 1.96\sigma, \mu + 1.96\sigma] (frequently approximated as μ±2σ\mu \pm 2\sigma).

  • Conditions for Interval Construction:

    • Quantitative Variables (A Priori Condition): Sample size must satisfy N30N \ge 30

    • Qualitative Variables (A Posteriori Conditions): Expected success/failure counts must satisfy N×P5N \times P \ge 5 and N(1P)5N(1 - P) \ge 5 (or using sample estimate N×f5N \times f \ge 5 and N(1f)5N(1 - f) \ge 5).

  • Factors Governing Estimation Precision:

    • The margin of error block (Uα×Standard ErrorU_\alpha \times \text{Standard Error}) determines the interval precision.

    • Larger sample size (NN)     \implies smaller precision block     \implies narrower interval     \implies higher estimation precision.

    • Smaller interval length     \implies higher estimation precision.

    • Higher alpha risk (α\alpha)     \implies smaller UαU_\alpha threshold     \implies narrower numerical interval width     \implies increased precision of interval bounds (at the cost of increased risk of excluding the true parameter).

Methodology of Statistical Hypothesis Testing

  • Primary Objective:

    • To make objective mathematical decisions about population characteristics based on sample evidence, while rigorously controlling decision error risk.

    • Example application question: Is there an association between systolic blood pressure and body mass index in adult men?

  • The 6 Standard Steps of a Statistical Test:

    1. Formulate Hypotheses:

      • Null Hypothesis (H0H_0): Hypothesis of no effect, no difference, or equality between populations (e.g., μ1=μ2\mu_1 = \mu_2).

      • Alternative Hypothesis (H1H_1): New hypothesis adopted if H0H_0 is rejected; asserts a real difference/effect exists.

    2. Select Significance Level (α\alpha):

      • Usually set at α=5%\alpha = 5\% (0.050.05). Sets the maximum acceptable threshold for committing a Type I error.

    3. Define Test Parameter / Statistic:

      • Identify appropriate probability distribution based on variable type (qualitative/quantitative), sample count, sample size, and mathematical conditions.

    4. Determine Critical Region (RC):

      • Set boundary values defining extreme results that have only α\le \alpha probability of occurring if H0H_0 were true (using distribution tables such as standard normal TER).

    5. Calculate Test Parameter:

      • Compute test statistic from observed sample data.

    6. Statistical Decision:

      • If Calculated Parameter \in Critical Region: Reject H0H_0 at risk α=5%\le \alpha = 5\%. The observed difference is unlikely under H0H_0 (pαp \le \alpha). There is a statistically significant difference in the population.

      • If Calculated Parameter \notin Critical Region: Fail to reject H0H_0 at risk α=5%\alpha = 5\%. The observed difference is plausible under H0H_0 (p > \alpha). There is no statistically significant difference; observed differences are attributed to random sampling fluctuations.

  • Decision Errors and $p$-Value Concept:

    • Type I Error (Alpha Risk, α\alpha): Rejection of null hypothesis H0H_0 when H0H_0 is actually true.

    • Type II Error (Beta Risk, β\beta): Acceptance / failure to reject null hypothesis H0H_0 when H0H_0 is actually false.

      • β\beta is inversely related to sample size: smaller samples lead to larger β\beta risks.

    • pp-Value (pp): Exact probability of obtaining a test statistic at least as extreme as the observed value, assuming H0H_0 is true.

      • If pαp \le \alpha (e.g., p0.05p \le 0.05): Reject H0H_0 with risk pαp \le \alpha

      • If p > \alpha (e.g., p=0.30p = 0.30): Cannot reject H0H_0 because doing so carries a 30%30\% risk of committing a Type I error.

  • Sample Representativeness Testing Application:

    • Null hypothesis formulated as H0H_0: "The study sample is representative of the population for the disease."

    • If test fails to reject H0H_0, the sample is concluded to be representative of the target population.

Distribution Selection for Statistical Tests and Sample Size Principles

  • Framework for Probability Distribution Selection:

    • Selection depends on: variable type, sample sizes, and underlying approximation conditions.

  • Large Samples (N30N \ge 30):

    • Central Limit Theorem Property: Sums or averages of large numbers of independent random variables of any initial distribution follow an approximately Normal Distribution.

    • Testing Means: Use Normal Distribution (N30N \ge 30).

    • Testing Proportions: Use Normal Distribution (N30N \ge 30 and N×P5N \times P \ge 5, N(1P)5N(1 - P) \ge 5).

    • Testing Independence of 2 Qualitative Variables: Use Chi-Squared Distribution (χ2\chi^2) provided expected cell theoretical frequencies satisfy Theoretical Effectifs5\text{Theoretical Effectifs} \ge 5

  • Small Samples:

    • Standard normal approximations fail.

    • When theoretical underlying distribution laws are unknown or conditions fail, resort to Non-parametric tests.