Comprehensive Study Notes on Descriptive Statistics, Probability Theory, and Hypothesis Testing
Descriptive Epidemiology and Population Concepts
Descriptive Epidemiology Objectives:
Allows the description of specific characteristics within a given population, such as health-related problems.
Serves to estimate characteristics in a larger general population based on observed data from a study sample.
Crucial Limitation: Direct conclusion/extrapolation to the overall population is not strictly possible unless specific sampling conditions are met. Study results apply directly to the study sample, and researchers must be vigilant regarding potential biases.
Target Population vs. Study Sample:
Population (): Large-scale total group of individuals for which specific characteristics are sought to be studied. The basic element is the individual subject.
Sample (): Restricted subgroup belonging to the larger target population upon which measurements and studies are actually conducted to simplify feasibility.
Variability and Defining Parameters:
Inter-subject Variability: Represents differing responses or measurements among individual subjects to an identical prompt or parameter (e.g., individual season preferences: summer vs. winter).
Essential Context Boundaries: Any population characteristic or study variable must strictly be defined according to three mandatory dimensions:
Time (e.g., year 2025)
Place (e.g., Pau)
Person (e.g., P1 health science students)
Conditions for General Population Estimation:
To estimate a characteristic in a broader population from a sample, the sample must be sufficiently representative of that population.
Representativeness requires:
A sufficiently large sample size ().
Random sampling selection (drawing subjects at random).
Variable Types and Data Presentation
Classification of Variables:
Quantitative Variables (Numerical):
Discrete Quantitative Variables: Possess a finite, countable number of distinct values. Values are strictly whole/round numbers (e.g., number of P1 students = , never ; number of pin's won).
Continuous Quantitative Variables: Possess an infinite number of possible real values within a given range (e.g., age, height, weight, biological parameters, blood cholesterol, active ingredient quantity in medicine).
Qualitative Variables (Categorical):
Represent non-numerical categories or modalities (e.g., sex, hair color, profession, diabetes status [yes/no, Type 1/Type 2]).
For analytical convenience, numerical codes can be assigned to qualitative modalities (e.g., coding disease status as for diseased and for non-diseased).
Data Summarization Techniques:
Raw Data Series: Unstructured lists of values (e.g., ) provide no synthetic overview and prevent meaningful conclusions.
Summarized Information: Databases aggregate data into synthetic numerical parameters (central tendency and dispersion) or graphical representations to expose variable distributions.
Graphical Representation Methods:
Bar Chart (Diagramme en barres): Standard for qualitative/categorical variables, as well as discrete quantitative variables.
Histogram: Designed for continuous quantitative variables, plotting frequency or percentage against continuous class intervals.
Descriptive Statistics for Quantitative Variables
Measures of Central Tendency (Position):
Must always be expressed with their proper measurement units.
Mean ( or ): Arithmetic average of all observed values, divided by total observation count :
* **Median ()**: Value that divides an ordered sequence of numbers into two equal halves (50% of values are greater, 50% are lower).
* If dataset size is **odd**: Median position is .
* If dataset size is **even**: Median is the mean of the values at positions and .
* **Mode**: Most frequently occurring value in the dataset. A distribution may have no mode, a single unique mode, or multiple modes.
Distribution Shape and Central Tendency Relationships:
Symmetrical Distribution (Gaussian / Normal): Mode = Median = Mean.
Asymmetrical Distribution (Skewed): Mode, Median, and Mean take distinct separate values.
Comparative Evaluation of Central Tendency Indicators:
Mean ():
Advantages: Easy to calculate; widely recognized and universally understood.
Disadvantages: Highly sensitive to extreme values/outliers; poorly represents heterogeneous or bimodal populations.
Median ():
Advantages: Unaffected by extreme values; robust indicator for asymmetrical distributions.
Disadvantages: Ignores total data distribution details (extreme ranges); manually tedious to determine when is large.
Mode:
Advantages: Effectively represents heterogeneous or multimodal populations.
Disadvantages: Manually tedious to identify for large ; fluctuates based on grouping class widths.
Measures of Dispersion (Variation):
Informs on the spread or variation of values around the central tendency. Must always retain measurement units.
Range (): Simple difference between maximum and minimum observed values:
* **Variance ( or )**: Mean of squared deviations from the arithmetic mean. Quantifies observation spread around the mean. Unit is **squared** ().
* Population Variance:
* Sample Variance:
* **Standard Deviation ( or )**: Square root of the variance. Restores the original measurement unit.
* Population Standard Deviation:
* Sample Standard Deviation:
* **Percentiles (Centiles)**: The \-th percentile is the value below which of observations lie.
* **Quartiles**: Split data into 4 equal parts (, [Median], ).
* **Terciles**: Split data into 3 equal parts (, ).
* **Deciles**: Split data into 10 equal parts (, , ..., ).
Summary Recommendation for Parameter Reporting:
For Symmetrical Distributions: Report Mean + Standard Deviation.
For Asymmetrical Distributions: Report Median + Percentiles (e.g., Interquartile Range).
Qualitative Variable Description:
Frequency (): Proportion of individuals in a specific category relative to total sample size :
* Where is the category effectif (count) and is the total population count. The sum of all category frequencies strictly equals .
Detailed Practical Examples and QCM Applications
Example 1: Cookie Consumption Analysis:
Scenario: P1 students discuss daily cookie consumption. eat cookies/day, eat cookie/day, and eat cookies/day.
Ordered Dataset: ().
Mean Calculation:
* *Median Calculation*: (even). Midpoint between 5th value () and 6th value ():
* *Mode*: (occurs times).
* *Distribution Shape*: Asymmetrical (Mode = , Median = , Mean = ).
* *Range*:
* *Variance Calculation*:
* *Standard Deviation Calculation*:
* *Quartiles*: , , .
Example 2: Kahoot Pin's Competition QCM:
Scenario: Kahoot competition where 1st prize is a pin's. A random selection of P1 students won respectively: pin's.
Item A: "The variable 'number of pin's won' is a qualitative variable." -> FALSE. It is a discrete quantitative variable because it represents countable whole numbers.
Item B: "The mean is ." -> TRUE.
* *Item C*: "The median is ." -> **FALSE**. The unit is incorrect (, not euros). Numerical calculation: midpoint between 6th () and 7th () value is .
* *Item D*: "The mode is " -> **FALSE**. Value occurs times, whereas occurs times. The mode is .
* *Item E*: "The variance is " -> **FALSE**. The variance magnitude is , but the required unit is . Detailed formula:
Fundamentals of Probability Theory
Basic Terminology:
Trial / Random Experiment (Épreuve): An experiment with an uncertain, random outcome (e.g., rolling a die, drawing a random patient).
Elementary Event: A single outcome of a trial (e.g., rolling a specific number like ).
Sample Space (): Set of all possible elementary events.
Event: A subset of the sample space corresponding to a defined condition (e.g., rolling an even number ).
Fundamental Probability Rules:
The probability of any event satisfies
Complementary Event (): Event occurring when does not occur. Rule:
* **Event Inclusion ()**: If event is completely contained in event , then:
Union and Intersection Operations:
Union (): Event " or or both occur". General additive rule:
* **Mutually Exclusive / Incompatible Events**: If and cannot occur simultaneously (), then , simplifying to:
* **Intersection ()**: Event "both and occur simultaneously". General multiplicative rule:
* **Independent Events**: If the occurrence of does not alter the likelihood of , then , simplifying to:
Conditional Probability and Bayes' Theorem
Conditional Probability Definition:
Denotes probability of event given that event has occurred ( or ). Expressed in natural language by phrases such as "given that" or "sachant que".
Symmetry of Intersection:
Law of Total Probability:
For a partition formed by event and its complement , any event can be decomposed into :
Bayes' Theorem:
Allows updating conditional probabilities (inverting conditions):
Worked Example: Probability Tree Analysis:
Variables: = eating fruits & vegetables at least , = being in shape.
Given Tree Values:
Question Calculations:
Probability of eating fruits : ().
Probability of being in shape given not eating fruits : ().
Probability of being in shape AND eating fruits :
* Total Probability of being in shape :
* Probability of eating fruits given being in shape :
Random Variables and Standard Distributions
Discrete Random Variables:
Quantitative: Take discrete integer values.
Qualitative: Modalities assigned numerical markers (e.g., diseased = , healthy = ).
Characterized by a discrete probability distribution , expectation/mean , and variance .
Common Discrete Laws: Bernoulli, Binomial, Poisson distributions.
Continuous Random Variables:
Take real number continuous values within an interval (e.g., biological parameters, drug dosage).
Evaluated via probability density functions (area under the curve represents probability) and cumulative distribution functions.
Standard Normal Distribution (Gauss Distribution, ):
Continuous bell-shaped curve, symmetrical around the mean , with standard deviation .
For any standard normal variable, Mean = Median = Mode.
Total area under the density curve equals .
Standard Normal Distribution Symmetry and Table Usage
Notation and General Normal Variable Transformation:
If a continuous variable follows a normal distribution with mean and standard deviation , it is written as .
Key Probability Area Identities ( Tail Rules):
Complementary two-tailed area definition: The value defines central confidence area and two outer tail areas total ( in each tail).
P(U < -U_\alpha) = P(U > +U_\alpha) = \frac{\alpha}{2}
P(U > -U_\alpha) = P(U \le +U_\alpha) = 1 - \frac{\alpha}{2}
P(-U_\alpha < U < 0) = \frac{1 - \alpha}{2}
P(-U_\alpha < U < +U_\alpha) = 1 - \alpha
P(-U_{\alpha 1} < U < +U_{\alpha 2}) = 1 - \frac{\alpha_1}{2} - \frac{\alpha_2}{2}
P(+U_{\alpha 1} < U < +U_{\alpha 2}) = \frac{\alpha_1 - \alpha_2}{2}
Reduced Normal Table ( vs mapping):
Worked Numerical Table Example:
Find P(U < 0.45). Given , looking up table yields corresponding two-tailed risk
Applying formula P(U < 0.45) = 1 - \frac{\alpha}{2} = 1 - \frac{0.65}{2} = 1 - 0.325 = 0.675
Other Related Statistical Distributions:
Chi-squared () distribution.
Student's t-distribution.
Fisher's F-distribution.
Sampling Fluctuation, Estimation, and Confidence Intervals
Notation Summary (Population vs. Sample):
Population Parameters (Unknown Truth): Mean , Variance , Standard Deviation , Proportion .
Sample Statistics (Observed Estimates): Mean , Variance , Standard Deviation , Proportion .
Sampling Fluctuation Principles:
Drawing different random samples from the same population leads to varying estimates () due to random chance.
Representativeness relies on sufficient sample size and random drawing.
Confidence Intervals (IC):
Constructs a value range expected to contain the true unknown population parameter with probability (typically when ).
For a normal distribution, of individual observations fall within (frequently approximated as ).
Conditions for Interval Construction:
Quantitative Variables (A Priori Condition): Sample size must satisfy
Qualitative Variables (A Posteriori Conditions): Expected success/failure counts must satisfy and (or using sample estimate and ).
Factors Governing Estimation Precision:
The margin of error block () determines the interval precision.
Larger sample size () smaller precision block narrower interval higher estimation precision.
Smaller interval length higher estimation precision.
Higher alpha risk () smaller threshold narrower numerical interval width increased precision of interval bounds (at the cost of increased risk of excluding the true parameter).
Methodology of Statistical Hypothesis Testing
Primary Objective:
To make objective mathematical decisions about population characteristics based on sample evidence, while rigorously controlling decision error risk.
Example application question: Is there an association between systolic blood pressure and body mass index in adult men?
The 6 Standard Steps of a Statistical Test:
Formulate Hypotheses:
Null Hypothesis (): Hypothesis of no effect, no difference, or equality between populations (e.g., ).
Alternative Hypothesis (): New hypothesis adopted if is rejected; asserts a real difference/effect exists.
Select Significance Level ():
Usually set at (). Sets the maximum acceptable threshold for committing a Type I error.
Define Test Parameter / Statistic:
Identify appropriate probability distribution based on variable type (qualitative/quantitative), sample count, sample size, and mathematical conditions.
Determine Critical Region (RC):
Set boundary values defining extreme results that have only probability of occurring if were true (using distribution tables such as standard normal TER).
Calculate Test Parameter:
Compute test statistic from observed sample data.
Statistical Decision:
If Calculated Parameter Critical Region: Reject at risk . The observed difference is unlikely under (). There is a statistically significant difference in the population.
If Calculated Parameter Critical Region: Fail to reject at risk . The observed difference is plausible under (p > \alpha). There is no statistically significant difference; observed differences are attributed to random sampling fluctuations.
Decision Errors and $p$-Value Concept:
Type I Error (Alpha Risk, ): Rejection of null hypothesis when is actually true.
Type II Error (Beta Risk, ): Acceptance / failure to reject null hypothesis when is actually false.
is inversely related to sample size: smaller samples lead to larger risks.
-Value (): Exact probability of obtaining a test statistic at least as extreme as the observed value, assuming is true.
If (e.g., ): Reject with risk
If p > \alpha (e.g., ): Cannot reject because doing so carries a risk of committing a Type I error.
Sample Representativeness Testing Application:
Null hypothesis formulated as : "The study sample is representative of the population for the disease."
If test fails to reject , the sample is concluded to be representative of the target population.
Distribution Selection for Statistical Tests and Sample Size Principles
Framework for Probability Distribution Selection:
Selection depends on: variable type, sample sizes, and underlying approximation conditions.
Large Samples ():
Central Limit Theorem Property: Sums or averages of large numbers of independent random variables of any initial distribution follow an approximately Normal Distribution.
Testing Means: Use Normal Distribution ().
Testing Proportions: Use Normal Distribution ( and , ).
Testing Independence of 2 Qualitative Variables: Use Chi-Squared Distribution () provided expected cell theoretical frequencies satisfy
Small Samples:
Standard normal approximations fail.
When theoretical underlying distribution laws are unknown or conditions fail, resort to Non-parametric tests.