Normal Distribution
Normal Distribution
Introduction to Normal Distribution
This module covers the normal distribution, a fundamental concept in statistics, including how transformations affect its parameters, properties of its density curve, the empirical rule, z-scores, and the use of the standard normal table for probability calculations, as well as graphical assessment using Q-Q plots.
Shifting vs. Scaling
These are transformations applied to data that affect its statistical properties:
Shifting (Center and Position): Adding or subtracting a constant value to each data point.
Effect on Mean: The mean changes by the amount of the constant added or subtracted.
Effect on Standard Deviation: The standard deviation (and thus spread) remains unchanged.
Example (Bonus Marks): If is the original mark and a bonus of % (interpreted as percentage points) is given, making the new mark . The new mean score will be , while the standard deviation will remain the same, .
Scaling (Spread and Shape): Multiplying or dividing each data point by a constant value.
Effect on Mean: The mean is multiplied or divided by the same constant.
Effect on Standard Deviation: The standard deviation is also multiplied or divided by the absolute value of the same constant.
Effect on Variance: The variance is multiplied or divided by the square of the constant.
Formulas provided for standard deviation and variance when scaling by a factor :
New standard deviation:
New variance:
Example (Scaling Marks): If is the original mark and each mark is scaled up by %, meaning . The new mean score will be , and the new standard deviation will be . For example, if a student got %, their new mark would be %.
Density Curve
A density curve is a graphical representation of the distribution of a continuous variable. It must satisfy specific conditions:
Non-negativity: A density curve is always on or above the horizontal axis (i.e., its values are never negative, ).
Total Area: The total area between the horizontal axis and under the density curve must equal (or %), representing the total probability or proportion of observations.
Probability as Area: The probability of an observation falling within an interval , denoted as P(a < x < b), is equal to the area under the curve between and . This also represents the percentage of observations that will fall in that interval.
Continuous Variables Properties: For continuous variables, certain probability statements hold:
The probability of a single exact value is zero: .
The inclusion or exclusion of endpoints does not change the probability for an interval: P(a < x < b) = P(a \le x \le b) = P(a < x \le b) = P(a \le x < b).
Example (Density Curve from Page 5):
Verification of Area = 1 (by geometry): Assuming a uniform distribution from to with a height (density) of . The area is .
Probability that is less than : P(X < 1) is the area from to . .
Probability that is greater than : P(X > 1.5) is the area from to . .
Empirical Rule (68-95-99.7% Rule)
This rule applies specifically to data that follows a normal distribution, describing the approximate percentage of observations that fall within a certain number of standard deviations from the mean:
Approximately % of observations fall within standard deviation () of the mean ().
Approximately % of observations fall within standard deviations () of the mean ().
Approximately % of observations fall within standard deviations () of the mean ().
Example (Standardized Exam Completion Time):
Given: Mean ( minutes), Standard Deviation ( minutes).
Percentage of students completing the exam in under an hour ( minutes):
minutes is one standard deviation below the mean ().
According to the empirical rule, % of students complete between and minutes ().
The remaining are outside this range, split equally below and above . So, complete in under minutes.
Percentage of students completing the exam between and minutes:
This interval is from one standard deviation below the mean to the mean ( to ).
This represents exactly half of the % interval: .
Time interval for the central % of students:
The central % of observations fall within standard deviations of the mean ().
The interval is .
Therefore, the central % of students would be found between and minutes.
Standardized Score / z-score
The z-score is a measure of relative standing that indicates how many standard deviations an observation is away from the mean.
Relative Standing: It positions a data point relative to the rest of the distribution.
No Units: Z-scores are unitless, allowing comparison across different distributions.
Interpretation: A z-score tells