Class Notes on Variances, Standard Deviations, and Measures of Variation

Class Schedule Adjustments

  • The instructor has moved the class schedule forward due to snow day delays.

  • Homework Assignment #3 is now due on the 12th.

  • Class will review material on Thursday instead of holding an exam.

  • Communication about adjustments has been sent via email to all students.

Attendance Roll Call

  • Students present: Owen, Brent, Destiny, Jacob, Eliana, Katie, Ronnie, Avon, Payton, Penn.

  • Noted students are now back home due to various activities.

Hallmarks Information

  • Hallmarks #1 and #2 are due today by midnight.

  • The Pearson platform will close submissions at midnight.

Measures of Variation - Overview

  • Section 32 covers measures of variation.

Range

  • Definition: The range is a basic measure of variation calculated by subtracting the minimum value from the maximum value in a dataset.

  • Formula: extRange=extMaxvalue−extMinvalueext{Range} = ext{Max value} - ext{Min value}

  • Characteristics:

    • Reflects only the extremes (maximum and minimum values).

    • Not resistant to outliers (a single extreme value can dramatically affect the range).

  • Example Calculation:

    • For a data set: 20, …, 75

    • Max = 75, Min = 20

    • extRange=75−20=55ext{Range} = 75 - 20 = 55

    • The dataset has 11 values, suggesting some variance.

Standard Deviation

  • Definition: A more comprehensive measure of variation that indicates how much individual data values deviate from the mean.

  • Notation:

    • Sample Standard Deviation: denoted by lowercase 's'

    • Population Standard Deviation: denoted by lowercase 'σ' (sigma).

  • Importance: Indicates the spread of values in relation to the mean; essential for understanding the dispersion in a dataset.

  • Mean: The central value around which the dataset is analyzed.

Calculation of Standard Deviation
  1. Calculate the Mean: xˉ=racextSumofallxn\bar{x} = rac{ ext{Sum of all } x}{n}

  2. Determining Deviations: For each data value, subtract the mean and square the result to eliminate negatives:

    • Example:

      • If mean height is 66 inches, an individual height of 74 inches yields:

      • Deviations: 74 - 66 = 8

      • Squared deviation: 82=648^2 = 64

  3. Sum of Squared Deviations

  4. Divide by n-1 (for sample variance):

    • Formula for Sample Variance: s2=racextSumofSquaredDeviationsn−1s^2 = rac{ ext{Sum of Squared Deviations}}{n - 1}

  5. Square Root:

    • s=extsqrt(s2)s = ext{sqrt}(s^2)

  • Characteristics:

    • Mean of deviations can be zero; thus squaring removes this.

    • Standard deviation gives us a numerical value indicating the average distance of each data point from the mean.

    • Units are the same as the data set's original units.

Differences between Sample and Population Statistics

  • Sample standard deviation tends to underestimate the population standard deviation.

  • Using n−1n - 1 in samples compensates for the bias.

  • Population standard deviation uses just nn because all data points are included and no bias from sample selection occurs.

Variance

  • Definition: Variance is the square of the standard deviation.

  • Sample Variance: (s2)(s^2)

  • Population Variance: (σ2)(\sigma^2)

  • Units of variance are squared units of the original dataset (e.g., if data in inches, variance in square inches).

Outliers and Variation

  • Outliers can dramatically affect both range and standard deviation.

  • Comparison between sample variance and population variance can guide estimations for datasets lacking full information.

Empirical Rule and Chebyshev's Theorem

Empirical Rule (Normal Distribution)

  • States that:

    • Approximately 68% of the data falls within 1 standard deviation of the mean.

    • Approximately 95% of the data falls within 2 standard deviations.

    • Approximately 99.7% of the data falls within 3 standard deviations.

Chebyshev's Theorem (Non-Normal Distributions)

  • For any dataset, at least:

    • (1−rac1k2)(1 - rac{1}{k^2}) of the data is contained within k standard deviations from the mean.

    • For example:

    • For k=2: (1−rac14)=0.75(1 - rac{1}{4}) = 0.75 (at least 75% within 2 standard deviations)

    • For k=3: (1−rac19)=0.89(1 - rac{1}{9}) = 0.89 (at least 89% within 3 standard deviations)

Coefficient of Variation (CV)

  • Formula: CV=racsxˉimes100CV = rac{s}{\bar{x}} imes 100, where s is the standard deviation and xˉ\bar{x} is the mean.

  • Purpose: It allows for the comparison of the degree of variation from one dataset to another, expressed as a percentage.

  • A CV greater than 1% indicates significant differences between sample variations.

Practical Example - Analysis of Celebrity Net Worths

  • Tasks involved:

    • Find the range, variance, and standard deviation for a dataset of celebrity net worths (in billions).

    • Understand limitations of findings (such as typicality) based on the exclusive nature of the sample data.

  • Round off considerations: Round results to the appropriate units based on the nature of the data.

Blood Platelet Count Example Using Empirical Rule

  • Given:

    • Mean = 255.3

    • Standard Deviation = 65.9

  • Using the empirical rule:

    • 95% of women will have platelet counts between:

    • 123.5 (2 SDs below) and 387.1 (2 SDs above).

  • Applying Chebyshev’s Theorem for different non-normal distributions.

Summary

  • Today's focus was on measures of variation, calculation of range, standard deviation, variance, and understanding practical applications and implications of these statistics within various contexts.

  • Next class will inquire about measures of relative standing and introduction to box plots.