Comprehensive Study Notes: Skewed Distributions, Quartiles, and Summary Statistics

Measures of Center and Spread for Symmetric and Skewed Distributions

  • Symmetric Distributions:

    • Center: Described using the mean.

    • Spread: Described using the standard deviation.

    • Data Distribution: The empirical rule is applied to determine where the majority of the dataset exists.

  • Skewed Distributions:

    • Center: Described using the median.

    • Spread: Described using the Interquartile Range (IQR\text{IQR}).

    • Data Distribution: The empirical rule, mean, and standard deviation are not appropriate or useful for skewed distributions or skewed histograms.

  • Determining Distribution Shape by Comparing Mean and Median:

    • Rather than relying strictly on visual estimation of histograms, compare the numerical values of the mean and median.

    • Perfectly Symmetric: The mean is exactly equal to the median (Mean=Median\text{Mean} = \text{Median}).

    • Basically Symmetric: The mean is almost equal to the median (Mean≈Median\text{Mean} \approx \text{Median}).

    • Skewed: The mean is several units away from the median (Mean≠Median\text{Mean} \neq \text{Median}).

    • Skewed Left: The mean is significantly smaller than the median (Mean<Median\text{Mean} < \text{Median}).

    • Skewed Right: The mean is significantly larger than the median (Mean>Median\text{Mean} > \text{Median}).

  • Numerical Example of Skewness:

    • A histogram dataset has a mean of 250250 and a median of 226226.

    • Because 250≠226250 \neq 226, the distribution is not symmetric.

    • Because the mean (250250) is larger than the median (226226), the distribution is skewed right, meaning data points are more spread out to the right of the tallest peak than to the left.

Questions & Discussion: Z-Scores for US Birth Weights

  • Question: When evaluating a birth weight measurement relative to all US births within a symmetric distribution, which mean and standard deviation parameters should be selected?

  • Response:

    • Mean: Use the overall US birth mean of 35623562\,g\text{g} (or relative unit), rather than the mean specifically for infants born one month early.

    • Standard Deviation: Use the population standard deviation of 500500.

    • Z-Score Application: The z-score formula calculates how many standard deviations a specific data point is away from the mean, allowing one to classify the observation as usual or unusual.

Calculating the Median and Summary Statistics

  • Procedure for Finding the Median by Hand:

    1. Arrange the data values in order from smallest to largest.

    2. Identify the direct middle of the ordered list.

    3. If the dataset contains an odd number of observations (nn), the median is the single value located exactly in the center at position n+12\frac{n+1}{2}.

    4. If the dataset contains an even number of observations (nn), the median is the arithmetic average of the two middle numbers.

  • Resistance Property:

    • The median is resistant to extreme values and outliers, whereas the mean is sensitive to outliers.

  • Class A Dataset Analysis:

    • Dataset size: n=9n = 9 data points.

    • Median: The 5th5\text{th} ordered data point, which equals 7777.

    • Mean: Approximately 7979.

    • Comparison: The mean (≈79\approx 79) and median (7777) are not equal, indicating the distribution is skewed.

  • Class B Dataset Analysis:

    • Dataset size: n=10n = 10 data points.

    • Median: The average of the two middle observations (5th5\text{th} and 6th6\text{th} values), which equals 83.583.5

    • Mean: 81.481.4.

    • Comparison: The mean (81.481.4) and median (83.583.5) are not equal, indicating the distribution is skewed.

  • StatCrunch Software Workflow for Summary Statistics:

    1. Open a blank spreadsheet or table.

    2. Enter raw data into labeled columns (e.g., Class A, Class B). Data does not need to be sorted manually; the software sorts automatically.

    3. Navigate to Summary Stats →\rightarrow Columns.

    4. Hold the Ctrl key on the keyboard to select multiple columns simultaneously.

    5. Click Compute to generate summary statistics (mean, median, range, quartiles).

Assessing Skewness and Range in Datasets

  • Evaluating Degree of Skewness:

    • When comparing two distributions, evaluate which is more dramatically skewed or whether one is sufficiently close to symmetric to permit the use of symmetric rules.

  • New York Income Case Example:

    • Histograms representing income in New York (scaled in thousands of dollars per year) show dramatic right-skewness.

    • The mean and median lie in significantly different positions on the histogram.

    • Because the mean and median are vastly different, the visual skewness is confirmed numerically. The median must be used to represent a typical income value.

  • The Range:

    • Definition: The total distance spanned by the dataset.

    • Formula:         Range=Maximum Value−Minimum Value\text{Range} = \text{Maximum Value} - \text{Minimum Value}

    • Class A Range Calculation: For Class A data with a maximum value of 9898 and a minimum value of 5656:         Range=98−56=42\text{Range} = 98 - 56 = 42

    • Properties: Range defines a window 4242 units wide containing all data points. Range does not indicate whether a distribution is symmetric or skewed and is severely affected by outliers.

Quartiles and Interquartile Range (IQR)

  • Quartiles:

    • Quartiles partition an ordered dataset into four equal segments, each containing 25%25\% of the data.

    • Median (Q2Q_2): Splits the total dataset in half (50%50\% below, 50%50\% above).

    • First Quartile (Q1Q_1): The median of the lower half of the dataset (left side of the overall median).

    • Third Quartile (Q3Q_3): The median of the upper half of the dataset (right side of the overall median).

  • Manual Quartile Step-by-Step Calculation (Class A Example):

    • Class A has n=9n = 9 data values with overall median 7777.

    • Finding Q1Q_1:

      • Look at the 44 data points to the left of 7777.

      • Average the middle two values of these 44 observations to get Q_1 = 71$.\n * 71splitsthelowerhalf:splits the lower half:2datapointssitbelowdata points sit below71,and, and2datapointssitbetweendata points sit between71andand77$.

    • Finding Q3Q_3:

      • Look at the 44 data points to the right of 7777 (values: 8181, 8787, 9595, 9898).

      • Average the middle two values (8787 and 9595) to get Q_3 = 91$.\n * 91splitstheupperhalf:splits the upper half:2datapointssitbetweendata points sit between77andand91,and, and2datapointssitabovedata points sit above91$.

    • Central Spread: Between Q1=71Q_1 = 71 and Q3=91Q_3 = 91 lies approximately 50%50\% (or slightly more for odd nn) of the dataset. Values falling between Q1Q_1 and Q3Q_3 represent usual or typical values.

  • Class B Quartiles and IQR:

    • Q1=75Q_1 = 75

    • Q3=86Q_3 = 86

    • Interquartile Range Formula:         IQR=Q3−Q1\text{IQR} = Q_3 - Q_1

    • Class B IQR Calculation:         IQR=86−75=11\text{IQR} = 86 - 75 = 11

  • Generating IQR in StatCrunch:

    1. Navigate to Summary Stats →\rightarrow Columns.

    2. Select the target column.

    3. Under the Statistics selection box, scroll down and click IQR.

    4. Hold the Ctrl key while clicking to keep default summary statistics selected alongside IQR.

    5. Click Compute.

Comparing Center Spread: Empirical Rule versus IQR

  • Graphical Representation of Quartile Spacing:

    • Symmetric Distributions: Standard deviation intervals are evenly spaced on both sides of the mean.

    • Skewed Distributions: Distance between quartiles is uneven.

      • For example, the distance from the median to Q1Q_1 may be significantly narrower than the distance from the median to Q3Q_3.

    • Quartile Proportions: Regardless of symmetry or skewness, the data is strictly divided into fixed 25%25\% intervals:

      • 25%25\% of data values ≤Q1\le Q_1

      • 25%25\% of data values between Q1Q_1 and Median

      • 25%25\% of data values between Median and Q3Q_3

      • 25%25\% of data values ≥Q3\ge Q_3

  • Comprehensive Sample Dataset Comparison:

    • Dataset size: n=8n = 8 observations.

    • Calculated summary metrics:

      • Mean≠Median\text{Mean} \neq \text{Median}

      • Median=53.75\text{Median} = 53.75

      • Q1=50.5Q_1 = 50.5

      • Q3=61Q_3 = 61

      • IQR=61−50.5=10.5\text{IQR} = 61 - 50.5 = 10.5

      • Standard Deviation=7.78\text{Standard Deviation} = 7.78

    • Methodology Comparison:

      • Empirical Rule Approach (Incorrect for skewed data): Uses standard deviation (7.787.78) to establish a spread window of roughly 1616 total units (≈8\approx 8 units above the mean and ≈8\approx 8 units below the mean).

      • IQR Approach (Correct for skewed data): Uses central interquartile range (Q3−Q1Q_3 - Q_1) to establish a central spread window of 10.510.5 units around the median.