Unit 1D: Describing Quantitative Data with Numbers

Page 1: Introduction to Numerical Summaries

Section 1D focuses on numerical summaries of quantitative data, building upon the graphical methods (dotplots, stemplots, histograms) explored in the previous section. The objective is to provide precise descriptions of the center and variability of distributions.

Learning Targets
  • Find the median of a distribution of quantitative data.
  • Calculate the mean of a distribution of quantitative data.
  • Find the range of a distribution of quantitative data.
  • Calculate and interpret the standard deviation of a distribution of quantitative data.
  • Find the interquartile range (IQR) of a distribution of quantitative data.
  • Identify outliers in a distribution of quantitative data.
  • Choose appropriate measures of center and variability to summarize a distribution.
  • Make and interpret boxplots of quantitative data.
  • Use boxplots and summary statistics to compare distributions.
Illustrative Context: Flint, Michigan Lead Levels

Recall the context of the Flint, Michigan water crisis, where the water source was switched from Lake Huron to the Flint River. Data from 71 randomly selected dwellings showed lead levels in parts per billion (ppb):

  • Data Set: 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 1, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 4, 4, 4, 4, 4, 7, 8, 8, 9, 10, 10, 10, 11, 13, 18, 20, 21, 22, 29, 42, 42, 104.
  • Distribution Shape: Right-skewed and single-peaked (unimodal).
  • Outlier: The dwelling with a lead level of 104 ppb104\,ppb appears to be an outlier.
The Mode

The mode of a distribution is the most frequently occurring data value.

  • Example: For the Flint data, the mode is 0 ppb0\,ppb.
  • Limitation: The mode is often not a good measure of the center because it can fall anywhere in a distribution, there can be multiple modes, or there can be no mode at all. In the Flint case, 0 ppb0\,ppb is not representative of how much lead a typical dwelling has in its water.

Page 2: Measuring Center - The Median

Definition: Median

The median is the midpoint of a distribution—the number such that about half the observations are smaller and about half are larger.

To find the median:

  1. Arrange the data values from smallest to largest.
  2. If the number of data values nn is odd, the median is the middle value in the ordered list.
  3. If the number of data values nn is even, the median is the average of the two middle values in the ordered list.
Example: Population Density of Central America

Data on population density (people per km2km^2) for seven countries:

  • Countries: Belize, Costa Rica, El Salvador, Guatemala, Honduras, Nicaragua, Panama.
  • Values: 17, 100, 308, 158, 82, 48, 52.
  • Ordered List: 17, 48, 52, 82, 100, 158, 308.
  • Calculation: Total values n=7n = 7 (odd). The middle value is 82.
  • Result: Median = 82 people/km282\,people/km^2.
Example: Percent Air in Chip Bags

A group of chip enthusiasts collected data on the percentage of air in 14 popular brands of chips.

  • Data (sorted): 19, 28, 34, 39, 41, 41, 45, 46, 47, 48, 49, 50, 50, 59.
  • Calculation: Total values n=14n = 14 (even). The middle two values are 45 and 46.
  • Median Formula: 45+462=45.5%\frac{45 + 46}{2} = 45.5\%
  • Interpretation: About half of the chip brands have less than 45.5% air, and about half have more.

Page 3: Measuring Center - The Mean

Definition: The Mean

The mean is the arithmetic average of all individual data values. To find the mean, add all values and divide by the total number of values (nn).

Formula for Sample Mean (xˉ\bar{x}):xˉ=sum of data valuesnumber of data values=x1+x2+⋯+xnn=∑xin\bar{x} = \frac{\text{sum of data values}}{\text{number of data values}} = \frac{x_1 + x_2 + \dots + x_n}{n} = \frac{\sum x_i}{n}

  • xˉ\bar{x} is pronounced "x-bar."
  • ∑\sum (sigma) is the summation notation meaning "add them all up."
  • xix_i refers to individual observations.
Example: Percent Air in Chip Bags (Mean Calculation)

(a) Overall Mean:xˉ=46+59+48+19+47+41+41+39+45+28+50+50+49+3414=59614≈42.57%\bar{x} = \frac{46 + 59 + 48 + 19 + 47 + 41 + 41 + 39 + 45 + 28 + 50 + 50 + 49 + 34}{14} = \frac{596}{14} \approx 42.57\%

(b) Mean Without Possible Outlier: The bag of Fritos (19% air) is a possible outlier. Recalculating for the other 13 brands: xˉ=57713≈44.38%\bar{x} = \frac{577}{13} \approx 44.38\%

  • Observation: The outlier (Fritos) decreased the mean by 1.81 percentage points.

Page 4: Statistics vs. Parameters and Population Mean

Notation Clarification
  • Sample Mean (Statistic): Denoted by xˉ\bar{x}. Used for data from a sample.
  • Population Mean (Parameter): Denoted by μ\mu (mu). Used when data from the entire population is available.
Definition: Statistic and Parameter
  • Statistic: A number describing a characteristic of a sample.
  • Parameter: A number describing a characteristic of a population.
  • Mnemonic: S-S (Statistic-Sample), P-P (Parameter-Population).
Example: Population Mean Density

Calculating the population mean (μ\mu) for the seven Central American countries: μ=17+48+52+82+100+158+3087≈109.286 people/km2\mu = \frac{17 + 48 + 52 + 82 + 100 + 158 + 308}{7} \approx 109.286\,people/km^2

Page 5: Properties of the Mean and the Seesaw Activity

Resistance
  • Definition: A statistical measure is resistant if it is not affected much by extreme data values.
  • Mean: Not resistant. It is easily pulled by outliers or skewness.
  • Median: Is resistant. In the chip example, removing the 19% outlier changed the median from 45.5 to 46 (negligible change).
Activity: Interpreting the Mean as a Balance Point

Investigation of the mean's physical interpretation using a 12-inch ruler and pennies (seesaw experiment):

  1. Uniform Stack: 5 pennies at the 6-inch mark. The balance point (pencil) is exactly at 6, which is the mean of 6, 6, 6, 6, and 6.
  2. Shifting One Penny: Moving one penny to the 8-inch mark requires moving another to the 4-inch mark (or others) to keep the balance at 6. The mean remains 6.
  3. Physical Meaning: The mean is called the balance point of a distribution because the sum of the distances from the values to the mean on one side equals the sum of the distances on the other side.

Page 6: Comparing Mean and Median

The relationship between the mean and median is determined by the shape of the distribution and the presence of outliers.

Effect of Shape (Figure 1.10)
  1. Skewed to the Left: The mean is pulled toward the long tail.     * Relationship: Mean<Median\text{Mean} < \text{Median}.     * Example: Quiz scores (Mean=17.97\text{Mean} = 17.97, Median=19\text{Median} = 19).
  2. Roughly Symmetric: The mean and median are roughly equal.     * Relationship: Mean≈Median\text{Mean} \approx \text{Median}.     * Example: Head circumference (Mean=22.96\text{Mean} = 22.96, Median=23.05\text{Median} = 23.05).
  3. Skewed to the Right: The mean is pulled toward the long tail.     * Relationship: Mean>Median\text{Mean} > \text{Median}.     * Example: Runs scored (Mean=4.14\text{Mean} = 4.14, Median=3\text{Median} = 3).
Summary Rule
  • If roughly symmetric: Mean and Median are similar.
  • If strongly skewed: Mean is pulled in the direction of the skewness (tails), while the Median is not.
  • The Median is resistant to outliers; the Mean is not.
Real-World Example: MLB Salaries (2022)
  • Context: Heavily right-skewed. Most players earn near the minimum (700,000700,000), while stars (e.g., Max Scherzer, Mike Trout) earn over 20 million20\,million.
  • Median Salary: ≈1.2 million\approx 1.2\,million (represents the "typical" player).
  • Mean Salary: ≈4.4 million\approx 4.4\,million (influenced by high-paid superstars).
  • Usage: The mean is useful for calculating totals. Total payout = (Mean)×(Number of players)=(4.4 million)×(975)≈4.3 billion(\text{Mean}) \times (\text{Number of players}) = (4.4\,million) \times (975) \approx 4.3\,billion.

Page 7: Check Your Understanding (Pumpkin Weights)

Data (pounds): 3.6, 4.0, 9.6, 14.0, 11.0, 12.4, 13.0, 2.0, 6.0, 6.6, 15.0, 3.4, 12.7, 9.6, 4.0, 6.1, 6.0, 2.8, 5.4, 11.9, 5.4, 31.0, 33.0.

  1. Find the Median:     * Sorted: 2.0, 2.8, 3.4, 3.6, 4.0, 4.0, 5.4, 5.4, 6.0, 6.0, 6.1, [6.6], 9.6, 9.6, 11.0, 11.9, 12.4, 12.7, 13.0, 14.0, 15.0, 31.0, 33.0.     * Count n=23n = 23. The median is the 12th value: 6.6 lbs.
  2. Calculate the Mean: Sum of values = 228.5. xˉ=228.523≈9.93 lbs\bar{x} = \frac{228.5}{23} \approx 9.93\,lbs.
  3. Explanation: The distribution is strongly right-skewed with two high outliers (31.0 and 33.0). These large values pull the mean higher than the median because the mean is not resistant.

Page 8: Measuring Variability - The Range

Definition: Range

The range is the distance between the minimum and maximum values. Range=Maximum−Minimum\text{Range} = \text{Maximum} - \text{Minimum}

Comparison Example: PVC Pipe Suppliers

Two suppliers (A and B) provide samples with centers at 600 mm600\,mm, but different variability.

  • Supplier A: Range = 601.5−598.5=3.0 mm601.5 - 598.5 = 3.0\,mm.
  • Supplier B: Range = 604.0−596.0=8.0 mm604.0 - 596.0 = 8.0\,mm.
  • Conclusion: Supplier A is more consistent.
Limitations of Range
  • Single Number: It is not an interval (don't say "the range is 17 to 308").
  • Not Resistant: It depends entirely on the most extreme values, which may be outliers.
  • Film Strip Example: Machine A and Machine B produce film strips with the same range (70.2−69.8=0.4 mm70.2 - 69.8 = 0.4\,mm), but Machine B's data values vary more from the center than Machine A's.

Page 9: Measuring Variability - The Standard Deviation

Definition: Standard Deviation (sxs_x)

The standard deviation measures the "typical" distance of data values from the mean.

Standard Deviation Formula:sx=∑(xi−xˉ)2n−1s_x = \sqrt{\frac{\sum(x_i - \bar{x})^2}{n - 1}}

Sample Variance (sx2s_x^2): The value before taking the square root. Measured in squared units, which is less helpful for interpretation. sx2=∑(xi−xˉ)2n−1s_x^2 = \frac{\sum(x_i - \bar{x})^2}{n - 1}

How to Calculate Standard Deviation by Hand
  1. Find the mean (xˉ\bar{x}).
  2. Calculate the deviation of each value: deviation=value−mean\text{deviation} = \text{value} - \text{mean}.
  3. Square each deviation.
  4. Add all squared deviations and divide by n−1n - 1 (Sample Variance).
  5. Take the square root (Standard Deviation).

Page 10: Standard Deviation Example (Close Friends)

Data: 1, 2, 2, 2, 3, 3, 3, 3, 4, 4, 6. (n=11n = 11)

  1. Mean: xˉ=3311=3\bar{x} = \frac{33}{11} = 3.

  2. Table of Calculations:

    Value (xx)Deviation (x−xˉx - \bar{x})Squared Deviation ((x−xˉ)2(x - \bar{x})^2)
    11−3=−21 - 3 = -24
    22−3=−12 - 3 = -11
    22−3=−12 - 3 = -11
    22−3=−12 - 3 = -11
    33−3=03 - 3 = 00
    33−3=03 - 3 = 00
    33−3=03 - 3 = 00
    33−3=03 - 3 = 00
    44−3=14 - 3 = 11
    44−3=14 - 3 = 11
    66−3=36 - 3 = 39
    Sum018
  3. Variance: sx2=1811−1=1.80s_x^2 = \frac{18}{11 - 1} = 1.80.

  4. Standard Deviation: sx=1.80≈1.34 close friendss_x = \sqrt{1.80} \approx 1.34\,close\,friends.

  5. Interpretation: The number of close friends these students have typically varies from the mean by about 1.34 close friends.

Note on Notation: Population standard deviation is σ\sigma (sigma), calculated by dividing by NN instead of n−1n - 1.

Page 11: Properties of the Standard Deviation

  • sxs_x is always ≥0\ge 0: It is 0 only when all data values are identical.
  • Greater Variation: Larger distances from the mean result in a larger sxs_x.
  • Non-Resistant: It is even more sensitive to outliers than the mean because deviations are squared.
  • Context: It should only be used when the mean is the chosen measure of center.
Effect of Adding a Value at the Mean

If a 12th value (equal to the mean of 3) is added to the friend data:

  • The mean remains 3.
  • The sum of squared deviations remains 18.
  • The denominator becomes 12−1=1112 - 1 = 11.
  • New sx=1811≈1.28s_x = \sqrt{\frac{18}{11}} \approx 1.28. The standard deviation decreases because the new value has a distance of 0 from the mean.

Page 12: Measuring Variability - The Interquartile Range (IQR)

To avoid the impact of extreme values, focus on the middle of the distribution using quartiles.

Definition: Quartiles

Quartiles divide an ordered data set into four groups of roughly equal size.

  • First Quartile (Q1Q_1): The median of the data values to the left of the actual median.
  • Second Quartile: The Median.
  • Third Quartile (Q3Q_3): The median of the data values to the right of the actual median.
Definition: Interquartile Range (IQR)

The IQR is the distance between the first and third quartiles. IQR=Q3−Q1\text{IQR} = Q_3 - Q_1

Example: Charity Collection

Data: $21, 22, 22, | 25, 26, 28, | 29, 31, 31, | 34, 37, 39 (n=12n=12)

  • Median: Average of 28 and 29 = $28.50.
  • Q1Q_1: Median of left half (19, 22, 22, 25, 26, 28) = $23.50.
  • Q3Q_3: Median of right half (29, 31, 31, 34, 37, 39) = $32.50.
  • IQR: 32.50−23.50=9.0032.50 - 23.50 = 9.00.

Page 13: IQR Example and Interpretation

Example: Percent Air in Chip Bags (IQR)
  • Sorted Data: 19, 28, 34, [39], 41, 41, 45, | 46, 47, 48, [49], 50, 50, 59.
  • Median: 45.5.
  • Q1Q_1: 39.
  • Q3Q_3: 49.
  • IQR: 49−39=10%49 - 39 = 10\%.
  • Interpretation: The middle half of the distribution of percent air in these 14 bags of chips has a range of 10%.
Property of IQR
  • Resistant: It is not affected by outliers since it only considers the middle 50% of data.
  • Calculation Tip: When the median is a data value (as in the Friend example), ignore it when finding Q1Q_1 and Q3Q_3.

Page 14: Choosing Summary Statistics

The choice of statistics depends on shape and outliers.

Guidelines
  • Symmetric / No Outliers: Use the Mean (xˉ\bar{x}) and Standard Deviation (sxs_x).
  • Skewed / Outliers: Use the Median and IQR.
  • General Note: Range is used only as a last resort because it provides very little information about the distribution apart from the extrema.
Example: Flint Lead Levels (Choosing Statistics)
  • Data Characteristics: Strongly right-skewed with an obvious outlier at 104 ppb104\,ppb.
  • Choice: Median and IQR.
  • Values: Median = 3 ppb3\,ppb; Q1=2,Q3=7Q_1 = 2, Q_3 = 7, so IQR=7−2=5 ppb\text{IQR} = 7 - 2 = 5\,ppb.

Page 15: Technology Corner - Calculating Summary Statistics

TI-83/84 Instructions
  1. Enter data into list L1 via STAT -> Edit.
  2. Go to STAT -> CALC -> 1-Var Stats.
  3. Specify List: L1 and leave FreqList blank.
  4. Select Calculate and press ENTER.
  5. The screen reports: xˉ\bar{x} (mean), ∑x\sum x, ∑x2\sum x^2, sxs_x (sample SD), σx\sigma_x (population SD), nn (sample size), and the five-number summary (min, Q1Q_1, Med, Q3Q_3, max).

Caution: Different software (like Minitab) may use slightly different rules for finding quartiles, resulting in small variations in Q1Q_1 and Q3Q_3 values compared to a TI-84.

Page 16: Identifying Outliers

How extreme must a value be to be considered an outlier? The most common method uses the IQR as a "ruler."

Standard Rule: The 1.5×IQR1.5 \times \text{IQR} Rule

An observation is an outlier if it falls:

  • Below Q1−1.5×IQRQ_1 - 1.5 \times \text{IQR}
  • Above Q3+1.5×IQRQ_3 + 1.5 \times \text{IQR}
Example: LeBron James Scoring (First 16 Seasons)

Sorted Data: 20.9, 25.3, 25.3, 26.4, 26.7, 26.8, 27.1, 27.1, 27.2, 27.3, 27.4, 27.5, 28.4, 29.7, 30.0, 31.4.

  • Median: 27.15.
  • Q1Q_1: 26.55.
  • Q3Q_3: 27.95.
  • IQR: 27.95−26.55=1.4027.95 - 26.55 = 1.40.
  • Lower Cutoff: 26.55−1.5(1.40)=24.4526.55 - 1.5(1.40) = 24.45.
  • Upper Cutoff: 27.95+1.5(1.40)=30.0527.95 + 1.5(1.40) = 30.05.
  • Outliers Identified:     * 20.920.9 (Rookie season) is a low outlier (less than 24.45).     * 31.431.4 is a high outlier (greater than 30.05).

Page 17-18: Alternative Outlier Rules and Why Outliers Matter

The 2×SD2 \times \text{SD} Rule

Some identify outliers as values more than 2 standard deviations from the mean.

  • Context: Mean = 27.156, SD = 2.328.
  • Cutoffs: 27.156±2(2.328)→22.50 to 31.8127.156 \pm 2(2.328) \rightarrow 22.50 \text{ to } 31.81.
  • Result: By this rule, only 20.9 is an outlier. 31.4 is not.
  • Preference: Always use the 1.5×IQR1.5 \times \text{IQR} rule unless otherwise directed, as it is based on resistant statistics.
Reasons to Identify Outliers
  1. Inaccurate Data: Recording errors or equipment malfunctions.
  2. Remarkable Occurrences: They highlight extraordinary individuals (e.g., Serena Williams' earnings).
  3. Influence: They heavily skew the mean, range, and standard deviation.

Page 19-20: Boxplots (Box-and-Whisker Plots)

Definition: Five-number Summary

Consists of the Minimum, Q1Q_1, Median, Q3Q_3, and Maximum.

How to Make a Boxplot
  1. Find the five-number summary.
  2. Identify outliers using the 1.5×IQR1.5 \times \text{IQR} rule.
  3. Draw a horizontal axis with the name of the variable and units.
  4. Draw a box from Q1Q_1 to Q3Q_3.
  5. Mark the Median with a vertical line inside the box.
  6. Mark any outliers with a symbol like an asterisk (∗*).
  7. Draw whiskers from the box to the smallest and largest values that are not outliers.
Example: Burger King Large Fries Weights

Data: 165, 172, 173, 176, 178, 179, 179, 180, 181, 181, 183, 183, 184, 186, 187.

  • Summary: Min=165, Q1=176Q_1=176, Med=180, Q3=183Q_3=183, Max=187.
  • IQR: 183−176=7183 - 176 = 7.
  • Low Outlier Cutoff: 176−10.5=165.5176 - 10.5 = 165.5.
  • Result: The 165g order is an outlier. The lower whisker goes to 172g.
  • Conclusion: Since Q1=176Q_1 = 176, at least 75% of orders weighed more than the advertised 173g. Suspicion of serving size exaggeration is not supported.
Limitations of Boxplots
  • They do not show individual data values.
  • They can hide features like peaks, gaps, and clusters (e.g., bimodal Old Faithful eruptions look unimodal in a boxplot).

Page 21-23: Comparing Distributions with Boxplots

When using boxplots to compare groups, always address: Shape, Outliers, Center, and Variability (SOCV) in context.

Example: Tablet Ratings (Apple vs. Samsung)
  • Shape: Both are left-skewed.
  • Outliers: Apple has two low outliers (73 and 76). Samsung has none.
  • Center: Apple has a slightly higher median rating (84) than Samsung (83).
  • Variability: Samsung ratings are much more variable. Samsung's IQR (11) is nearly 4 times larger than Apple's IQR (3). Apple's ratings are more consistent.
Critical Tip for AP Exam
  • Use statistical terms precisely. Do not say "mean" if you calculated the "median."
  • Range, Q1Q_1, Q3Q_3, and IQR are single numbers, not regions or intervals.
  • Always compare explicitly: use words like "greater than," "less than," or "about the same as."

Page 24: Questions & Discussion / Team Challenge

Dispute: Did Mr. Starnes Stack His AP Statistics Class?

Context: Mr. Starnes and Ms. McGrail teach AP Statistics. Mr. Starnes' students averaged 8 points higher on a test. Ms. McGrail suspects he assigned better students (higher GPAs) to his own class.

Data (GPA):

  • McGrail: 3.300, 2.900, 2.850, 3.100, 2.860, 2.900, 3.400, 3.338, 3.245, 3.085, 3.200, 3.000, 2.800, 2.900, 3.100, 2.600, 3.600, 3.200, 2.700, 3.560.
  • Starnes: 3.200, 3.500, 2.800, 2.900, 3.950, 3.980, 2.900, 2.900, 3.000, 3.000, 2.800, 3.750, 3.800, 3.200, 2.900, 3.100.

Analysis Strategy: Students are encouraged to use technology to create parallel boxplots and summary statistics to compare the GPA distributions and determine if one class is significantly "stronger."