Unit 1D: Describing Quantitative Data with Numbers
Page 1: Introduction to Numerical Summaries
Section 1D focuses on numerical summaries of quantitative data, building upon the graphical methods (dotplots, stemplots, histograms) explored in the previous section. The objective is to provide precise descriptions of the center and variability of distributions.
Learning Targets
- Find the median of a distribution of quantitative data.
- Calculate the mean of a distribution of quantitative data.
- Find the range of a distribution of quantitative data.
- Calculate and interpret the standard deviation of a distribution of quantitative data.
- Find the interquartile range (IQR) of a distribution of quantitative data.
- Identify outliers in a distribution of quantitative data.
- Choose appropriate measures of center and variability to summarize a distribution.
- Make and interpret boxplots of quantitative data.
- Use boxplots and summary statistics to compare distributions.
Illustrative Context: Flint, Michigan Lead Levels
Recall the context of the Flint, Michigan water crisis, where the water source was switched from Lake Huron to the Flint River. Data from 71 randomly selected dwellings showed lead levels in parts per billion (ppb):
- Data Set: 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 1, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 3, 4, 4, 4, 4, 4, 7, 8, 8, 9, 10, 10, 10, 11, 13, 18, 20, 21, 22, 29, 42, 42, 104.
- Distribution Shape: Right-skewed and single-peaked (unimodal).
- Outlier: The dwelling with a lead level of appears to be an outlier.
The Mode
The mode of a distribution is the most frequently occurring data value.
- Example: For the Flint data, the mode is .
- Limitation: The mode is often not a good measure of the center because it can fall anywhere in a distribution, there can be multiple modes, or there can be no mode at all. In the Flint case, is not representative of how much lead a typical dwelling has in its water.
Page 2: Measuring Center - The Median
Definition: Median
The median is the midpoint of a distribution—the number such that about half the observations are smaller and about half are larger.
To find the median:
- Arrange the data values from smallest to largest.
- If the number of data values is odd, the median is the middle value in the ordered list.
- If the number of data values is even, the median is the average of the two middle values in the ordered list.
Example: Population Density of Central America
Data on population density (people per ) for seven countries:
- Countries: Belize, Costa Rica, El Salvador, Guatemala, Honduras, Nicaragua, Panama.
- Values: 17, 100, 308, 158, 82, 48, 52.
- Ordered List: 17, 48, 52, 82, 100, 158, 308.
- Calculation: Total values (odd). The middle value is 82.
- Result: Median = .
Example: Percent Air in Chip Bags
A group of chip enthusiasts collected data on the percentage of air in 14 popular brands of chips.
- Data (sorted): 19, 28, 34, 39, 41, 41, 45, 46, 47, 48, 49, 50, 50, 59.
- Calculation: Total values (even). The middle two values are 45 and 46.
- Median Formula:
- Interpretation: About half of the chip brands have less than 45.5% air, and about half have more.
Page 3: Measuring Center - The Mean
Definition: The Mean
The mean is the arithmetic average of all individual data values. To find the mean, add all values and divide by the total number of values ().
Formula for Sample Mean ():
- is pronounced "x-bar."
- (sigma) is the summation notation meaning "add them all up."
- refers to individual observations.
Example: Percent Air in Chip Bags (Mean Calculation)
(a) Overall Mean:
(b) Mean Without Possible Outlier: The bag of Fritos (19% air) is a possible outlier. Recalculating for the other 13 brands:
- Observation: The outlier (Fritos) decreased the mean by 1.81 percentage points.
Page 4: Statistics vs. Parameters and Population Mean
Notation Clarification
- Sample Mean (Statistic): Denoted by . Used for data from a sample.
- Population Mean (Parameter): Denoted by (mu). Used when data from the entire population is available.
Definition: Statistic and Parameter
- Statistic: A number describing a characteristic of a sample.
- Parameter: A number describing a characteristic of a population.
- Mnemonic: S-S (Statistic-Sample), P-P (Parameter-Population).
Example: Population Mean Density
Calculating the population mean () for the seven Central American countries:
Page 5: Properties of the Mean and the Seesaw Activity
Resistance
- Definition: A statistical measure is resistant if it is not affected much by extreme data values.
- Mean: Not resistant. It is easily pulled by outliers or skewness.
- Median: Is resistant. In the chip example, removing the 19% outlier changed the median from 45.5 to 46 (negligible change).
Activity: Interpreting the Mean as a Balance Point
Investigation of the mean's physical interpretation using a 12-inch ruler and pennies (seesaw experiment):
- Uniform Stack: 5 pennies at the 6-inch mark. The balance point (pencil) is exactly at 6, which is the mean of 6, 6, 6, 6, and 6.
- Shifting One Penny: Moving one penny to the 8-inch mark requires moving another to the 4-inch mark (or others) to keep the balance at 6. The mean remains 6.
- Physical Meaning: The mean is called the balance point of a distribution because the sum of the distances from the values to the mean on one side equals the sum of the distances on the other side.
Page 6: Comparing Mean and Median
The relationship between the mean and median is determined by the shape of the distribution and the presence of outliers.
Effect of Shape (Figure 1.10)
- Skewed to the Left: The mean is pulled toward the long tail. * Relationship: . * Example: Quiz scores (, ).
- Roughly Symmetric: The mean and median are roughly equal. * Relationship: . * Example: Head circumference (, ).
- Skewed to the Right: The mean is pulled toward the long tail. * Relationship: . * Example: Runs scored (, ).
Summary Rule
- If roughly symmetric: Mean and Median are similar.
- If strongly skewed: Mean is pulled in the direction of the skewness (tails), while the Median is not.
- The Median is resistant to outliers; the Mean is not.
Real-World Example: MLB Salaries (2022)
- Context: Heavily right-skewed. Most players earn near the minimum (), while stars (e.g., Max Scherzer, Mike Trout) earn over .
- Median Salary: (represents the "typical" player).
- Mean Salary: (influenced by high-paid superstars).
- Usage: The mean is useful for calculating totals. Total payout = .
Page 7: Check Your Understanding (Pumpkin Weights)
Data (pounds): 3.6, 4.0, 9.6, 14.0, 11.0, 12.4, 13.0, 2.0, 6.0, 6.6, 15.0, 3.4, 12.7, 9.6, 4.0, 6.1, 6.0, 2.8, 5.4, 11.9, 5.4, 31.0, 33.0.
- Find the Median: * Sorted: 2.0, 2.8, 3.4, 3.6, 4.0, 4.0, 5.4, 5.4, 6.0, 6.0, 6.1, [6.6], 9.6, 9.6, 11.0, 11.9, 12.4, 12.7, 13.0, 14.0, 15.0, 31.0, 33.0. * Count . The median is the 12th value: 6.6 lbs.
- Calculate the Mean: Sum of values = 228.5. .
- Explanation: The distribution is strongly right-skewed with two high outliers (31.0 and 33.0). These large values pull the mean higher than the median because the mean is not resistant.
Page 8: Measuring Variability - The Range
Definition: Range
The range is the distance between the minimum and maximum values.
Comparison Example: PVC Pipe Suppliers
Two suppliers (A and B) provide samples with centers at , but different variability.
- Supplier A: Range = .
- Supplier B: Range = .
- Conclusion: Supplier A is more consistent.
Limitations of Range
- Single Number: It is not an interval (don't say "the range is 17 to 308").
- Not Resistant: It depends entirely on the most extreme values, which may be outliers.
- Film Strip Example: Machine A and Machine B produce film strips with the same range (), but Machine B's data values vary more from the center than Machine A's.
Page 9: Measuring Variability - The Standard Deviation
Definition: Standard Deviation ()
The standard deviation measures the "typical" distance of data values from the mean.
Standard Deviation Formula:
Sample Variance (): The value before taking the square root. Measured in squared units, which is less helpful for interpretation.
How to Calculate Standard Deviation by Hand
- Find the mean ().
- Calculate the deviation of each value: .
- Square each deviation.
- Add all squared deviations and divide by (Sample Variance).
- Take the square root (Standard Deviation).
Page 10: Standard Deviation Example (Close Friends)
Data: 1, 2, 2, 2, 3, 3, 3, 3, 4, 4, 6. ()
Mean: .
Table of Calculations:
Value () Deviation () Squared Deviation () 1 4 2 1 2 1 2 1 3 0 3 0 3 0 3 0 4 1 4 1 6 9 Sum 0 18 Variance: .
Standard Deviation: .
Interpretation: The number of close friends these students have typically varies from the mean by about 1.34 close friends.
Note on Notation: Population standard deviation is (sigma), calculated by dividing by instead of .
Page 11: Properties of the Standard Deviation
- is always : It is 0 only when all data values are identical.
- Greater Variation: Larger distances from the mean result in a larger .
- Non-Resistant: It is even more sensitive to outliers than the mean because deviations are squared.
- Context: It should only be used when the mean is the chosen measure of center.
Effect of Adding a Value at the Mean
If a 12th value (equal to the mean of 3) is added to the friend data:
- The mean remains 3.
- The sum of squared deviations remains 18.
- The denominator becomes .
- New . The standard deviation decreases because the new value has a distance of 0 from the mean.
Page 12: Measuring Variability - The Interquartile Range (IQR)
To avoid the impact of extreme values, focus on the middle of the distribution using quartiles.
Definition: Quartiles
Quartiles divide an ordered data set into four groups of roughly equal size.
- First Quartile (): The median of the data values to the left of the actual median.
- Second Quartile: The Median.
- Third Quartile (): The median of the data values to the right of the actual median.
Definition: Interquartile Range (IQR)
The IQR is the distance between the first and third quartiles.
Example: Charity Collection
Data: $21, 22, 22, | 25, 26, 28, | 29, 31, 31, | 34, 37, 39 ()
- Median: Average of 28 and 29 = $28.50.
- : Median of left half (19, 22, 22, 25, 26, 28) = $23.50.
- : Median of right half (29, 31, 31, 34, 37, 39) = $32.50.
- IQR: .
Page 13: IQR Example and Interpretation
Example: Percent Air in Chip Bags (IQR)
- Sorted Data: 19, 28, 34, [39], 41, 41, 45, | 46, 47, 48, [49], 50, 50, 59.
- Median: 45.5.
- : 39.
- : 49.
- IQR: .
- Interpretation: The middle half of the distribution of percent air in these 14 bags of chips has a range of 10%.
Property of IQR
- Resistant: It is not affected by outliers since it only considers the middle 50% of data.
- Calculation Tip: When the median is a data value (as in the Friend example), ignore it when finding and .
Page 14: Choosing Summary Statistics
The choice of statistics depends on shape and outliers.
Guidelines
- Symmetric / No Outliers: Use the Mean () and Standard Deviation ().
- Skewed / Outliers: Use the Median and IQR.
- General Note: Range is used only as a last resort because it provides very little information about the distribution apart from the extrema.
Example: Flint Lead Levels (Choosing Statistics)
- Data Characteristics: Strongly right-skewed with an obvious outlier at .
- Choice: Median and IQR.
- Values: Median = ; , so .
Page 15: Technology Corner - Calculating Summary Statistics
TI-83/84 Instructions
- Enter data into list L1 via
STAT->Edit. - Go to
STAT->CALC->1-Var Stats. - Specify
List: L1and leaveFreqListblank. - Select
Calculateand pressENTER. - The screen reports: (mean), , , (sample SD), (population SD), (sample size), and the five-number summary (min, , Med, , max).
Caution: Different software (like Minitab) may use slightly different rules for finding quartiles, resulting in small variations in and values compared to a TI-84.
Page 16: Identifying Outliers
How extreme must a value be to be considered an outlier? The most common method uses the IQR as a "ruler."
Standard Rule: The Rule
An observation is an outlier if it falls:
- Below
- Above
Example: LeBron James Scoring (First 16 Seasons)
Sorted Data: 20.9, 25.3, 25.3, 26.4, 26.7, 26.8, 27.1, 27.1, 27.2, 27.3, 27.4, 27.5, 28.4, 29.7, 30.0, 31.4.
- Median: 27.15.
- : 26.55.
- : 27.95.
- IQR: .
- Lower Cutoff: .
- Upper Cutoff: .
- Outliers Identified: * (Rookie season) is a low outlier (less than 24.45). * is a high outlier (greater than 30.05).
Page 17-18: Alternative Outlier Rules and Why Outliers Matter
The Rule
Some identify outliers as values more than 2 standard deviations from the mean.
- Context: Mean = 27.156, SD = 2.328.
- Cutoffs: .
- Result: By this rule, only 20.9 is an outlier. 31.4 is not.
- Preference: Always use the rule unless otherwise directed, as it is based on resistant statistics.
Reasons to Identify Outliers
- Inaccurate Data: Recording errors or equipment malfunctions.
- Remarkable Occurrences: They highlight extraordinary individuals (e.g., Serena Williams' earnings).
- Influence: They heavily skew the mean, range, and standard deviation.
Page 19-20: Boxplots (Box-and-Whisker Plots)
Definition: Five-number Summary
Consists of the Minimum, , Median, , and Maximum.
How to Make a Boxplot
- Find the five-number summary.
- Identify outliers using the rule.
- Draw a horizontal axis with the name of the variable and units.
- Draw a box from to .
- Mark the Median with a vertical line inside the box.
- Mark any outliers with a symbol like an asterisk ().
- Draw whiskers from the box to the smallest and largest values that are not outliers.
Example: Burger King Large Fries Weights
Data: 165, 172, 173, 176, 178, 179, 179, 180, 181, 181, 183, 183, 184, 186, 187.
- Summary: Min=165, , Med=180, , Max=187.
- IQR: .
- Low Outlier Cutoff: .
- Result: The 165g order is an outlier. The lower whisker goes to 172g.
- Conclusion: Since , at least 75% of orders weighed more than the advertised 173g. Suspicion of serving size exaggeration is not supported.
Limitations of Boxplots
- They do not show individual data values.
- They can hide features like peaks, gaps, and clusters (e.g., bimodal Old Faithful eruptions look unimodal in a boxplot).
Page 21-23: Comparing Distributions with Boxplots
When using boxplots to compare groups, always address: Shape, Outliers, Center, and Variability (SOCV) in context.
Example: Tablet Ratings (Apple vs. Samsung)
- Shape: Both are left-skewed.
- Outliers: Apple has two low outliers (73 and 76). Samsung has none.
- Center: Apple has a slightly higher median rating (84) than Samsung (83).
- Variability: Samsung ratings are much more variable. Samsung's IQR (11) is nearly 4 times larger than Apple's IQR (3). Apple's ratings are more consistent.
Critical Tip for AP Exam
- Use statistical terms precisely. Do not say "mean" if you calculated the "median."
- Range, , , and IQR are single numbers, not regions or intervals.
- Always compare explicitly: use words like "greater than," "less than," or "about the same as."
Page 24: Questions & Discussion / Team Challenge
Dispute: Did Mr. Starnes Stack His AP Statistics Class?
Context: Mr. Starnes and Ms. McGrail teach AP Statistics. Mr. Starnes' students averaged 8 points higher on a test. Ms. McGrail suspects he assigned better students (higher GPAs) to his own class.
Data (GPA):
- McGrail: 3.300, 2.900, 2.850, 3.100, 2.860, 2.900, 3.400, 3.338, 3.245, 3.085, 3.200, 3.000, 2.800, 2.900, 3.100, 2.600, 3.600, 3.200, 2.700, 3.560.
- Starnes: 3.200, 3.500, 2.800, 2.900, 3.950, 3.980, 2.900, 2.900, 3.000, 3.000, 2.800, 3.750, 3.800, 3.200, 2.900, 3.100.
Analysis Strategy: Students are encouraged to use technology to create parallel boxplots and summary statistics to compare the GPA distributions and determine if one class is significantly "stronger."