Descriptive Statistics: Methods for Displaying and Summarizing Quantitative Data
Course Announcements & Administrative Details
Activity 1 Grading:
Activity 1 grades are posted.
A rubric is available to review specific questions where points were lost.
Solutions have been posted to allow comparison of calculations and numerical results.
Quiz 1:
Quiz 1 will be unlocked online for student access.
Chapter 2 Overview:
Chapter 2 covers a wide variety of graphical data representations and spans multiple class sessions over two weeks (approximately three lecture days).
Initial focus sections cover six primary graph types:
Stem-and-leaf graphs (stem plots)
Line graphs
Bar graphs
Histograms
Frequency polygons
Time series graphs
Stem-and-Leaf Graphs (Stem Plots)
Definition & Structure:
A stem-and-leaf graph (or stem plot) splits each data value into two components:
Leaf: Always contains the final significant digit.
Stem: Contains all leading digits preceding the final significant digit.
Example: For the number : this is the total(end/final number)
Construction Rules:
Stems are listed vertically from top to bottom in ascending order (smallest to largest).
Leaves are listed horizontally to the right of their corresponding stem in ascending order.
Decimal points are omitted in the plot layout unless explicitly specified by a key.
If a stem value has no corresponding data, the stem number must still be listed, and the leaf space must be left blank. Placing a as a leaf indicates a data value ending in (such as ), rather than an empty stem.
Outliers:
An outlier is an observation of data that does not fit the rest of the data set (a value far removed from the remaining cluster).
Stem-and-leaf plots allow for immediate visual identification of outliers and concentration of values.
Worked Example 1: Susan Bean's Spring Pre-Calculus Exam Scores:
Data Set:
Plot Construction:
Stem : leaf ()
Stem : leaves ()
Stem : leaves ()
Stem : leaves ()
Stem : leaves ()
Stem : leaves ()
Stem : leaves ()
Stem : leaf ()
Worked Example 2: Distances from Home to Local Supermarkets:
Data Set:
Plot Construction:
Stem : leaves
Stem : leaves
Stem : leaves
Stem : leaves
Stem : leaves
Stem : leaves
Stems : [blank]
Stem : leaf
Analysis: The data shows a concentration between and . The distance is located far away from the rest of the data, making a potential outlier.
Worked Example 3: Miles Per Gallon (MPG) Ratings for 30 Cars:
Data Set: MPG ratings recorded for vehicles.
Analysis: Stem plot values show a continuous grouping without large gaps between stems. Because all data points lie close together, there are no outliers present.
Line Graphs
Structure & Elements:
Horizontal Axis (-axis): Represents data values or measured categories.
Vertical Axis (-axis): Represents frequency (or relative frequency).
Plotting Points: Each data point is plotted at .
Line Segments: Sequential points are connected with straight line segments.
Labeling Requirement: Both axes must be explicitly labeled with titles and numbers. Omitting axis labels results in lost points.
Worked Example 1: Teenager Chore Reminders:
Data pair : reminders has a frequency of
Data pair : reminder has a frequency of
Points are plotted at and and connected via a line segment.
Worked Example 2: Car Repairs Per Year Survey:
Data Set ( people surveyed on annual car repairs):
times in shop: frequency
time in shop: frequency
times in shop: frequency
times in shop: frequency
Setup Procedure:
Draw and axes.
Label horizontal axis: "Number of times in shop".
Label vertical axis: "Frequency".
Scale horizontal axis with values: .
Scale vertical axis in increments of up to ().
Plot points: , , , .
Connect consecutive points with straight line segments.
Worked Example 3: Store Visits Prior to Major Purchase:
Survey of people measuring store visits prior to a purchase.
Line graph constructed identically using store visit categories on the -axis and frequency on the -axis.
Bar Graphs
Characteristics:
Used for categorical or discrete grouped data.
Crucial Property: The bars in a bar graph are separated from each other and do not touch.
Bars can be rendered vertically (up and down) or horizontally (sideways).
Worked Example 1: Facebook User Age Distribution:
Age group –:
Age group –:
Separated bars display percentages relative to age brackets.
Worked Example 2: US Public Schools Advanced Placement (AP) Examinees (Class of 2011):
Variables: Race/Ethnicity categories (-axis, coded through ) vs. AP Examinee Population percentages (-axis).
Category Percentages:
Category :
Category :
Category :
Category :
Category :
Category :
Setup: Vertical scale formatted by s (). Six distinct, non-touching bars plotted to their respective height percentages.
Worked Example 3: Student Birthdays by Season (Ms. Ramirez's Math Class):
Axes: Horizontal axis labeled "Season" (Winter, Spring, Summer, Fall); Vertical axis labeled "Proportion of Population".
Vertical Scale: Increments of
Results: Fall represents the tallest bar; Winter represents the shortest bar. All four seasonal bars remain non-touching.
Two-Way Tables, Marginal Distributions, and Conditional Distributions
Two-Way Table Structure:
Categorizes data simultaneously across two qualitative or quantitative variables.
Sample Data: Pet Ownership by Gender
Dogs: Men = , Women = , Total =
Cats: Men = , Women = , Total =
Fish: Men = , Women = , Total =
Column Totals: Men = , Women =
Grand Total:
Marginal Distribution Calculation:
Measures the total frequency or percentage of a single variable category relative to the overall total sample size ().
Sum Check: ( of the total population).
Conditional Distribution Calculation:
Measures proportions within a specific subpopulation (restricting focus to one row or column, ignoring all other data).
Subpopulation Example: Subpopulation of Men ().
Sum Check: ( of the male subpopulation).
Histograms
Overview & Key Properties:
The most important and frequently utilized graph type in statistics.
Unlike bar graphs, the bars in a histogram must touch.
Vertical Axis (-axis): Frequency or Relative Frequency.
Horizontal Axis (-axis): Quantitative data values (continuous or discrete scales).
Advantage over bar graphs: Captures order, intervals, continuous scale progression, and shape of numeric data distributions.
Method 1: Continuous Data using Class Boundaries (Male Semi-Professional Soccer Players Heights):
Data Range: Minimum height = , Maximum height =
Step 1: Choose Starting and Ending Points:
Avoid starting on an exact data value so values do not land on bar boundaries.
Step 2: Calculate Interval Width and Bar Boundaries:
Select desired number of bars (e.g., bars).
Round up to a convenient integer width:
Step 3: Establish Boundaries by Repeatedly Adding Width :
Bar 1: to
Bar 2: to
Bar 3: to
Bar 4: to
Bar 5: to
Bar 6: to
Bar 7: to
Bar 8: to
Step 4: Relative Frequency Heights:
: relative frequency =
: relative frequency =
: relative frequency =
: relative frequency =
: relative frequency =
: relative frequency =
: relative frequency =
: relative frequency =
Method 2: Discrete Data (Books Bought by 50 Part-Time College Students at ABC College):
Data is discrete (countable whole numbers through ).
-axis starts slightly before at and proceeds by integer steps ().
Frequencies:
book ( to ): frequency =
books ( to ): frequency =
books ( to ): frequency =
books ( to ): frequency =
books ( to ): frequency =
books ( to ): frequency =
Method 3: Continuous Data without Decimals using Left-Boundary Rule (Weekend Video Game Hours):
Data Range: Minimum = , Maximum =
Bin Setup: Intervals set at whole number steps of ().
Left-Boundary Inclusion Rule: If a data value falls exactly on an interval boundary, it is included in the bar where it serves as the left boundary (i.e., interval contains ).
Bar : Includes , excludes . Data points: (Frequency = ).
Bar : Includes , excludes . Data points: (Frequency = ).
Bar : Includes , excludes . Data points: (Frequency = ).
Bar : Includes , excludes . Data points: (Frequency = ).
Bar : Includes , excludes . Remaining data points (Frequency = ).
Frequency Polygons
Definition & Characteristics:
A graph constructed similarly to a line graph, with the specific modification that both ends of the line are anchored directly down to the horizontal axis (-axis), forming a closed polygon.
Highly effective for visually overlaying and comparing multiple data distributions simultaneously (e.g., comparing calculus final exam test scores against overall calculus final course grades).
Construction Procedure:
Calculate the midpoint for each class interval:
Plot points where .
Anchor the polygon: Add one extra midpoint step before the lowest class interval and one extra midpoint step after the highest class interval, setting their frequencies to . Connect the line down to the axis at these anchor points.
Worked Example: US Presidents' Ages at Inauguration:
Midpoint Calculations and Frequencies:
Interval –: Midpoint = , Frequency =
Interval –: Midpoint = , Frequency =
Interval –: Midpoint = , Frequency =
Interval –: Midpoint = , Frequency =
Interval –: Midpoint = , Frequency =
Interval –: Midpoint = , Frequency =
Pattern: Midpoints increase by increments of
Anchoring:
Lower anchor: Subtract from first midpoint (Frequency = ).
Upper anchor: Add to last midpoint (Frequency = ).
Plotting Sequence: Connect points sequentially across $x$-axis midpoints .
Time Series Graphs
Definition & Purpose:
A graph generated from a paired data set where one variable explicitly represents time (such as hours, days, months, or years) and the other variable represents a measured attribute.
Used to observe trends, fluctuations, and progression over time.
Worked Example: Annual Consumer Price Index (CPI) Over 10 Years:
Paired Variables: Year (-axis) paired with Annual Consumer Price Index (-axis).
Data Point Example: In year , the CPI was . Plotted at .
Structure: Years are spaced sequentially along the horizontal axis, CPI values along the vertical axis, and consecutive years are linked with line segments.
Questions & Discussion
Question: If there is a gap in stem numbers with no data values, do you still write the stem number?
Response: Yes, always list all intermediate stem numbers vertically. Leave the leaf side completely blank. Do not write as a leaf because represents an actual data value ending in zero.
Question: Should decimal points be written in the leaf section of a stem plot?
Response: No, decimal points do not need to be written in the leaf column.
Question: If a graph has a stem with no leaves, but a higher stem contains values, is that higher value considered an outlier?
Response: Yes, if a data value is separated from the main concentration of data by empty stems/gaps, it is considered an outlier.
Question: Why use a histogram instead of a bar graph?
Response: A histogram is superior for numeric data because the horizontal axis follows an ordered quantitative scale. In a bar graph, categories are often unordered qualitative groups without sequential numeric progression.