Descriptive Statistics: Graphical Displays and Location Measures
Graphical Representations of Frequency Distributions
Overview of graphical methods used to describe frequency distributions:
Stem-and-leaf graphs (stemplots)
Line graphs
Bar graphs
Histograms
Stem-and-Leaf Graphs (Stemplots):
Definition: A graphical method used to display the frequency distribution of a quantitative data set.
Anatomy of a Data Value:
Leaf: The final single digit of a data value.
Stem: All leading digits preceding the final digit.
Construction Procedure:
Construct a two-column table where each stem is listed in order down the left column and the corresponding leaves are listed in the right column.
Arrange the leaves in each row in ascending numerical order.
Outlier (Extreme Value): A data value that does not fit the overall pattern of the rest of the data because it is either unusually large or unusually small.
Worked Example — Exam Scores ():
Data set of 30 exam scores: 70, 62, 81, 75, 77, 78, 91, 63, 99, 75, 89, 80, 80, 78, 85, 95, 48, 75, 78, 81, 78, 57, 82, 72, 68, 70, 91, 78, 91, 63.
Corresponding Stem-and-Leaf Plot breakdown:
Stem 4 | Leaf: 8 | Class: 40–49 | Frequency: 1
Stem 5 | Leaf: 7 | Class: 50–59 | Frequency: 1
Stem 6 | Leaves: 2, 3, 3, 8 | Class: 60–69 | Frequency: 4
Stem 7 | Leaves: 0, 0, 2, 5, 5, 5, 7, 8, 8, 8, 8, 8 | Class: 70–79 | Frequency: 12
Stem 8 | Leaves: 0, 0, 1, 1, 2, 5, 9 | Class: 80–89 | Frequency: 7
Stem 9 | Leaves: 1, 1, 1, 5, 9 | Class: 90–99 | Frequency: 5
Side-by-Side Stem-and-Leaf Plots:
Definition and Function: A plot used to display and compare two data sets simultaneously.
Requirement: The two data sets must be measured on comparable scales so that they share a single column of stems in the center.
Worked Example — San Francisco Giants Wins and Losses (2010–2019):
Data Set (10 seasons):
2010: 92 Wins, 70 Losses
2011: 86 Wins, 76 Losses
2012: 94 Wins, 68 Losses
2013: 76 Wins, 86 Losses
2014: 88 Wins, 74 Losses
2015: 84 Wins, 78 Losses
2016: 87 Wins, 75 Losses
2017: 64 Wins, 98 Losses
2018: 73 Wins, 89 Losses
2019: 77 Wins, 85 Losses
Plot Structure (Wins leaves on left, shared Stem in center, Losses leaves on right):
Wins Leaves: 4 | Stem: 6 | Losses Leaves: 8
Wins Leaves: 7, 6, 3 | Stem: 7 | Losses Leaves: 0, 4, 5, 6, 8
Wins Leaves: 8, 7, 6, 4 | Stem: 8 | Losses Leaves: 5, 6, 9
Wins Leaves: 4, 2 | Stem: 9 | Losses Leaves: 8
Line Graphs:
Axes Setup:
Horizontal Axis: Represents the categories or specific data values/intervals for which frequencies are expressed.
Vertical Axis: Represents the frequency of each category or data value.
Construction Steps: Plot a dot at the frequency height directly above each category/value on the horizontal axis, then connect adjacent dots using straight line segments.
Example 1 — Categorical Data (Pets):
Categories: Cats (Frequency = 3), Dogs (Frequency = 15), Reptiles (Frequency = 5), Fish (Frequency = 9).
Example 2 — Continuous Rainfall Intervals:
Interval 2.95–4.97 inches: Frequency = 6
Interval 4.97–6.99 inches: Frequency = 7
Interval 6.99–9.01 inches: Frequency = 15
Interval 9.01–11.03 inches: Frequency = 8
Interval 11.03–13.05 inches: Frequency = 9
Interval 13.05–15.07 inches: Frequency = 5
Bar Graphs:
Definition and Comparison: Follows the same structural setup as a line graph, except that vertical rectangular bars are drawn up to the height of each plotted point instead of connecting the points with line segments.
Histograms, Frequency Polygons, and Time Series Graphs
Histograms:
Definition: A specialized type of bar graph used for continuous quantitative distributions.
Key Attributes:
Adjacent bars touch without gaps between them.
Horizontal Axis: Represents continuous class intervals of equal width.
Vertical Axis: Represents absolute frequencies or relative frequencies.
Construction: Can be constructed from raw data or an organized frequency table.
Boundary Adjustment Rule:
Class boundaries must be adjusted so that no observed data value falls directly on a boundary line.
For integer data sets, adjust boundaries by subtracting from the lower class limit and adding to the upper class limit (forming decimal limits).
Step-by-Step Construction Procedure:
Determine the class width and total number of classes.
Construct a comprehensive frequency table.
Adjust class boundaries to create contiguous, non-overlapping intervals.
Plot adjusted class boundaries on the horizontal axis and draw adjacent bars to the frequency height.
Complete Worked Example — Exam Scores ():
Class Intervals, Tallies, Absolute Frequencies, and Relative Frequencies:
Class 40–49: Tally = 1, Frequency = 1, Relative Frequency =
Class 50–59: Tally = 1, Frequency = 1, Relative Frequency =
Class 60–69: Tally = 4, Frequency = 4, Relative Frequency =
Class 70–79: Tally = 12, Frequency = 12, Relative Frequency =
Class 80–89: Tally = 7, Frequency = 7, Relative Frequency =
Class 90–99: Tally = 5, Frequency = 5, Relative Frequency =
Total Sample Size () = 30
Boundary Calculations:
Gap between class limits =
Half-gap adjustment =
Adjusted Class Boundaries: 39.5, 49.5, 59.5, 69.5, 79.5, 89.5, 99.5
Graphing Calculator Step-by-Step Procedure (TI-83/84):
Press
Y=. PressCLEARto delete any existing equations.Press
STAT, then select1:EDIT. Clear existing data inL1orL2by moving the cursor to the list header, pressingDEL(orCLEAR), and pressingENTER/ arrowing down.Enter the lower class limits or representative class values into
L1: 40, 50, 60, 70, 80, 90.Enter the corresponding class frequencies into
L2: 1, 1, 4, 12, 7, 5.Press
WINDOWand set the parameters:Xmin= 39.5Xmax= 99.5Xscl=Ymin= -5Ymax= 15Yscl= 1Xres= 1
Press
2ndY=(STAT PLOT). Select4:Plotsoffand pressENTERto turn off all existing plots.Press
2ndY=, select1:Plot1, and pressENTER. Arrow down toTYPEand select the 3rd icon (histogram plot). PressENTER.Arrow down to
Xlistand specifyL1(2nd1). Arrow down toFreqand specifyL2(2nd2).Press
GRAPH. Press theTRACEkey and use the left/right arrow keys to examine individual class intervals and bar heights.
Frequency Polygons:
Definition: A line-graph representation of a frequency distribution derived directly from a histogram.
Construction Steps:
Plot a point at the midpoint of the upper edge of each histogram bar.
Remove or erase the histogram bars.
Connect adjacent plotted midpoints with straight line segments.
Time Series Graphs:
Definition: A line graph where the horizontal axis represents specific progression points in time (e.g., years, months, days).
Example — San Francisco 49ers Regular Season Wins (2007–2011):
Season 2007: 5 Wins
Season 2008: 7 Wins
Season 2009: 8 Wins
Season 2010: 6 Wins
Season 2011: 13 Wins
Measures of Data Location
Percentiles:
Definition: The percentile of a data set is a value such that approximately of the data values are less than or equal to it, and approximately of the data values are greater than or equal to it.
Key Property: A percentile does not need to be an actual value contained within the observed data set.
Example Interpretation: Being at the percentile means the value is greater than or equal to of all values in the data set, and of the data set lies below or at this value.
Quartiles and Five-Number Summary:
First Quartile (): Equivalent to the percentile.
Median ( / Second Quartile): Equivalent to the percentile.
Third Quartile (): Equivalent to the percentile.
Minimum: The smallest individual data value in the set.
Maximum: The largest individual data value in the set.
Five-Number Summary: The ordered set consisting of:
Measures of Spread Derived from Quartiles:
Range: The total distance between maximum and minimum values:
Interquartile Range (IQR): The range of the middle of the data:
Procedure for Calculating Median, , and :
Sort all data values in ascending order from smallest to largest.
Determine the Median ():
If the total number of data values () is odd, the median is the single middle value.
If the total number of data values () is even, the median is the arithmetic average of the two middle values.
Determine the First Quartile ():
Isolate the lower half of the sorted data set (all values located strictly below the median position).
is the median of this lower half.
Determine the Third Quartile ():
Isolate the upper half of the sorted data set (all values located strictly above the median position).
is the median of this upper half.
Worked Example — 18-Element Data Set:
Raw Data Set: 67, 94, 85, 79, 53, 78, 87, 98, 86, 67, 90, 77, 84, 98, 66, 92, 79, 89 ( values).
Sorted Data Set: 53, 66, 67, 67, 77, 78, 79, 79, 84, 85, 86, 87, 89, 90, 92, 94, 98, 98
Calculation Steps:
Minimum = 53
Maximum = 98
Median Calculation (, even): The 9th value is 84 and the 10th value is 85.
First Quartile (): Lower half consists of 9 values: 53, 66, 67, 67, 77, 78, 79, 79, 84. The middle (5th) value is 77.
Third Quartile (): Upper half consists of 9 values: 85, 86, 87, 89, 90, 92, 94, 98, 98. The middle (5th) value is 90.
Complete Five-Number Summary: 53, 77, 84.5, 90, 98