Descriptive Statistics: Graphical Displays and Location Measures

Graphical Representations of Frequency Distributions

  • Overview of graphical methods used to describe frequency distributions:

    • Stem-and-leaf graphs (stemplots)

    • Line graphs

    • Bar graphs

    • Histograms

  • Stem-and-Leaf Graphs (Stemplots):

    • Definition: A graphical method used to display the frequency distribution of a quantitative data set.

    • Anatomy of a Data Value:

    • Leaf: The final single digit of a data value.

    • Stem: All leading digits preceding the final digit.

    • Construction Procedure:

    • Construct a two-column table where each stem is listed in order down the left column and the corresponding leaves are listed in the right column.

    • Arrange the leaves in each row in ascending numerical order.

    • Outlier (Extreme Value): A data value that does not fit the overall pattern of the rest of the data because it is either unusually large or unusually small.

    • Worked Example — Exam Scores (n=30n = 30):

    • Data set of 30 exam scores: 70, 62, 81, 75, 77, 78, 91, 63, 99, 75, 89, 80, 80, 78, 85, 95, 48, 75, 78, 81, 78, 57, 82, 72, 68, 70, 91, 78, 91, 63.

    • Corresponding Stem-and-Leaf Plot breakdown:

      • Stem 4 | Leaf: 8 | Class: 40–49 | Frequency: 1

      • Stem 5 | Leaf: 7 | Class: 50–59 | Frequency: 1

      • Stem 6 | Leaves: 2, 3, 3, 8 | Class: 60–69 | Frequency: 4

      • Stem 7 | Leaves: 0, 0, 2, 5, 5, 5, 7, 8, 8, 8, 8, 8 | Class: 70–79 | Frequency: 12

      • Stem 8 | Leaves: 0, 0, 1, 1, 2, 5, 9 | Class: 80–89 | Frequency: 7

      • Stem 9 | Leaves: 1, 1, 1, 5, 9 | Class: 90–99 | Frequency: 5

  • Side-by-Side Stem-and-Leaf Plots:

    • Definition and Function: A plot used to display and compare two data sets simultaneously.

    • Requirement: The two data sets must be measured on comparable scales so that they share a single column of stems in the center.

    • Worked Example — San Francisco Giants Wins and Losses (2010–2019):

    • Data Set (10 seasons):

      • 2010: 92 Wins, 70 Losses

      • 2011: 86 Wins, 76 Losses

      • 2012: 94 Wins, 68 Losses

      • 2013: 76 Wins, 86 Losses

      • 2014: 88 Wins, 74 Losses

      • 2015: 84 Wins, 78 Losses

      • 2016: 87 Wins, 75 Losses

      • 2017: 64 Wins, 98 Losses

      • 2018: 73 Wins, 89 Losses

      • 2019: 77 Wins, 85 Losses

    • Plot Structure (Wins leaves on left, shared Stem in center, Losses leaves on right):

      • Wins Leaves: 4 | Stem: 6 | Losses Leaves: 8

      • Wins Leaves: 7, 6, 3 | Stem: 7 | Losses Leaves: 0, 4, 5, 6, 8

      • Wins Leaves: 8, 7, 6, 4 | Stem: 8 | Losses Leaves: 5, 6, 9

      • Wins Leaves: 4, 2 | Stem: 9 | Losses Leaves: 8

  • Line Graphs:

    • Axes Setup:

    • Horizontal Axis: Represents the categories or specific data values/intervals for which frequencies are expressed.

    • Vertical Axis: Represents the frequency of each category or data value.

    • Construction Steps: Plot a dot at the frequency height directly above each category/value on the horizontal axis, then connect adjacent dots using straight line segments.

    • Example 1 — Categorical Data (Pets):

    • Categories: Cats (Frequency = 3), Dogs (Frequency = 15), Reptiles (Frequency = 5), Fish (Frequency = 9).

    • Example 2 — Continuous Rainfall Intervals:

    • Interval 2.95–4.97 inches: Frequency = 6

    • Interval 4.97–6.99 inches: Frequency = 7

    • Interval 6.99–9.01 inches: Frequency = 15

    • Interval 9.01–11.03 inches: Frequency = 8

    • Interval 11.03–13.05 inches: Frequency = 9

    • Interval 13.05–15.07 inches: Frequency = 5

  • Bar Graphs:

    • Definition and Comparison: Follows the same structural setup as a line graph, except that vertical rectangular bars are drawn up to the height of each plotted point instead of connecting the points with line segments.

Histograms, Frequency Polygons, and Time Series Graphs

  • Histograms:

    • Definition: A specialized type of bar graph used for continuous quantitative distributions.

    • Key Attributes:

    • Adjacent bars touch without gaps between them.

    • Horizontal Axis: Represents continuous class intervals of equal width.

    • Vertical Axis: Represents absolute frequencies or relative frequencies.

    • Construction: Can be constructed from raw data or an organized frequency table.

    • Boundary Adjustment Rule:

    • Class boundaries must be adjusted so that no observed data value falls directly on a boundary line.

    • For integer data sets, adjust boundaries by subtracting 0.50.5 from the lower class limit and adding 0.50.5 to the upper class limit (forming decimal limits).

    • Step-by-Step Construction Procedure:

    1. Determine the class width and total number of classes.

    2. Construct a comprehensive frequency table.

    3. Adjust class boundaries to create contiguous, non-overlapping intervals.

    4. Plot adjusted class boundaries on the horizontal axis and draw adjacent bars to the frequency height.

    • Complete Worked Example — Exam Scores (n=30n = 30):

    • Class Intervals, Tallies, Absolute Frequencies, and Relative Frequencies:

      • Class 40–49: Tally = 1, Frequency = 1, Relative Frequency = 130\frac{1}{30}

      • Class 50–59: Tally = 1, Frequency = 1, Relative Frequency = 130\frac{1}{30}

      • Class 60–69: Tally = 4, Frequency = 4, Relative Frequency = 430=215\frac{4}{30} = \frac{2}{15}

      • Class 70–79: Tally = 12, Frequency = 12, Relative Frequency = 1230=25\frac{12}{30} = \frac{2}{5}

      • Class 80–89: Tally = 7, Frequency = 7, Relative Frequency = 730\frac{7}{30}

      • Class 90–99: Tally = 5, Frequency = 5, Relative Frequency = 530=16\frac{5}{30} = \frac{1}{6}

      • Total Sample Size (nn) = 30

    • Boundary Calculations:

      • Gap between class limits = 11

      • Half-gap adjustment = 0.50.5

      • Adjusted Class Boundaries: 39.5, 49.5, 59.5, 69.5, 79.5, 89.5, 99.5

    • Graphing Calculator Step-by-Step Procedure (TI-83/84):

    1. Press Y=. Press CLEAR to delete any existing equations.

    2. Press STAT, then select 1:EDIT. Clear existing data in L1 or L2 by moving the cursor to the list header, pressing DEL (or CLEAR), and pressing ENTER / arrowing down.

    3. Enter the lower class limits or representative class values into L1: 40, 50, 60, 70, 80, 90.

    4. Enter the corresponding class frequencies into L2: 1, 1, 4, 12, 7, 5.

    5. Press WINDOW and set the parameters:

      • Xmin = 39.5

      • Xmax = 99.5

      • Xscl = 99.539.56=10\frac{99.5 - 39.5}{6} = 10

      • Ymin = -5

      • Ymax = 15

      • Yscl = 1

      • Xres = 1

    6. Press 2nd Y= (STAT PLOT). Select 4:Plotsoff and press ENTER to turn off all existing plots.

    7. Press 2nd Y=, select 1:Plot1, and press ENTER. Arrow down to TYPE and select the 3rd icon (histogram plot). Press ENTER.

    8. Arrow down to Xlist and specify L1 (2nd 1). Arrow down to Freq and specify L2 (2nd 2).

    9. Press GRAPH. Press the TRACE key and use the left/right arrow keys to examine individual class intervals and bar heights.

  • Frequency Polygons:

    • Definition: A line-graph representation of a frequency distribution derived directly from a histogram.

    • Construction Steps:

    1. Plot a point at the midpoint of the upper edge of each histogram bar.

    2. Remove or erase the histogram bars.

    3. Connect adjacent plotted midpoints with straight line segments.

  • Time Series Graphs:

    • Definition: A line graph where the horizontal axis represents specific progression points in time (e.g., years, months, days).

    • Example — San Francisco 49ers Regular Season Wins (2007–2011):

    • Season 2007: 5 Wins

    • Season 2008: 7 Wins

    • Season 2009: 8 Wins

    • Season 2010: 6 Wins

    • Season 2011: 13 Wins

Measures of Data Location

  • Percentiles:

    • Definition: The kthk^{\text{th}} percentile of a data set is a value such that approximately k%k\% of the data values are less than or equal to it, and approximately (100k)%(100 - k)\% of the data values are greater than or equal to it.

    • Key Property: A percentile does not need to be an actual value contained within the observed data set.

    • Example Interpretation: Being at the 70th70^{\text{th}} percentile means the value is greater than or equal to 70%70\% of all values in the data set, and 70%70\% of the data set lies below or at this value.

  • Quartiles and Five-Number Summary:

    • First Quartile (Q1Q_1): Equivalent to the 25th25^{\text{th}} percentile.

    • Median (Q2Q_2 / Second Quartile): Equivalent to the 50th50^{\text{th}} percentile.

    • Third Quartile (Q3Q_3): Equivalent to the 75th75^{\text{th}} percentile.

    • Minimum: The smallest individual data value in the set.

    • Maximum: The largest individual data value in the set.

    • Five-Number Summary: The ordered set consisting of:     Minimum,Q1,Median,Q3,Maximum\text{Minimum}, Q_1, \text{Median}, Q_3, \text{Maximum}

  • Measures of Spread Derived from Quartiles:

    • Range: The total distance between maximum and minimum values:     range=maxmin\text{range} = \text{max} - \text{min}

    • Interquartile Range (IQR): The range of the middle 50%50\% of the data:     IQR=Q3Q1\text{IQR} = Q_3 - Q_1

  • Procedure for Calculating Median, Q1Q_1, and Q3Q_3:

    1. Sort all data values in ascending order from smallest to largest.

    2. Determine the Median (Q2Q_2):

    • If the total number of data values (nn) is odd, the median is the single middle value.

    • If the total number of data values (nn) is even, the median is the arithmetic average of the two middle values.

    1. Determine the First Quartile (Q1Q_1):

    • Isolate the lower half of the sorted data set (all values located strictly below the median position).

    • Q1Q_1 is the median of this lower half.

    1. Determine the Third Quartile (Q3Q_3):

    • Isolate the upper half of the sorted data set (all values located strictly above the median position).

    • Q3Q_3 is the median of this upper half.

  • Worked Example — 18-Element Data Set:

    • Raw Data Set: 67, 94, 85, 79, 53, 78, 87, 98, 86, 67, 90, 77, 84, 98, 66, 92, 79, 89 (n=18n = 18 values).

    • Sorted Data Set:     53, 66, 67, 67, 77, 78, 79, 79, 84, 85, 86, 87, 89, 90, 92, 94, 98, 98

    • Calculation Steps:

    • Minimum = 53

    • Maximum = 98

    • Median Calculation (n=18n = 18, even):       The 9th value is 84 and the 10th value is 85.       Median=84+852=84.5\text{Median} = \frac{84 + 85}{2} = 84.5

    • First Quartile (Q1Q_1):       Lower half consists of 9 values: 53, 66, 67, 67, 77, 78, 79, 79, 84.       The middle (5th) value is 77.       Q1=77Q_1 = 77

    • Third Quartile (Q3Q_3):       Upper half consists of 9 values: 85, 86, 87, 89, 90, 92, 94, 98, 98.       The middle (5th) value is 90.       Q3=90Q_3 = 90

    • Complete Five-Number Summary: 53, 77, 84.5, 90, 98