Describing Data with Graphs - Chapter Note

Course Overview and Fundamental Definitions

  • Course Information: Introduction to Probability and Statistics, Chapter 1: Describing Data with Graphs (Fall 2026).

  • Instructor: Dr. Reza Pakyari.

  • Textbook Reference: Introduction to Probability and Statistics, 14th edition (2013), authored by Robert J. Beaver, Barbara M. Beaver, and William Mendenhall.

  • Variable: A characteristic that changes or varies over time and/or for different individuals or objects under consideration.

    • Examples of Variables: Hair color, white blood cell count, time to failure of a computer component.

  • Experimental Unit: The individual or object on which a variable is measured.

  • Measurement: The result obtained when a variable is actually measured on an experimental unit.

  • Data: A set of measurements, which can represent either a sample or a population.

  • Concrete Measurement Examples:

    • Example 1:

    • Variable: Hair color

    • Experimental Unit: Person

    • Typical Measurements: Brown, black, blonde, etc.

    • Example 2:

    • Variable: Time until a light bulb burns out

    • Experimental Unit: Light bulb

    • Typical Measurements: 1500hours1500\,\text{hours}, 1535.5hours1535.5\,\text{hours}, etc.

Classification of Data and Variables

  • Classification by Number of Variables Measured per Experimental Unit:

    • Univariate Data: Exactly one variable is measured on a single experimental unit.

    • Bivariate Data: Two variables are measured on a single experimental unit.

    • Multivariate Data: More than two variables are measured on a single experimental unit.

  • Classification by Variable Type:

Variable Types Hierarchy
  • Qualitative Variables (Categorical): Measure a quality or characteristic on each experimental unit.

    • Examples:

    • Hair color (black, brown, blonde, gray, etc.)

    • Make of car (Dodge, Honda, Ford, Kia, etc.)

    • Gender (male, female)

    • Place of birth (Qatar, Egypt, Palestine, etc.)

  • Quantitative Variables: Measure a numerical quantity on each experimental unit.

    • Discrete Quantitative Variable: Can assume only a finite or countable number of values.

    • Example 1: Number of oranges measured for each orange tree in a grove.

    • Example 2: Number of cars entering a college campus on a particular day.

    • Continuous Quantitative Variable: Can assume the infinitely many values corresponding to the points on a line interval.

    • Example: Time until a light bulb burns out.

Graphing Qualitative Data

  • Data Distribution: Describes what values of the variable have been measured and how often each value has occurred.

  • Measures of Frequency ("How Often"):

    • Frequency (ff): The absolute count of observations falling into a specific category.

    • Relative Frequency: Defined as the fraction or proportion of total observations falling into a category:     Relative Frequency=fn\text{Relative Frequency} = \frac{f}{n}     where nn represents the total sample size.

    • Percent: The relative frequency expressed as a percentage of the total count:     Percent=100×Relative Frequency=100×(fn)\text{Percent} = 100 \times \text{Relative Frequency} = 100 \times \left(\frac{f}{n}\right)

  • Worked Example: M&M Candies Distribution (n=25n = 25 total candies):

    • Red: Tally = 111, Frequency = 33, Relative Frequency = 325=0.12\frac{3}{25} = 0.12, Percent = 12%12\%

    • Blue: Tally = 1111 1, Frequency = 66, Relative Frequency = 625=0.24\frac{6}{25} = 0.24, Percent = 24%24\%

    • Green: Tally = 1111, Frequency = 44, Relative Frequency = 425=0.16\frac{4}{25} = 0.16, Percent = 16%16\%

    • Orange: Tally = 1111, Frequency = 55, Relative Frequency = 525=0.20\frac{5}{25} = 0.20, Percent = 20%20\%

    • Brown: Tally = 111, Frequency = 33, Relative Frequency = 325=0.12\frac{3}{25} = 0.12, Percent = 12%12\%

    • Yellow: Tally = 1111, Frequency = 44, Relative Frequency = 425=0.16\frac{4}{25} = 0.16, Percent = 16%16\%

    • Total: Frequency = 2525, Relative Frequency = 2525=1.00\frac{25}{25} = 1.00, Percent = 100%100\%

  • Graphical Displays for Qualitative Variables:

    • Bar Chart: Uses vertical or horizontal bars whose heights or lengths correspond to category frequencies or relative frequencies.

    • Pie Chart: Uses a circle divided into sectors whose areas are proportional to the relative frequency of each category.

    • Pareto Chart: A specialized bar chart in which the categorical bars are arranged in descending order from largest frequency to smallest frequency.

Pareto Chart of M&M Colors

Graphing Quantitative Data across Categories and Over Time

  • Quantitative Variables Across Categories: A single quantitative variable measured across different population segments or categorical classifications can be visually summarized using pie or bar charts.

    • Example (Big Mac Index):

    • Switzerland: Cost = $4.90\$4.90

    • United States (U.S.): Cost = $2.90\$2.90

    • South Africa: Cost = $1.86\$1.86

Cost Comparison of Big Mac
  • Time Series Data: A single quantitative variable measured sequentially over time.

    • Graphical Display: Typically visualized using a line chart or time series plot.

    • Example (Consumer Price Index over Consecutive Months):

    • September: 178.10178.10

    • October: 177.60177.60

    • November: 177.50177.50

    • December: 177.30177.30

    • January: 177.60177.60

    • February: 178.00178.00

    • March: 178.60178.60

Consumer Price Index Time Series

Graphical Displays for Small Quantitative Datasets

  • Dot Plots:

    • The simplest graphical representation for quantitative data.

    • Measurements are plotted as individual points along a horizontal numerical scale.

    • Identical measurements are stacked vertically above one another.

    • Example: For dataset S={4,5,5,6,7}\mathbf{S} = \{4, 5, 5, 6, 7\}, single dots are positioned at 44, 66, and 77, while two vertical dots are stacked at 55

  • Stem and Leaf Plots:

    • Displays the actual numerical values of data while displaying the shape of the distribution.

    • Construction Procedure:

    1. Divide each numerical measurement into two components: the stem (leading digits) and the leaf (trailing digit).

    2. List all possible stem values in a vertical column, separated by a vertical boundary line.

    3. For each data value, place its corresponding leaf digit in the row next to its matching stem.

    4. Sort the leaf digits in ascending order from left to right within each stem row.

    5. Include a explicit key defining the unit scale for stems and leaves.

    • Example 1: Walking Shoe Prices (\) for 18 Brands

    • Raw Dataset: 90,70,70,70,75,70,65,68,60,74,70,95,75,70,68,65,40,6590, 70, 70, 70, 75, 70, 65, 68, 60, 74, 70, 95, 75, 70, 68, 65, 40, 65

    • Stem-and-Leaf Representation (Leaf Unit = 11):

      • Stem 4: 00

      • Stem 5: (no values)

      • Stem 6: 0,5,5,5,8,80, 5, 5, 5, 8, 8

      • Stem 7: 0,0,0,0,0,0,4,5,50, 0, 0, 0, 0, 0, 4, 5, 5

      • Stem 8: (no values)

      • Stem 9: 0,50, 5

    • Example 2: Birth Weights of 18 Full-Term Newborn Babies

    • Raw Dataset: 9.1,4.1,7.0,7.0,7.5,7.0,6.5,6.8,6.0,7.4,7.0,9.5,7.5,7.0,6.8,6.5,4.2,6.59.1, 4.1, 7.0, 7.0, 7.5, 7.0, 6.5, 6.8, 6.0, 7.4, 7.0, 9.5, 7.5, 7.0, 6.8, 6.5, 4.2, 6.5

    • Stem-and-Leaf Representation (Leaf Unit = 0.10.1):

      • Stem 4: 1,21, 2

      • Stem 5: (no values)

      • Stem 6: 0,5,5,5,8,80, 5, 5, 5, 8, 8

      • Stem 7: 0,0,0,0,0,4,5,50, 0, 0, 0, 0, 4, 5, 5

      • Stem 8: (no values)

      • Stem 9: 1,51, 5

Relative Frequency Histograms

  • Definition: A relative frequency histogram for a quantitative dataset is a bar graph where each bar's base spans a continuous class subinterval and its height represents the relative frequency (proportion of total observations) falling within that subinterval.

  • Step-by-Step Construction Procedure:

    1. Divide the total range of data into 55 to 1212 subintervals of equal length.

    2. Calculate the approximate class subinterval width:      Approximate Width=RangeNumber of Subintervals=Maximum ValueMinimum ValueNumber of Subintervals\text{Approximate Width} = \frac{\text{Range}}{\text{Number of Subintervals}} = \frac{\text{Maximum Value} - \text{Minimum Value}}{\text{Number of Subintervals}}

    3. Round the computed width up to a convenient numerical value.

    4. Apply the method of left inclusion: each class interval includes its left boundary endpoint but excludes its right boundary endpoint ([a,b)[a, b)).

    5. Create a frequency table listing classes, tally counts, frequencies, and relative frequencies.

    6. Construct the histogram by placing class intervals on the horizontal axis and relative frequencies on the vertical axis.

  • Statistical Interpretation of Bar Heights:

    • The bar height represents the proportion of overall observations falling into that specific class subinterval.

    • The bar height equals the empirical probability that a single observation selected at random from the dataset will fall into that subinterval.

  • Comprehensive Example: Ages of 50 Tenured Faculty Members

    • Dataset (n=50n = 50): 34,48,70,63,52,52,35,50,37,43,53,43,52,44,42,31,36,48,43,26,58,62,49,34,48,53,39,45,34,59,34,66,40,59,36,41,35,36,62,34,38,28,43,50,30,43,32,44,58,5334, 48, 70, 63, 52, 52, 35, 50, 37, 43, 53, 43, 52, 44, 42, 31, 36, 48, 43, 26, 58, 62, 49, 34, 48, 53, 39, 45, 34, 59, 34, 66, 40, 59, 36, 41, 35, 36, 62, 34, 38, 28, 43, 50, 30, 43, 32, 44, 58, 53

    • Range: Minimum age = 2626, Maximum age = 7070.

    • Interval Setup: 66 classes chosen.

    • Minimum Width Calculation:     Minimum Width=70266=4467.33\text{Minimum Width} = \frac{70 - 26}{6} = \frac{44}{6} \approx 7.33

    • Selected Convenient Class Width: 88

    • Starting Point: 2525

    • Statistical Table:

    • 25 to <3325 \text{ to } < 33: Tally = 1111, Frequency = 55, Relative Frequency = 550=0.10\frac{5}{50} = 0.10 (10%10\%

    • 33 to <4133 \text{ to } < 41: Tally = 1111 1111 1111, Frequency = 1414, Relative Frequency = 1450=0.28\frac{14}{50} = 0.28 (28%28\%

    • 41 to <4941 \text{ to } < 49: Tally = 1111 1111 111, Frequency = 1313, Relative Frequency = 1350=0.26\frac{13}{50} = 0.26 (26%26\%

    • 49 to <5749 \text{ to } < 57: Tally = 1111 1111, Frequency = 99, Relative Frequency = 950=0.18\frac{9}{50} = 0.18 (18%18\%

    • 57 to <6557 \text{ to } < 65: Tally = 1111 11, Frequency = 77, Relative Frequency = 750=0.14\frac{7}{50} = 0.14 (14%14\%

    • 65 to <7365 \text{ to } < 73: Tally = 11, Frequency = 22, Relative Frequency = 250=0.04\frac{2}{50} = 0.04 (4%4\%

Faculty Ages Relative Frequency Histogram
  • Distribution Analysis & Probability Calculations:

    • Distribution Shape: Skewed right (unimodal with a tail extending toward higher ages).

    • Outliers: None present in this distribution.

    • Proportion of Tenured Faculty Younger than 41:     Proportion=5+1450=1950=0.38\text{Proportion} = \frac{5 + 14}{50} = \frac{19}{50} = 0.38

    • Probability of Selecting a Faculty Member Aged 49 or Older:     P(Age49)=9+7+250=1850=0.36P(\text{Age} \ge 49) = \frac{9 + 7 + 2}{50} = \frac{18}{50} = 0.36

Interpreting Graphs: Location, Spread, Shape, and Outliers

  • Location and Spread:

    • Location: Indicates where the central portion of the distribution is centered along the horizontal numerical axis.

    • Spread: Indicates the degree of variability or dispersion of data measurements around the center.

  • Distribution Shapes:

    • Mound-Shaped and Symmetric: The right and left halves of the graph are approximate mirror images of one another.

    • Skewed Right: A distribution with a longer tail extending toward unusually large numerical values.

    • Skewed Left: A distribution with a longer tail extending toward unusually small numerical values.

    • Bimodal: A distribution displaying two distinct local peaks.

  • Outliers:

    • Definition: Measurements that deviate significantly from the general pattern of the remaining data.

    • Data Entry Error Example: A quality control technician measures gear diameters (cm\text{cm}) across 15 production items: 1.991,1.891,1.991,1.988,1.993,1.989,1.990,1.988,1.988,1.993,1.991,1.989,1.989,1.993,1.990,1.9941.991, 1.891, 1.991, 1.988, 1.993, 1.989, 1.990, 1.988, 1.988, 1.993, 1.991, 1.989, 1.989, 1.993, 1.990, 1.994

    • Outlier Identification: The second measurement, 1.8911.891, represents an obvious typing error outlier compared to all other gear diameters centered tightly around 1.990cm1.990\,\text{cm}.

Key Concepts Summary

  • Data Generation Principles:

    • Experimental units, variables, and discrete measurements.

    • Samples versus underlying populations.

    • Univariate, bivariate, and multivariate measurement structures.

  • Types of Variables:

    • Qualitative (Categorical).

    • Quantitative (Discrete vs. Continuous).

  • Graphical Techniques for Univariate Data:

    • Qualitative: Pie charts, Bar charts, Pareto charts.

    • Quantitative: Pie and bar charts (categorical comparison), Line charts (time series), Dot plots, Stem and leaf plots, Relative frequency histograms.

  • Descriptive Characteristics of Distributions:

    • Distribution shapes: symmetric, skewed left, skewed right, unimodal, bimodal.

    • Computing proportions and empirical probabilities within specified subintervals.

    • Detection and evaluation of outliers.