2.1 STEM-AND-LEAF GRAPHS (STEMPLOTS), LINE GRAPHS, AND BAR GRAPHS
Stem-and-leaf graphs (stemplots) overview
Purpose: organize data by splitting each observation into a stem and a leaf.
Leaves are the final significant digit; stems capture the remaining leading digits.
Example conventions:
The observation 23 has stem 2 and leaf 3.
The observation 432 has stem 43 and leaf 2.
Construction steps:
Write stems in a vertical line from smallest to largest.
Draw a vertical line to the right of the stems.
Write leaves in increasing order next to their corresponding stem.
Example task: creating a stem-and-leaf graph for a given dataset (as in the text).
Line graphs
Used to display data that change over an ordered category or time (e.g., frequency of reminders per week).
Data are typically shown in a table (Table 2.7) and a figure (Figure 2.2).
Interpretation focus: trend over time, rise/fall patterns, and comparisons across categories.
Bar graphs
Used to display categorical or grouped numerical data.
Data example (Facebook users by age group, end of 2011):
Age groups and user counts:
13–25: 65,082,280 users; 45%
26–44: 53,300,200 users; 36%
45–64: 27,885,100 users; 19%
Bar height corresponds to the measure shown (counts or proportions).
Construction points: order of categories, equal-width bars, clear axis labels, and a legend if needed.
2.2 HISTOGRAMS, FREQUENCY POLYGONS, AND TIME SERIES GRAPHS
Histograms
A histogram consists of contiguous boxes (bars) with both a horizontal and a vertical axis.
Horizontal axis: represents the data (e.g., a measurement such as distance).
Vertical axis: represents frequency, relative frequency, percent frequency, or probability.
The histogram conveys the shape of the data, the center, and the spread, similarly to a stem-and-leaf plot.
Construction steps:
Decide the number of bars (classes) representing the data.
Frequency polygons
A line graph that connects the midpoints of class intervals (the class marks) to show the distribution shape.
Useful for comparing distributions or showing the overall pattern of a frequency distribution.
Time series graphs
Plot values of a variable across time (e.g., Annual CPI over years).
Axes typically include Time (years) on the horizontal axis and the measured value on the vertical axis.
Interpretation emphasizes trends, cycles, and long-run changes over time.
2.3 MEASURES OF THE LOCATION OF THE DATA
Median
Definition: a number that separates ordered data into halves; the center of the data in the sense of positional rank.
Properties:
For ordered data, half of the values are ≤ the median and half are ≥ the median.
The median need not be an observed value.
Example (data):
Data: 1, 11.5, 6, 7.2, 4, 8, 9, 10, 6.8, 8.3, 2, 2, 10, 1
Ordered: 1, 1, 2, 2, 4, 6, 6.8, 7.2, 8, 8.3, 9, 10, 10, 11.5
With n = 14 (even), the median is between the 7th and 8th values:
Quartiles (Q1, Q2, Q3)
Q2 is the median (second quartile).
Q1: the median of the lower half of the data.
Q3: the median of the upper half of the data.
IQR (interquartile range):
IQR use and outliers:
The IQR helps assess the spread of the middle 50% of data.
Potential outliers are values below or above .
Percentiles
A percentile indicates the relative standing of a data value when data are ordered.
Interpretation: e.g., the 15th percentile means 15% of data values are ≤ that value.
Two percentile concepts in the text:
Percentile formula for kth percentile (i.e., position-based):
Ordering data from smallest to largest, let
If is a positive integer, the kth percentile is the data value at position .
If is not an integer, round up and down to the nearest integers and average the two data values at those positions.
Percentile of a given value within a data set (rank-based percentile):
Let be the position counting from the bottom to the data value of interest, the number of data values equal to that value, and the total number of data.
Compute the percentile as:
Outliers (context)
Outliers are data points that stand away from the main body of the data; they may be errors or important insights.
2.4 BOX PLOTS
Box plots (box-and-whisker plots) provide a graphical summary of data concentration and spread.
Five-number summary used: minimum, Q1, median, Q3, maximum.
Construction idea:
A box spans from Q1 to Q3 with a line inside for the median.
Whiskers extend to the minimum and maximum data values (unless outliers are defined separately).
Example application: heights of 40 students (data listed in the text) illustrating how a box plot summarizes spread and center visually.
2.5 MEASURES OF THE CENTER OF THE DATA
The center of a data set is described by measures of location, primarily:
Mean (average):
For a sample of n values,
For a population, the mean is ar{x} ext{ or } ext{often }oldsymbol{bc} depending on notation, but the key idea is the same: the arithmetic average.
Median: discussed in 2.3.
When to prefer mean vs. median
The median is generally more robust to extreme values and outliers; hence it is a better measure of center for skewed data.
The mean is the most common measure of center and is used in many statistical formulas and inferential procedures.
Mode
The most frequent value(s) in the data.
There can be more than one mode (multimodal). A data set with two modes is called bimodal.
The mode can be used for qualitative data as well (e.g., color categories).
Grouped data: mean of grouped frequency tables
When only grouped data are available, the exact mean cannot be computed; estimate via the mean of the frequency table.
Key quantities:
= frequency of the interval
= midpoint of the interval
2.6 SKEWNESS AND THE MEAN, MEDIAN, AND MODE
Symmetrical distribution
A distribution is symmetrical if there exists a vertical line around which the left and right sides are mirror images.
In a perfectly symmetrical distribution, mean = median = mode.
Skewed distributions
Left-skewed (skewed to the left): tail on the left; mean < median < mode.
Right-skewed (skewed to the right): tail on the right; mode < median < mean.
Illustrative examples from the text
Example 1 (left-skew): data 4,5,6,6,6,7,7,7,7,8; mean = 6.3, median = 6.5, mode = 7. Interpretation: mean is pulled toward the tail more than the median or mode.
Example 2 (right-skew): data 6,7,7,7,7,8,8,8,9,10; mean = 7.7, median = 7.5, mode = 7. Interpretation: the mean is the largest of the three, reflecting the right tail.
Practical takeaway
When a distribution is skewed, the mean is more sensitive to outliers and the tail than the median or mode.
2.7 MEASURES OF THE SPREAD OF THE DATA
Spread (variability) is a key characteristic of a data set.
Standard deviation (the most common measure of spread)
Purpose: quantify how far data values are from their mean.
Interpretation: smaller SD indicates data are tightly clustered around the mean; larger SD indicates more dispersion.
The standard deviation is always nonnegative (positive or zero).
The SD is central to many statistical methods and helps assess whether a particular value is close to or far from the mean.
Formulas (conceptual; see below for explicit forms):
For a sample: $$s = \