Describing, Exploring, and Comparing Data: Measures of Center and Variation
Measures of Center
Definition of Measure of Center: A measure of center is a value at the center or middle of a data set. It provides a representative value to summarize the entire data set.
The Mean
Definition: The mean (or arithmetic mean) of a set of data is the measure of center found by adding all of the data values and dividing the total by the number of data values. It is the most important numerical measurement used to describe data and corresponds to what is commonly called an average.
Important Properties of the Mean:
Sample means drawn from the same population tend to vary less than other measures of center.
The mean of a data set uses every single data value.
Disadvantage (Sensitivity to Outliers): A single extreme value (outlier) can substantially change the value of the mean. Therefore, the mean is not resistant.
Definition of Resistance: A statistic is resistant if the presence of extreme values (outliers) does not cause it to change very much.
Notation and Formulas for the Mean:
Σ denotes the sum of a set of data values.
x is the variable used to represent individual data values.
n represents the total number of data values in a sample.
N represents the total number of data values in a population.
English letters represent sample statistics; Greek letters represent population parameters.
Sample Mean Formula (Formula 3-1):xˉ=nΣx=number of data valuessum of all data values
(Pronounced "x-bar").
Population Mean Formula:μ=NΣx
(Denoted by lower-case Greek letter mu, μ).
Worked Example 1: Calculating the Mean
Problem: Find the mean of the first five male pulse rates from a body data set: 84, 74, 50, 60, 52 (all in beats per minute, or BPM).
Definition: The median of a data set is the measure of center that is the middle value when the original data values are arranged in order of increasing (or decreasing) magnitude. Approximately half of the data values are less than the median and half are greater than the median.
Important Properties of the Median:
The median does not change significantly when extreme values are included; therefore, the median is a resistant measure of center.
The median does not directly use every data value (e.g., changing the largest value to a much higher value leaves the median unchanged).
Notation:
Sample median is denoted by x~ (pronounced "x-tilde"), M, or Med.
There is no universally accepted notation or special symbol for a population median.
Procedure for Calculating the Median:
Sort the data values in ascending order.
Odd number of values (n is odd): The median is the exact middle number in the sorted list.
Even number of values (n is even): The median is the mean of the two middle numbers in the sorted list.
Worked Example 2: Median with an Odd Number of Data Values
Problem: Find the median of five male pulse rates: 84, 74, 50, 60, 52\text{ BPM}.
Step 1: Arrange in ascending order: 50,52,60,74,84.
Step 2: Count n=5 (odd). The middle (3rd) value is 60.0BPM.
Result: Median = 60.0BPM (Note: this differs from the mean of 64.0BPM).
Worked Example 3: Median with an Even Number of Data Values
Problem: Find the median after adding a sixth pulse rate (62BPM) to the dataset: 84, 74, 50, 60, 52, 62\text{ BPM}.
Step 1: Arrange in ascending order: 50,52,60,62,74,84.
Step 2: Count n=6 (even). Locate the two middle numbers: 60 and 62
Step 3: Calculate the mean of the two middle numbers:
Median=260+62=2122=61.0BPM
Result: Median = 61.0BPM.
The Mode
Definition: The mode of a data set is the value(s) that occurs with the greatest frequency.
Important Properties of the Mode:
The mode can be computed for both quantitative data and qualitative (categorical) data consisting of names, labels, or categories.
It is the only measure of center that can be used with qualitative data.
A data set can have no mode, one mode, or multiple modes.
Classifications of Mode:
Single Mode: One value occurs most frequently.
Bimodal: Two values occur with the same maximum frequency; each value is a mode.
Multimodal: More than two values occur with the same maximum frequency; each value is a mode.
No Mode: No data value is repeated.
Worked Examples: Mode
Example 4a (Single Mode): Dataset 58,58,58,58,60,60,62,64
Mode = 58BPM (occurs most frequently, 4 times).
Example 4b (Two Modes / Bimodal): Dataset 58,58,58,60,60,60,62,64
Modes = 58BPM and 60BPM (both occur 3 times).
Example 4c (No Mode): Dataset 58,60,64,68,72
Mode = No mode (no value is repeated).
Calculating the Mean from a Frequency Distribution
When data are summarized in a frequency table, the original raw values are unknown. Therefore, the calculated mean is an approximation.
Formula 3-2 (Mean from a Frequency Distribution):xˉ=ΣfΣ(f⋅x)
Where:
f = frequency of each class.
x = class midpoint of each class.
Σf=n = sum of frequencies (total sample size).
Σ(f⋅x) = sum of the products of each class frequency and midpoint.
Worked Example: Mean from Frequency Distribution
Pulse Rate (BPM)
Frequency (f)
Class Midpoint (x)
f⋅x
40–54
15
47
15×47=705
55–69
63
62
63×62=3906
70–84
62
77
62×77=4774
85–99
11
92
11×92=1012
100–114
2
107
2×107=214
Totals
Σf=153
Σ(f⋅x)=10,611
Calculation:xˉ=ΣfΣ(f⋅x)=15310,611=69.4BPM
Comparison: The result 69.4BPM is an approximation. The actual mean calculated using all raw individual male pulse rates is 69.6BPM.
Calculating a Weighted Mean
When data values (x) are assigned different weights (w), a weighted mean is calculated.
Formula 3-3 (Weighted Mean):xˉ=ΣwΣ(w⋅x)
Procedure: Multiply each weight w by its corresponding value x, sum these products, and divide by the total sum of the weights Σw
Worked Example: Grade-Point Average (GPA)
Problem: Compute the first-semester Grade-Point Average (GPA) for a student taking five courses:
Course 1: Grade A (3 credits)
Course 2: Grade A (4 credits)
Course 3: Grade B (3 credits)
Course 4: Grade C (3 credits)
Course 5: Grade F (1 credit)
Quality Points Assignment:A=4, B=3, C=2, D=1, F=0
Weights (w): Credits = 3,4,3,3,1; Total Σw=3+4+3+3+1=14
Substitution into Formula 3-5:s=5(5−1)5(21,336)−(320)2=5(4)106,680−102,400=204280=214=14.6BPM
Variance
Definition: The variance of a set of values is a measure of variation equal to the square of the standard deviation.
Symbols and Formulas:
Sample Variance (s2): Square of the sample standard deviation ss2=n−1Σ(x−xˉ)2
Population Variance (σ2): Square of the population standard deviation σσ2=NΣ(x−μ)2
Important Properties of Variance:
Units: Units of variance are the squares of the original units (e.g., if original data are in feet, variance units are ft2; if seconds, sec2).
Increases dramatically with inclusion of outliers (not resistant).
Value is never negative; equals zero only when all data values are identical.
Estimation Bias: The sample variance s2 is an unbiased estimator of the population variance σ2 (it targets the true population value). In contrast, sample standard deviation s is a biased estimator of σ
Coefficient of Variation (CV)
Definition: The coefficient of variation (or CV) for a set of non-negative sample or population data, expressed as a percentage, describes the standard deviation relative to the mean.
Formulas:
Sample Coefficient of Variation:CV=xˉs⋅100%
Population Coefficient of Variation:CV=μσ⋅100%
When to Use CV:
To compare variation from two samples/populations with different units or scales (e.g., comparing pulse rates in BPM to heights in cm).
To compare variation between groups with vastly different means.
Note: Direct comparison of standard deviations is appropriate only when sample means are approximately equal and measurement scales/units are identical.
Worked Example: Comparing Variation (Pulse Rates vs. Heights)
Problem: Compare variation between 153 male pulse rates (xˉ=69.6BPM, s=11.3BPM) and their heights (xˉ=174.12cm, s=7.10cm).
Male Pulse Rates CV:CV=69.6BPM11.3BPM⋅100%=16.2%
Male Heights CV:CV=174.12cm7.10cm⋅100%=4.1%
Conclusion: Male pulse rates (CV=16.2%) exhibit significantly greater variation than male heights (CV=4.1%).