Comprehensive Study Notes on Describing & Comparing Data
Measures of Center
Measures of center are statistics that identify a single value as representative of an entire distribution.
Definitions of Central Tendency
Mean ( for sample, for population): The arithmetic average calculated by summing all data values and dividing by the total count of values.
Median (\tilde{x} or ): The middle data value when all observations are arranged in order of magnitude (ascending or descending).
Mode: The data value or values that appear with the greatest frequency.
Midrange: The measure of center that is the value midway between the maximum and minimum values in the dataset.
Formulas for Calculating Measures of Center
Sample Mean:
Median Location:
If sample size is odd, the median is the single value located at position .
If sample size is even, the median is the arithmetic mean of the two middle values located at positions and .
Midrange:
Rounding Rule for Measures of Center
Round calculated measures of center to one more decimal place than present in the raw data values.
Examples and Calculations for Raw Data
Example 1: Position of Median
If data values, the median is located at position:
If data values, the median is located at position:
(average of the 10th and 11th data values)
Example 2: Data Set Analysis
Dataset ():
Step 1: Sort data in ascending order
Mean:
Median: Position is . The 5th value is and the 6th value is .
Mode: The values and both occur twice. The dataset is bimodal with modes and .
Midrange:
Calculating Mean from Frequency Distributions
When data are summarized in a frequency distribution, individual raw values are lost. The class midpoint represents all values within each class.
Formula for Frequency Distribution Mean:
Class Midpoint Formula:
Example: Exam Scores Distribution ( Students)

Class : Frequency , Midpoint ,
Class : Frequency , Midpoint ,
Class : Frequency , Midpoint ,
Class : Frequency , Midpoint ,
Class : Frequency , Midpoint ,
Sum of Frequencies
Sum of Products
Calculated Mean:
Weighted Mean (Semester Grade Point Average - GPA)
A weighted mean is used when individual values carry varying degrees of importance or weight.
Formula:
Grade Point Values: , , , , .

Example Calculation:
Physics: Grade B (Value ), Weight ,
Foreign Language: Grade A (Value ), Weight ,
English: Grade C (Value ), Weight ,
History: Grade B (Value ), Weight ,
Sum of Weights
Sum of Products
Semester GPA:
Measures of Variation
Measures of variation quantify the degree of spread or dispersion among data values.
Definitions of Variation Statistics
Range: The difference between the highest data value and the lowest data value.
Sample Standard Deviation (): A measure of variation of values relative to the sample mean.
Sample Variance (): The square of the sample standard deviation.
Population Standard Deviation (): A measure of variation of all values relative to the population mean.
Population Variance (): The square of the population standard deviation.
Summary of Notations
Parameter / Statistic | Sample Notation | Population Notation |
|---|---|---|
Mean | ||
Standard Deviation | ||
Variance | ||
Size / Total Count |
Formulas for Measures of Variation
Range:
Sample Standard Deviation:
Sample Standard Deviation (Shortcut Computation Formula):
Sample Variance:
Rounding Rule for Measures of Variation
Round values of variation to one more decimal place than present in the original data.
Step-by-Step Examples of Variation Calculations
Example 1: Bank Customer Waiting Times
Sample of 10 waiting times (minutes):
Range:
Mean:
Variance Calculation Table:
Sample Variance:
Sample Standard Deviation:
Example 2: Female Pulse Rates Frequency Table ()

Class : , Midpoint
Class : , Midpoint
Class : , Midpoint
Class : , Midpoint
Class : , Midpoint
Class : , Midpoint
Class : , Midpoint
Total Sample Size:
Sum of Products
Sum of Squared Products
Standard Deviation Calculation:
Relative Spread, Rules of Thumb, and Empirical Distribution
Coefficient of Variation (CV)
The coefficient of variation describes the standard deviation relative to the mean, allowing direct comparison of variation between datasets with different units or substantially different means.
Sample Formula:
Population Formula:
Example Comparison:
English Final Exam: Mean = , Standard Deviation =
History Final Exam: Mean = , Standard Deviation =
Comparison: The History final scores exhibit greater relative variation than the English final scores ().
Range Rule of Thumb
Estimating Standard Deviation:
Example Estimation (Bank wait times between and minutes):
Identifying Usual vs. Unusual Values:
Minimum Usual Value =
Maximum Usual Value =
Values inside are considered usual. Values outside this range are considered unusual.
The Empirical Rule (68–95–99.7 Rule)
Applies strictly to distributions that are bell-shaped (symmetric and normal).
Approximately of all data values fall within standard deviation of the mean ().
Approximately of all data values fall within standard deviations of the mean ().
Approximately of all data values fall within standard deviations of the mean ().
Example Calculation (Generator Voltage):
Generator output mean , standard deviation .
(a) Percentage between and :
This interval corresponds to . By the Empirical Rule, approximately 95% of voltage amounts fall in this range.
(b) Percentage between and :
This interval corresponds to . By the Empirical Rule, approximately 99.7% of voltage amounts fall in this range.
Interpretation: Virtually all (about ) generator voltage outputs lie between and , with lying within and .
Measures of Relative Standing & Boxplots
Measures of relative standing indicate the location of a value relative to other values within a data set.
z-Scores (Standardized Values)
A z-score represents the number of standard deviations a value is located above or below the mean.
Sample z-score Formula:
Population z-score Formula:
Usual vs. Unusual z-Scores:
Usual values:
Unusual values: or
Earthquake Example:
Magnitude dataset (): Mean , Standard Deviation s = 0.587$.\n * Convert magnitude x = 1.766 to a z-score:\n\nz = \frac{1.766 - 1.184}{0.587} = \frac{0.582}{0.587} \approx 0.99\n\n * Conclusion: Because z = 0.99-221.766 is considered **usual**.\n\n## Percentiles and Quartiles\n\n* **Percentile**: Measures of location that divide a set of ordered data into 100 equal groups.\n* **Interpretation Example**:\n * Statement: "A child is in the 75th percentile for height."\n * Meaning: The child is taller than 75\%25\% of them.\n\n### Percentile Formulas and Rules\n\n* **Finding Percentile of a Given Data Value x**:\n\n\text{Percentile of } x = \left( \frac{\text{Number of values less than } x}{\text{Total number of values } n} \right) \times 100\n\n* **Finding Data Value Corresponding to Percentile kP_k)**:\n * Compute Locator L:\n\nL = \left( \frac{k}{100} \right) \cdot n\n\n * **Decision Rule for Locator L**:\n 1. If LLP_kL\text{th} value in the ordered dataset.\n 2. If LP_kL\text{th}(L+1)\text{th} value in the ordered dataset.\n\n### Step-by-Step Dataset Examples (n = 16)\n\nSorted Dataset: 2, 18, 19, 22, 24, 26, 26, 35, 35, 35, 36, 38, 40, 46, 48, 65\n\n* **(a) Find the 32nd Percentile (P_{32})**:\n\nL = \left( \frac{32}{100} \right) \times 16 = 5.12\n\n * Round UP to 6P_{32} is the 6th value in the sorted list.\n * 6th value = 26.\n\n* **(b) Find the 10th Percentile (P_{10})**:\n\nL = \left( \frac{10}{100} \right) \times 16 = 1.6\n\n * Round UP to 2P_{10} is the 2nd value in the sorted list.\n * 2nd value = 18.\n\n* **(c) Find Median / 50th Percentile (P_{50})**:\n\nL = \left( \frac{50}{100} \right) \times 16 = 8\n\n * Since 8 is a whole number, average the 8th and 9th values.\n * 8th value = 3535$.
(d) Find 3rd Quartile / 75th Percentile ():
Since is a whole number, average the 12th and 13th values.
12th value = , 13th value = 40$.\n\nQ_3 = \frac{38 + 40}{2} = 39\n\n* **(e) Find Percentile for Value 22**:\n * Count values strictly less than 222, 18, 19 (3 values).\n\n\text{Percentile} = \left( \frac{3}{16} \right) \times 100 = 18.75\% \approx 19\text{th percentile}\n\n * Value 22P_{19}.\n\n## 5-Number Summary, Boxplots, and Outlier Calculations\n\n### The 5-Number Summary\n\nThe 5-number summary consists of five values that summarize a dataset:\n1. **Minimum Value**\n2. **First Quartile (Q_1 = P_{25})**\n3. **Median (Q_2 = P_{50})**\n4. **Third Quartile (Q_3 = P_{75})**\n5. **Maximum Value**\n\n* **Calculations for the Example Dataset**:\n * Minimum = 2\n * Q_1 = P_{25}L = \left(\frac{25}{100}\right) \times 16 = 4\frac{22 + 24}{2} = 23)\n * Median (Q_235\n * Q_3 = 39\n * Maximum = 65\n * **5-Number Summary**: `[2, 23, 35, 39, 65]`\n\n### Interquartile Range (IQR) and Outlier Boundary Identification\n\n* **Interquartile Range Formula**:\n\nIQR = Q_3 - Q_1\n\n * For example data: IQR = 39 - 23 = 16.\n\n* **Outlier Boundaries**:\n * **Lower Boundary**:\n\n\text{Lower Boundary} = Q_1 - 1.5 \times IQR\n\n * **Upper Boundary**:\n\n\text{Upper Boundary} = Q_3 + 1.5 \times IQR\n\n* **Example Outlier Analysis for Values 265**:\n * Lower Boundary:\n\n23 - 1.5(16) = 23 - 24 = -1\n\n * Upper Boundary:\n\n39 + 1.5(16) = 39 + 24 = 63\n\n * **Check value 22 \ge -12 lies within the boundary and is **not an outlier**.\n * **Check value 6565 > 6365 exceeds the upper boundary and **is an outlier**.\n\n### Boxplots\nA boxplot (or box-and-whisker diagram) is a graphical plot of a dataset that displays:\n* A box drawn from Q_1Q_3$.
A vertical line drawn inside the box at the Median ().
Line segments ("whiskers") extending from the box out to the minimum and maximum data values (or to the furthest non-outlier values in a modified boxplot).