Comprehensive Study Notes on Describing & Comparing Data

Measures of Center

Measures of center are statistics that identify a single value as representative of an entire distribution.

Definitions of Central Tendency

  • Mean (xˉ\bar{x} for sample, μ\mu for population): The arithmetic average calculated by summing all data values and dividing by the total count of values.

  • Median (\tilde{x} or MM): The middle data value when all observations are arranged in order of magnitude (ascending or descending).

  • Mode: The data value or values that appear with the greatest frequency.

  • Midrange: The measure of center that is the value midway between the maximum and minimum values in the dataset.

Formulas for Calculating Measures of Center

  • Sample Mean:

xˉ=∑xn\bar{x} = \frac{\sum x}{n}

  • Median Location:

Position of Median=n+12\text{Position of Median} = \frac{n + 1}{2}

  • If sample size nn is odd, the median is the single value located at position n+12\frac{n+1}{2}.

  • If sample size nn is even, the median is the arithmetic mean of the two middle values located at positions n2\frac{n}{2} and n2+1\frac{n}{2} + 1.

    • Midrange:

Midrange=Maximum value+Minimum value2\text{Midrange} = \frac{\text{Maximum value} + \text{Minimum value}}{2}

Rounding Rule for Measures of Center

  • Round calculated measures of center to one more decimal place than present in the raw data values.

Examples and Calculations for Raw Data

  • Example 1: Position of Median

    • If n=15n = 15 data values, the median is located at position:

15+12=8th data value\frac{15 + 1}{2} = 8\text{th data value}

  • If n=20n = 20 data values, the median is located at position:

20+12=10.5th position\frac{20 + 1}{2} = 10.5\text{th position}

    (average of the 10th and 11th data values)

  • Example 2: Data Set Analysis

    • Dataset (n=10n = 10): 17,12,13,11,14,18,17,18,16,1517, 12, 13, 11, 14, 18, 17, 18, 16, 15

    • Step 1: Sort data in ascending order

11,12,13,14,15,16,17,17,18,1811, 12, 13, 14, 15, 16, 17, 17, 18, 18

  • Mean:

∑x=11+12+13+14+15+16+17+17+18+18=141\sum x = 11 + 12 + 13 + 14 + 15 + 16 + 17 + 17 + 18 + 18 = 141

xˉ=14110=14.1\bar{x} = \frac{141}{10} = 14.1

  • Median:     Position is 10+12=5.5th\frac{10+1}{2} = 5.5\text{th}. The 5th value is 1515 and the 6th value is 1616.

Median=15+162=15.5\text{Median} = \frac{15 + 16}{2} = 15.5

  • Mode:     The values 1717 and 1818 both occur twice. The dataset is bimodal with modes 1717 and 1818.

  • Midrange:

Midrange=18+112=292=14.5\text{Midrange} = \frac{18 + 11}{2} = \frac{29}{2} = 14.5

Calculating Mean from Frequency Distributions

When data are summarized in a frequency distribution, individual raw values are lost. The class midpoint xmx_m represents all values within each class.

  • Formula for Frequency Distribution Mean:

xˉ=∑(f⋅xm)∑f\bar{x} = \frac{\sum (f \cdot x_m)}{\sum f}

  • Class Midpoint Formula:

xm=Lower Class Limit+Upper Class Limit2x_m = \frac{\text{Lower Class Limit} + \text{Upper Class Limit}}{2}

  • Example: Exam Scores Distribution (n=108n = 108 Students)

Exam score frequency distribution table showing class limits, frequency, midpoint, and product of midpoint and frequency
  • Class 90–9890 – 98: Frequency f=6f = 6, Midpoint xm=90+982=94x_m = \frac{90+98}{2} = 94, f⋅xm=6×94=564f \cdot x_m = 6 \times 94 = 564

  • Class 99–10799 – 107: Frequency f=22f = 22, Midpoint xm=99+1072=103x_m = \frac{99+107}{2} = 103, f⋅xm=22×103=2266f \cdot x_m = 22 \times 103 = 2266

  • Class 108–116108 – 116: Frequency f=43f = 43, Midpoint xm=108+1162=112x_m = \frac{108+116}{2} = 112, f⋅xm=43×112=4816f \cdot x_m = 43 \times 112 = 4816

  • Class 117–125117 – 125: Frequency f=28f = 28, Midpoint xm=117+1252=121x_m = \frac{117+125}{2} = 121, f⋅xm=28×121=3388f \cdot x_m = 28 \times 121 = 3388

  • Class 126–134126 – 134: Frequency f=9f = 9, Midpoint xm=126+1342=130x_m = \frac{126+134}{2} = 130, f⋅xm=9×130=1170f \cdot x_m = 9 \times 130 = 1170

  • Sum of Frequencies ∑f=6+22+43+28+9=108\sum f = 6 + 22 + 43 + 28 + 9 = 108

  • Sum of Products ∑(f⋅xm)=564+2266+4816+3388+1170=12204\sum (f \cdot x_m) = 564 + 2266 + 4816 + 3388 + 1170 = 12204

  • Calculated Mean:

xˉ=12204108=113.0\bar{x} = \frac{12204}{108} = 113.0

Weighted Mean (Semester Grade Point Average - GPA)

A weighted mean is used when individual values carry varying degrees of importance or weight.

  • Formula:

Weighted Mean=∑(x⋅w)∑w\text{Weighted Mean} = \frac{\sum (x \cdot w)}{\sum w}

  • Grade Point Values: A=4A = 4, B=3B = 3, C=2C = 2, D=1D = 1, F=0F = 0.

Weighted mean calculation table for course grade values and course credit hour weights
  • Example Calculation:

    • Physics: Grade B (Value x=3x = 3), Weight w=5 credit hoursw = 5\,\text{credit hours}, x⋅w=3×5=15x \cdot w = 3 \times 5 = 15

    • Foreign Language: Grade A (Value x=4x = 4), Weight w=4 credit hoursw = 4\,\text{credit hours}, x⋅w=4×4=16x \cdot w = 4 \times 4 = 16

    • English: Grade C (Value x=2x = 2), Weight w=3 credit hoursw = 3\,\text{credit hours}, x⋅w=2×3=6x \cdot w = 2 \times 3 = 6

    • History: Grade B (Value x=3x = 3), Weight w=3 credit hoursw = 3\,\text{credit hours}, x⋅w=3×3=9x \cdot w = 3 \times 3 = 9

    • Sum of Weights ∑w=5+4+3+3=15\sum w = 5 + 4 + 3 + 3 = 15

    • Sum of Products ∑(x⋅w)=15+16+6+9=46\sum (x \cdot w) = 15 + 16 + 6 + 9 = 46

    • Semester GPA:

GPA=4615≈3.07\text{GPA} = \frac{46}{15} \approx 3.07

Measures of Variation

Measures of variation quantify the degree of spread or dispersion among data values.

Definitions of Variation Statistics

  • Range: The difference between the highest data value and the lowest data value.

  • Sample Standard Deviation (ss): A measure of variation of values relative to the sample mean.

  • Sample Variance (s2s^2): The square of the sample standard deviation.

  • Population Standard Deviation (σ\sigma): A measure of variation of all values relative to the population mean.

  • Population Variance (σ2\sigma^2): The square of the population standard deviation.

Summary of Notations

Parameter / Statistic

Sample Notation

Population Notation

Mean

xˉ\bar{x}

μ\mu

Standard Deviation

ss

σ\sigma

Variance

s2s^2

σ2\sigma^2

Size / Total Count

nn

NN

Formulas for Measures of Variation

  • Range:

Range=Maximum value−Minimum value\text{Range} = \text{Maximum value} - \text{Minimum value}

  • Sample Standard Deviation:

s=∑(x−xˉ)2n−1s = \sqrt{\frac{\sum (x - \bar{x})^2}{n - 1}}

  • Sample Standard Deviation (Shortcut Computation Formula):

s=n∑x2−(∑x)2n(n−1)s = \sqrt{\frac{n \sum x^2 - (\sum x)^2}{n(n - 1)}}

  • Sample Variance:

s2=∑(x−xˉ)2n−1s^2 = \frac{\sum (x - \bar{x})^2}{n - 1}

Rounding Rule for Measures of Variation

  • Round values of variation to one more decimal place than present in the original data.

Step-by-Step Examples of Variation Calculations

  • Example 1: Bank Customer Waiting Times

    • Sample of 10 waiting times (minutes): 8.5,5.4,9.3,6.2,6.7,7.7,4.2,7.7,5.8,10.08.5, 5.4, 9.3, 6.2, 6.7, 7.7, 4.2, 7.7, 5.8, 10.0

    • Range:

Range=10.0−4.2=5.8 minutes\text{Range} = 10.0 - 4.2 = 5.8\,\text{minutes}

  • Mean:

xˉ=8.5+5.4+9.3+6.2+6.7+7.7+4.2+7.7+5.8+10.010=7.15 minutes\bar{x} = \frac{8.5 + 5.4 + 9.3 + 6.2 + 6.7 + 7.7 + 4.2 + 7.7 + 5.8 + 10.0}{10} = 7.15\,\text{minutes}

  • Variance Calculation Table:

(8.5−7.15)2=1.8225(8.5 - 7.15)^2 = 1.8225

(5.4−7.15)2=3.0625(5.4 - 7.15)^2 = 3.0625

(9.3−7.15)2=4.6225(9.3 - 7.15)^2 = 4.6225

(6.2−7.15)2=0.9025(6.2 - 7.15)^2 = 0.9025

(6.7−7.15)2=0.2025(6.7 - 7.15)^2 = 0.2025

(7.7−7.15)2=0.3025(7.7 - 7.15)^2 = 0.3025

(4.2−7.15)2=8.7025(4.2 - 7.15)^2 = 8.7025

(7.7−7.15)2=0.3025(7.7 - 7.15)^2 = 0.3025

(5.8−7.15)2=1.8225(5.8 - 7.15)^2 = 1.8225

(10.0−7.15)2=8.1225(10.0 - 7.15)^2 = 8.1225

∑(x−xˉ)2=29.865\sum (x - \bar{x})^2 = 29.865

  • Sample Variance:

s2=29.86510−1=29.8659≈3.32 minutes2s^2 = \frac{29.865}{10 - 1} = \frac{29.865}{9} \approx 3.32\,\text{minutes}^2

  • Sample Standard Deviation:

s=3.3183≈1.82 minutess = \sqrt{3.3183} \approx 1.82\,\text{minutes}

  • Example 2: Female Pulse Rates Frequency Table (n=40n = 40)

Frequency distribution table for female pulse rates
  • Class 60–6960 – 69: f=12f = 12, Midpoint xm=64.5x_m = 64.5

  • Class 70–7970 – 79: f=14f = 14, Midpoint xm=74.5x_m = 74.5

  • Class 80–8980 – 89: f=11f = 11, Midpoint xm=84.5x_m = 84.5

  • Class 90–9990 – 99: f=1f = 1, Midpoint xm=94.5x_m = 94.5

  • Class 100–109100 – 109: f=1f = 1, Midpoint xm=104.5x_m = 104.5

  • Class 110–119110 – 119: f=0f = 0, Midpoint xm=114.5x_m = 114.5

  • Class 120–129120 – 129: f=1f = 1, Midpoint xm=124.5x_m = 124.5

  • Total Sample Size: n=∑f=40n = \sum f = 40

  • Sum of Products ∑(f⋅xm)=3070\sum (f \cdot x_m) = 3070

  • Sum of Squared Products ∑(f⋅xm2)=241520\sum (f \cdot x_m^2) = 241520

  • Standard Deviation Calculation:

s=40(241520)−(3070)240(39)=9660800−94249001560=2359001560≈12.3 beats per minutes = \sqrt{\frac{40(241520) - (3070)^2}{40(39)}} = \sqrt{\frac{9660800 - 9424900}{1560}} = \sqrt{\frac{235900}{1560}} \approx 12.3\,\text{beats per minute}

Relative Spread, Rules of Thumb, and Empirical Distribution

Coefficient of Variation (CV)

The coefficient of variation describes the standard deviation relative to the mean, allowing direct comparison of variation between datasets with different units or substantially different means.

  • Sample Formula:

CV=sxˉ×100%CV = \frac{s}{\bar{x}} \times 100\%

  • Population Formula:

CV=σμ×100%CV = \frac{\sigma}{\mu} \times 100\%

  • Example Comparison:

    • English Final Exam: Mean = 8585, Standard Deviation = 55

CVEnglish=585×100%≈5.88%CV_{\text{English}} = \frac{5}{85} \times 100\% \approx 5.88\%

  • History Final Exam: Mean = 110110, Standard Deviation = 88

CVHistory=8110×100%≈7.27%CV_{\text{History}} = \frac{8}{110} \times 100\% \approx 7.27\%

  • Comparison: The History final scores exhibit greater relative variation than the English final scores (7.27%>5.88%7.27\% > 5.88\%).

Range Rule of Thumb

  • Estimating Standard Deviation:

s≈Range4s \approx \frac{\text{Range}}{4}

  • Example Estimation (Bank wait times between 4.24.2 and 10.010.0 minutes):

Range=10.0−4.2=5.8 minutes\text{Range} = 10.0 - 4.2 = 5.8\,\text{minutes}

s≈5.84=1.45 minutess \approx \frac{5.8}{4} = 1.45\,\text{minutes}

  • Identifying Usual vs. Unusual Values:

    • Minimum Usual Value = xˉ−2s\bar{x} - 2s

    • Maximum Usual Value = xˉ+2s\bar{x} + 2s

    • Values inside [xˉ−2s,xˉ+2s][\bar{x} - 2s, \bar{x} + 2s] are considered usual. Values outside this range are considered unusual.

The Empirical Rule (68–95–99.7 Rule)

Applies strictly to distributions that are bell-shaped (symmetric and normal).

  • Approximately 68%68\% of all data values fall within 11 standard deviation of the mean (μ±1σ\mu \pm 1\sigma).

  • Approximately 95%95\% of all data values fall within 22 standard deviations of the mean (μ±2σ\mu \pm 2\sigma).

  • Approximately 99.7%99.7\% of all data values fall within 33 standard deviations of the mean (μ±3σ\mu \pm 3\sigma).

  • Example Calculation (Generator Voltage):

    • Generator output mean μ=125.0 volts\mu = 125.0\,\text{volts}, standard deviation σ=0.3 volts\sigma = 0.3\,\text{volts}.

    • (a) Percentage between 124.4 volts124.4\,\text{volts} and 125.6 volts125.6\,\text{volts}:

125.0−2(0.3)=124.4 volts125.0 - 2(0.3) = 124.4\,\text{volts}

125.0+2(0.3)=125.6 volts125.0 + 2(0.3) = 125.6\,\text{volts}

    This interval corresponds to μ±2σ\mu \pm 2\sigma. By the Empirical Rule, approximately 95% of voltage amounts fall in this range.

  • (b) Percentage between 124.1 volts124.1\,\text{volts} and 125.9 volts125.9\,\text{volts}:

125.0−3(0.3)=124.1 volts125.0 - 3(0.3) = 124.1\,\text{volts}

125.0+3(0.3)=125.9 volts125.0 + 3(0.3) = 125.9\,\text{volts}

    This interval corresponds to μ±3σ\mu \pm 3\sigma. By the Empirical Rule, approximately 99.7% of voltage amounts fall in this range.

  • Interpretation: Virtually all (about 99.7%99.7\%) generator voltage outputs lie between 124.1 volts124.1\,\text{volts} and 125.9 volts125.9\,\text{volts}, with 95%95\% lying within 124.4 volts124.4\,\text{volts} and 125.6 volts125.6\,\text{volts}.

Measures of Relative Standing & Boxplots

Measures of relative standing indicate the location of a value relative to other values within a data set.

z-Scores (Standardized Values)

A z-score represents the number of standard deviations a value xx is located above or below the mean.

  • Sample z-score Formula:

z=x−xˉsz = \frac{x - \bar{x}}{s}

  • Population z-score Formula:

z=x−μσz = \frac{x - \mu}{\sigma}

  • Usual vs. Unusual z-Scores:

    • Usual values: −2≤z≤2-2 \le z \le 2

    • Unusual values: z<−2z < -2 or z>2z > 2

  • Earthquake Example:

    • Magnitude dataset (n=50n = 50): Mean xˉ=1.184\bar{x} = 1.184, Standard Deviation s = 0.587$.\n * Convert magnitude x = 1.766 to a z-score:\n\nz = \frac{1.766 - 1.184}{0.587} = \frac{0.582}{0.587} \approx 0.99\n\n * Conclusion: Because z = 0.99liesbetweenlies between-2andand2,anearthquakeofmagnitude, an earthquake of magnitude1.766 is considered **usual**.\n\n## Percentiles and Quartiles\n\n* **Percentile**: Measures of location that divide a set of ordered data into 100 equal groups.\n* **Interpretation Example**:\n * Statement: "A child is in the 75th percentile for height."\n * Meaning: The child is taller than 75\%ofchildreninthereferencegroup,andshorterthanorequaltoof children in the reference group, and shorter than or equal to25\% of them.\n\n### Percentile Formulas and Rules\n\n* **Finding Percentile of a Given Data Value x**:\n\n\text{Percentile of } x = \left( \frac{\text{Number of values less than } x}{\text{Total number of values } n} \right) \times 100\n\n* **Finding Data Value Corresponding to Percentile k((P_k)**:\n * Compute Locator L:\n\nL = \left( \frac{k}{100} \right) \cdot n\n\n * **Decision Rule for Locator L**:\n 1. If L∗∗isnotawholenumber∗∗:Round**is not a whole number**: RoundLUPtothenextinteger.ThevalueofUP to the next integer. The value ofP_kistheis theL\text{th} value in the ordered dataset.\n 2. If L∗∗isawholenumber∗∗:Thevalueof**is a whole number**: The value ofP_kistheaverageoftheis the average of theL\text{th}valueandthevalue and the(L+1)\text{th} value in the ordered dataset.\n\n### Step-by-Step Dataset Examples (n = 16)\n\nSorted Dataset: 2, 18, 19, 22, 24, 26, 26, 35, 35, 35, 36, 38, 40, 46, 48, 65\n\n* **(a) Find the 32nd Percentile (P_{32})**:\n\nL = \left( \frac{32}{100} \right) \times 16 = 5.12\n\n * Round UP to 6..P_{32} is the 6th value in the sorted list.\n * 6th value = 26.\n\n* **(b) Find the 10th Percentile (P_{10})**:\n\nL = \left( \frac{10}{100} \right) \times 16 = 1.6\n\n * Round UP to 2..P_{10} is the 2nd value in the sorted list.\n * 2nd value = 18.\n\n* **(c) Find Median / 50th Percentile (P_{50})**:\n\nL = \left( \frac{50}{100} \right) \times 16 = 8\n\n * Since 8 is a whole number, average the 8th and 9th values.\n * 8th value = 35,9thvalue=, 9th value =35$.

Median=35+352=35\text{Median} = \frac{35 + 35}{2} = 35

  • (d) Find 3rd Quartile / 75th Percentile (Q3=P75Q_3 = P_{75}):

L=(75100)×16=12L = \left( \frac{75}{100} \right) \times 16 = 12

  • Since 1212 is a whole number, average the 12th and 13th values.

  • 12th value = 3838, 13th value = 40$.\n\nQ_3 = \frac{38 + 40}{2} = 39\n\n* **(e) Find Percentile for Value 22**:\n * Count values strictly less than 22::2, 18, 19 (3 values).\n\n\text{Percentile} = \left( \frac{3}{16} \right) \times 100 = 18.75\% \approx 19\text{th percentile}\n\n * Value 22correspondstocorresponds toP_{19}.\n\n## 5-Number Summary, Boxplots, and Outlier Calculations\n\n### The 5-Number Summary\n\nThe 5-number summary consists of five values that summarize a dataset:\n1. **Minimum Value**\n2. **First Quartile (Q_1 = P_{25})**\n3. **Median (Q_2 = P_{50})**\n4. **Third Quartile (Q_3 = P_{75})**\n5. **Maximum Value**\n\n* **Calculations for the Example Dataset**:\n * Minimum = 2\n * Q_1 = P_{25}::L = \left(\frac{25}{100}\right) \times 16 = 4(average4thand5thvalues:(average 4th and 5th values:\frac{22 + 24}{2} = 23)\n * Median (Q_2)=) =35\n * Q_3 = 39\n * Maximum = 65\n * **5-Number Summary**: `[2, 23, 35, 39, 65]`\n\n### Interquartile Range (IQR) and Outlier Boundary Identification\n\n* **Interquartile Range Formula**:\n\nIQR = Q_3 - Q_1\n\n * For example data: IQR = 39 - 23 = 16.\n\n* **Outlier Boundaries**:\n * **Lower Boundary**:\n\n\text{Lower Boundary} = Q_1 - 1.5 \times IQR\n\n * **Upper Boundary**:\n\n\text{Upper Boundary} = Q_3 + 1.5 \times IQR\n\n* **Example Outlier Analysis for Values 2andand65**:\n * Lower Boundary:\n\n23 - 1.5(16) = 23 - 24 = -1\n\n * Upper Boundary:\n\n39 + 1.5(16) = 39 + 24 = 63\n\n * **Check value 2∗∗:Since**: Since2 \ge -1,thevalue, the value2 lies within the boundary and is **not an outlier**.\n * **Check value 65∗∗:Since**: Since65 > 63,thevalue, the value65 exceeds the upper boundary and **is an outlier**.\n\n### Boxplots\nA boxplot (or box-and-whisker diagram) is a graphical plot of a dataset that displays:\n* A box drawn from Q_1totoQ_3$.

    • A vertical line drawn inside the box at the Median (Q2Q_2).

    • Line segments ("whiskers") extending from the box out to the minimum and maximum data values (or to the furthest non-outlier values in a modified boxplot).