Ch. 3- Displaying & Describing Data Part II

Historical Foundations of Descriptive Statistics

  • Adolphe Quetelet emphasized the scientific importance of statistical aggregation, stating: "The determination of the average man is not merely a matter of speculative curiosity; it may be of the most important service to the science of man and the social system."

Portrait of Adolphe Quetelet

Measures of Location

  • Measures of location summarize an entire data set into a single representative value at the aggregate level.

  • These measures describe where data points are situated along a numerical or ordinal scale.

  • Measures of central tendency represent a specific subcategory of location measures that describe the "typical" or "central" value of a distribution.

  • All measures of central tendency are measures of location, but not all measures of location qualify as measures of central tendency.

The Mean (Arithmetic Average)

  • The mean is defined as the sum of all observed values on a variable divided by the total number of observations (NN).

  • Notation conventions:

    • Sample mean is represented by Xˉ\bar{X}.

    • Population mean is represented by μ\mu.

  • Analytical Properties:

    • Highly influenced by extreme values (outliers).

    • Readily susceptible to mathematical manipulation and algebraic transformations.

  • Mathematical Formula for Continuous and Pseudo-Continuous Data:   Xˉ=∑xN=1N∑x\bar{X} = \frac{\sum x}{N} = \frac{1}{N} \sum x

  • Applicable Variable Types:

    • Ordered-Categorical Variables

    • Pseudo-Continuous Variables

  • Mean of Dichotomous Variables:

    • For binary variables coded as 00 and 11, the mean equals the proportion (pp) of cases coded as 11:   Xˉ=N1N=px(1)\bar{X} = \frac{N_1}{N} = p_x(1)   where N1N_1 is the count of observations with a value of 11, and pp represents the proportion.

The Median

  • The median is the "middle point" or the 50th percentile of an ordered distribution.

  • Analytical Properties:

    • Robust against the influence of extreme outliers.

    • Not as easily manipulated mathematically compared to the mean.

  • Location Calculation:

    • To locate the position of the median in an ordered sequence of NN values:   Median location=N+12\text{Median location} = \frac{N + 1}{2}

  • Applicable Variable Types:

    • Ordered-Categorical Variables

    • Pseudo-Continuous Variables

The Mode

  • The mode is the most frequently occurring value observed in a data set for a given variable.

  • Analytical Properties:

    • Provides no information regarding the location of scores that do not fall at the mode, other than indicating that non-modal values occur less frequently.

    • Completely unaffected by outliers.

  • Classifications of Modality:

    • Unimodal: A distribution containing exactly 11 most frequently occurring value.

    • Bimodal: A distribution containing exactly 22 most frequently occurring values.

    • Trimodal: A distribution containing exactly 33 most frequently occurring values.

    • Multimodal: A distribution containing 33 or more most frequently occurring values.

  • Applicable Variable Types:

    • Nominal Variables

    • Dichotomous Variables

    • Ordered-Categorical / Pseudo-Continuous Variables

  • Examples of Mode Determination:

    • For data set X=[2,4,2,8,6]X = [2, 4, 2, 8, 6], the value 22 occurs twice. Thus, Mode=2\text{Mode} = 2 (Unimodal).

    • For data set X=[1,1,1,0,0,1]X = [1, 1, 1, 0, 0, 1], the value 11 occurs four times. Thus, Mode=1\text{Mode} = 1

Comparison of Central Tendency Measures in Skewed Distributions

  • In skewed distributions, the relative position of central tendency measures reflects the direction of skew:

    • Mode remains directly under the peak of the curve where frequency density is highest.

    • Median shifts toward the long tail to divide the distribution area in half.

    • Mean is pulled furthest into the long tail due to its sensitivity to extreme values.

Comparison of mean, median, and mode in a skewed distribution

Measures of Dispersion

  • Measures of dispersion describe the spread of data or the magnitude of variability present around central values.

Comparison of distribution curves showing differing levels of dispersion

Range and Interquartile Range

  • Range:

    • The total distance between the absolute maximum and minimum values in a data set.   Range=Xmax⁡−Xmin⁡\text{Range} = X_{\max} - X_{\min}

  • Interquartile Range (IQR):

    • The distance spanning the middle 50%50\% of the distribution between the 25th percentile (X25th%X_{25\text{th}\%}) and the 75th percentile (X75th%X_{75\text{th}\%}).   Interquartile Range=X75th%−X25th%\text{Interquartile Range} = X_{75\text{th}\%} - X_{25\text{th}\%}

Diagram illustrating Range and Interquartile Range across percentiles

Deviation Scores and Summary Variance Metrics

  • Deviation Score:

    • Represents the distance and direction of an individual score XiX_i from the sample mean Xˉ\bar{X}:   Deviation=Xi−Xˉ\text{Deviation} = X_i - \bar{X}

  • Sum of Deviations:

    • The algebraic sum of raw deviations around the mean is always equal to zero:   ∑(Xi−Xˉ)=0\sum (X_i - \bar{X}) = 0

  • Average Absolute Deviation:

    • Measures dispersion by taking the average of absolute deviations from the mean, preventing positive and negative values from canceling each other out:   Average Absolute Deviation=∑∣Xi−Xˉ∣N\text{Average Absolute Deviation} = \frac{\sum |X_i - \bar{X}|}{N}

    • The numerator represents the sum of absolute deviations.

  • Average Squared Deviation / Variance (s2s^2):

    • Measures dispersion by taking the mean of squared deviations from the mean.   s2=∑(Xi−Xˉ)2N=∑Xi2−(∑Xi)2NNs^2 = \frac{\sum (X_i - \bar{X})^2}{N} = \frac{\sum X_i^2 - \frac{(\sum X_i)^2}{N}}{N}

    • The numerator ∑(Xi−Xˉ)2\sum (X_i - \bar{X})^2 represents the sum of squared deviations.

  • Standard Deviation (ss):

    • The positive square root of variance, returning the measure of dispersion to the original units of measurement:   s=s2s = \sqrt{s^2}

Standard Notation Conventions

  • Variance Notation:

    • s2s^2 denotes sample variance.

    • σ2\sigma^2 denotes population variance.

  • Standard Deviation Notation:

    • ss denotes sample standard deviation.

    • σ\sigma denotes population standard deviation.

Interpreting Aggregate Statistics vs. General Propositions

  • Theoretical Distinction (Bakan, 1967, p. 35):

    • General-Type Proposition: Asserts something presumably true of each and every member belonging to a designable class.

    • Aggregate-Type Proposition: Asserts something presumably true of the class considered as an aggregate group, which may not hold true for every individual member of that group.

  • Empirical Example (DePaulo Study on Marital Status):

    • Longitudinal research demonstrates aggregate-level patterns where single individuals value meaningful work more highly than married individuals, and single people maintain greater social connectivity with parents, siblings, friends, neighbors, and coworkers.

    • The aggregate observation that marriage tends to make individuals more insular represents an aggregate-type proposition and cannot be fallaciously applied as a general-type proposition to every single individual.

American Psychological Association post discussing single vs married life outcomes

Visualizing Data and Outliers: Box and Whisker Plots

  • Box and whisker plots display distributional shape, central tendency, dispersion, and extreme outliers.

Key Components of a Box Plot

  • 50th Percentile: Median line located inside the central box.

  • 25th Percentile (Lower Hinge): Bottom boundary of the central box (X25th%X_{25\text{th}\%}).

  • 75th Percentile (Upper Hinge): Top boundary of the central box (X75th%X_{75\text{th}\%}).

  • Interquartile Range (IQR): Vertical height of the central box (X75th%−X25th%X_{75\text{th}\%} - X_{25\text{th}\%}).

  • Upper Whisker: Line extending to the highest observation within 1.5×IQR1.5 \times \text{IQR} above the upper hinge.

  • Lower Whisker: Line extending to the lowest observation within 1.5×IQR1.5 \times \text{IQR} below the lower hinge.

  • Outliers: Individual observations situated beyond the upper or lower whiskers (indicated by asterisks with case numbers).

Diagnostic Worked Example and Skew Analysis

  • Sample Data Set (N=15N = 15):   X=[1,3,6,7,7,8,10,11,15,17,18,19,21,55,67]X = [1, 3, 6, 7, 7, 8, 10, 11, 15, 17, 18, 19, 21, 55, 67]

  • Diagnostic Questions for Determining Skewness from a Box Plot:

    1. Are there any outliers present beyond the whiskers?

    2. Which half of the central box (above or below the median line) is wider?

  • Plot Diagnostic Results:

    • Case 1414 (value 5555) and Case 1515 (value 6767) are marked as extreme outliers well above the upper whisker.

    • The presence of prominent upper outliers confirms a positive (rightward) skew in the distribution.

Box and whisker plot demonstrating outliers and quartiles