Ch. 3- Displaying & Describing Data Part II
Historical Foundations of Descriptive Statistics
Adolphe Quetelet emphasized the scientific importance of statistical aggregation, stating: "The determination of the average man is not merely a matter of speculative curiosity; it may be of the most important service to the science of man and the social system."

Measures of Location
Measures of location summarize an entire data set into a single representative value at the aggregate level.
These measures describe where data points are situated along a numerical or ordinal scale.
Measures of central tendency represent a specific subcategory of location measures that describe the "typical" or "central" value of a distribution.
All measures of central tendency are measures of location, but not all measures of location qualify as measures of central tendency.
The Mean (Arithmetic Average)
The mean is defined as the sum of all observed values on a variable divided by the total number of observations ().
Notation conventions:
Sample mean is represented by .
Population mean is represented by .
Analytical Properties:
Highly influenced by extreme values (outliers).
Readily susceptible to mathematical manipulation and algebraic transformations.
Mathematical Formula for Continuous and Pseudo-Continuous Data:
Applicable Variable Types:
Ordered-Categorical Variables
Pseudo-Continuous Variables
Mean of Dichotomous Variables:
For binary variables coded as and , the mean equals the proportion () of cases coded as : where is the count of observations with a value of , and represents the proportion.
The Median
The median is the "middle point" or the 50th percentile of an ordered distribution.
Analytical Properties:
Robust against the influence of extreme outliers.
Not as easily manipulated mathematically compared to the mean.
Location Calculation:
To locate the position of the median in an ordered sequence of values:
Applicable Variable Types:
Ordered-Categorical Variables
Pseudo-Continuous Variables
The Mode
The mode is the most frequently occurring value observed in a data set for a given variable.
Analytical Properties:
Provides no information regarding the location of scores that do not fall at the mode, other than indicating that non-modal values occur less frequently.
Completely unaffected by outliers.
Classifications of Modality:
Unimodal: A distribution containing exactly most frequently occurring value.
Bimodal: A distribution containing exactly most frequently occurring values.
Trimodal: A distribution containing exactly most frequently occurring values.
Multimodal: A distribution containing or more most frequently occurring values.
Applicable Variable Types:
Nominal Variables
Dichotomous Variables
Ordered-Categorical / Pseudo-Continuous Variables
Examples of Mode Determination:
For data set , the value occurs twice. Thus, (Unimodal).
For data set , the value occurs four times. Thus,
Comparison of Central Tendency Measures in Skewed Distributions
In skewed distributions, the relative position of central tendency measures reflects the direction of skew:
Mode remains directly under the peak of the curve where frequency density is highest.
Median shifts toward the long tail to divide the distribution area in half.
Mean is pulled furthest into the long tail due to its sensitivity to extreme values.

Measures of Dispersion
Measures of dispersion describe the spread of data or the magnitude of variability present around central values.

Range and Interquartile Range
Range:
The total distance between the absolute maximum and minimum values in a data set.
Interquartile Range (IQR):
The distance spanning the middle of the distribution between the 25th percentile () and the 75th percentile ().

Deviation Scores and Summary Variance Metrics
Deviation Score:
Represents the distance and direction of an individual score from the sample mean :
Sum of Deviations:
The algebraic sum of raw deviations around the mean is always equal to zero:
Average Absolute Deviation:
Measures dispersion by taking the average of absolute deviations from the mean, preventing positive and negative values from canceling each other out:
The numerator represents the sum of absolute deviations.
Average Squared Deviation / Variance ():
Measures dispersion by taking the mean of squared deviations from the mean.
The numerator represents the sum of squared deviations.
Standard Deviation ():
The positive square root of variance, returning the measure of dispersion to the original units of measurement:
Standard Notation Conventions
Variance Notation:
denotes sample variance.
denotes population variance.
Standard Deviation Notation:
denotes sample standard deviation.
denotes population standard deviation.
Interpreting Aggregate Statistics vs. General Propositions
Theoretical Distinction (Bakan, 1967, p. 35):
General-Type Proposition: Asserts something presumably true of each and every member belonging to a designable class.
Aggregate-Type Proposition: Asserts something presumably true of the class considered as an aggregate group, which may not hold true for every individual member of that group.
Empirical Example (DePaulo Study on Marital Status):
Longitudinal research demonstrates aggregate-level patterns where single individuals value meaningful work more highly than married individuals, and single people maintain greater social connectivity with parents, siblings, friends, neighbors, and coworkers.
The aggregate observation that marriage tends to make individuals more insular represents an aggregate-type proposition and cannot be fallaciously applied as a general-type proposition to every single individual.

Visualizing Data and Outliers: Box and Whisker Plots
Box and whisker plots display distributional shape, central tendency, dispersion, and extreme outliers.
Key Components of a Box Plot
50th Percentile: Median line located inside the central box.
25th Percentile (Lower Hinge): Bottom boundary of the central box ().
75th Percentile (Upper Hinge): Top boundary of the central box ().
Interquartile Range (IQR): Vertical height of the central box ().
Upper Whisker: Line extending to the highest observation within above the upper hinge.
Lower Whisker: Line extending to the lowest observation within below the lower hinge.
Outliers: Individual observations situated beyond the upper or lower whiskers (indicated by asterisks with case numbers).
Diagnostic Worked Example and Skew Analysis
Sample Data Set ():
Diagnostic Questions for Determining Skewness from a Box Plot:
Are there any outliers present beyond the whiskers?
Which half of the central box (above or below the median line) is wider?
Plot Diagnostic Results:
Case (value ) and Case (value ) are marked as extreme outliers well above the upper whisker.
The presence of prominent upper outliers confirms a positive (rightward) skew in the distribution.
