Business Statistics: Summary Measures
Measures of Central Location
Central location describes how numerical data cluster around a middle value.
Arithmetic Mean: The primary measure of central location calculated by dividing the sum of observations by the count.
Sample Mean:
Population Mean:
Sensitive to extreme values or outliers.
Median: The middle value when observations are sorted in ascending order. Robust to outliers.
Odd sample size: Exact middle observation.
Even sample size: Average of the two middle observations.
Mode: The observation that occurs most frequently.
Can be unimodal, bimodal, multimodal, or have no mode.
Sole meaningful central measure for categorical data.
Weighted Mean: Account for observations with varying degrees of importance.
, where
Skewness and Central Measures:
Symmetric:
Positively Skewed (Right-tailed):
Negatively Skewed (Left-tailed):
Percentiles and Boxplots
Percentiles: Divide data into 100 equal parts. The percentile has approximately of observations below it and above it.
Quartiles:
First Quartile (): 25th percentile
Second Quartile (): 50th percentile (Median)
Third Quartile (): 75th percentile
Interquartile Range (IQR): The range of the middle of data.
Five-Number Summary: Minimum, , Median (), , and Maximum.
Boxplot Structure:

Box spans from to with a line at the median.
Whiskers extend to the minimum and maximum values within from the box edges.
Outliers: Values located further than from or .
Measures of Dispersion
Measures of dispersion evaluate the variability or spread of data.
Range: Difference between the maximum and minimum values.
Mean Absolute Deviation (MAD): Average of absolute deviations from the mean.
Sample MAD:
Population MAD:
Variance: Average of squared deviations from the mean.
Sample Variance:
Population Variance:
Standard Deviation: Positive square root of variance, returning units to the original scale.
Sample Standard Deviation:
Population Standard Deviation:
Coefficient of Variation (CV): Unitless measure of relative risk/variability per unit of mean.
Sample CV:
Population CV:
Mean-Variance Analysis and the Sharpe Ratio
Mean-Variance Analysis: Measures financial asset performance by evaluating reward (mean return) against risk (variance or standard deviation).
Sharpe Ratio: Measures excess return per unit of risk relative to a risk-free rate ().
Higher values indicate superior risk-adjusted return performance.
Analysis of Relative Location
Chebyshev's Theorem: For any dataset and , the proportion of values within standard deviations of the mean is at least .
: At least of data lie within
: At least of data lie within
Empirical Rule: Applies to symmetric, bell-shaped distributions.

Approximately of observations lie within
Approximately of observations lie within
Almost (roughly ) of observations lie within
z-Score: Standardized measure of relative position, representing distance from the mean in standard deviation units.
Sample z-score:
Population z-score:
Values with in bell-shaped distributions indicate potential outliers.
Measures of Association
Covariance: Quantifies the direction of a linear relationship between two variables.
Sample Covariance:
Population Covariance:
Sensitive to measurement units; cannot indicate relationship strength.
Correlation Coefficient: Unit-free measure of both direction and strength of a linear relationship ().
Sample Correlation:
Population Correlation:
: Perfect positive linear relationship
: No linear relationship
: Perfect negative linear relationship
Geometric Mean and Growth Rates
Geometric Mean Return: Multiplicative average used for compounding multiperiod investment returns.
Compound Average Growth Rate: Multiplicative average growth rate across time periods or observations.
Using rates:
Using end values () and start values ():