Categorical and Quantitative Data Analysis Study Guide
Categorical Variables and Displays
Definitions:
Categorical Variable: Places an individual case into one of several categories (e.g., gender, race, grade level).
Frequency: The total numerical count of individuals in a given category.
Relative Frequency: The fraction or percentage of individuals relative to the total, enabling comparisons between datasets of different sizes.
Distribution: Shows the values/categories a variable takes along with their corresponding frequencies or relative frequencies.
Tables for Categorical Data:
One-Way Table: Displays the distribution of a single categorical variable.
Two-Way Table: Displays the distribution of individuals classified by two categorical variables.
Distributions in Two-Way Tables:
Marginal Distribution: The distribution of values of one variable among all individuals described by the table (found in the row or column margins).
Conditional Distribution: The distribution of values of one variable for a specific, fixed value of another variable.
Graphing Two Categorical Variables:
Comparative Bar Graph: Compares categories using adjacent vertical bars with spaces between explanatory groups. Always uses relative frequencies unless group totals are equal.
Segmented Bar Graph: Displays each explanatory category as a single bar stacked to (), partitioned into segments corresponding to response categories.
Mosaic Plot: A modified segmented bar graph where the width of each bar is proportional to the relative frequency of that explanatory variable category.

Assessing Association:
Two categorical variables have an association if knowing the value of one variable helps predict the value of the other (i.e., conditional distributions differ).
Requirements for Describing an Association:
Include context.
Address every value of the explanatory variable.
Compare relative frequencies or percentages across categories.
Describing Quantitative Distributions (SCVU)
Quantitative Variable: A numerical variable for which calculating an arithmetic average makes logical sense.
SCVU Framework: Always state descriptions in context addressing four components:
Shape (S):
Roughly Symmetric: Upper and lower halves have approximately equal spread.
Mound-Shaped: Single central peak sloping symmetrically down on both sides.
Right-Skewed: Tail extends toward higher values; majority of observations are lower.
Left-Skewed: Tail extends toward lower values; majority of observations are higher.
Modality: Unimodal (one peak), Bimodal (two peaks), or Uniform (flat distribution).
Center (C):
Mean ( or ): Arithmetic average calculated as:
Median: Midpoint of a dataset where of observations are smaller and are larger.
Variability (V):
Range: Maximum minus minimum value ().
Interquartile Range (): Range of the middle of observations:
Five-Number Summary: Minimum, ( percentile), Median ( percentile), ( percentile), Maximum.
Standard Deviation (): Typical distance of values from the mean:
Unusual Features (U):
Outlier Rule: Observations falling beyond upper or lower fences:
Gaps and Clusters: Note empty intervals or distinct concentrations of data.
Resistance and Distribution Shapes
Resistance to Extreme Values:
Resistant Measures: Statistics that remain unchanged or barely change when extreme values or outliers are added (e.g., Median and ).
Non-Resistant Measures: Statistics significantly affected by extreme values or outliers (e.g., Mean, Range, and Standard Deviation ).
Effect of Skewness on Center:
Symmetric Distribution: Mean is approximately equal to median ().
Right-Skewed Distribution: Mean is pulled toward the higher tail, making .
Left-Skewed Distribution: Mean is pulled toward the lower tail, making .
Selection Rule: Use Median and for skewed data or distributions with outliers; use Mean and Standard Deviation for roughly symmetric distributions without outliers.
Quantitative Graph Types
Dotplots:
Plots individual data values as dots above a number line.
Preferred measures: Median for center, Range//clusters for variability.
Stemplots (Stem-and-Leaf Displays):
Separates observations into a stem (leading digits) and leaf (final single digit).
Requires a key specifying units (e.g., ).
Do not skip stem values even if empty to clearly demonstrate gaps.
Histograms:
Groups continuous quantitative data into adjacent bins/classes of equal width.
Values on class boundaries belong to the bin on the right.
Median class is found by accumulating bin counts from left to right until reaching the middle position.
Boxplots:
Visual display of the five-number summary and outliers.
Central box spans to with a line at the median; whiskers extend to the extreme non-outlier values, and outliers are plotted as isolated points.
Cannot display distribution modality (peaks) or sample size .