Categorical and Quantitative Data Analysis Study Guide

Categorical Variables and Displays

  • Definitions:

    • Categorical Variable: Places an individual case into one of several categories (e.g., gender, race, grade level).

    • Frequency: The total numerical count of individuals in a given category.

    • Relative Frequency: The fraction or percentage of individuals relative to the total, enabling comparisons between datasets of different sizes.

    • Distribution: Shows the values/categories a variable takes along with their corresponding frequencies or relative frequencies.

  • Tables for Categorical Data:

    • One-Way Table: Displays the distribution of a single categorical variable.

    • Two-Way Table: Displays the distribution of individuals classified by two categorical variables.

  • Distributions in Two-Way Tables:

    • Marginal Distribution: The distribution of values of one variable among all individuals described by the table (found in the row or column margins).

    • Conditional Distribution: The distribution of values of one variable for a specific, fixed value of another variable.

  • Graphing Two Categorical Variables:

    • Comparative Bar Graph: Compares categories using adjacent vertical bars with spaces between explanatory groups. Always uses relative frequencies unless group totals are equal.

    • Segmented Bar Graph: Displays each explanatory category as a single bar stacked to 11 (100%100\%), partitioned into segments corresponding to response categories.

    • Mosaic Plot: A modified segmented bar graph where the width of each bar is proportional to the relative frequency of that explanatory variable category.

Mosaic Plot Example
  • Assessing Association:

    • Two categorical variables have an association if knowing the value of one variable helps predict the value of the other (i.e., conditional distributions differ).

    • Requirements for Describing an Association:

    1. Include context.

    2. Address every value of the explanatory variable.

    3. Compare relative frequencies or percentages across categories.

Describing Quantitative Distributions (SCVU)

  • Quantitative Variable: A numerical variable for which calculating an arithmetic average makes logical sense.

  • SCVU Framework: Always state descriptions in context addressing four components:

    • Shape (S):

    • Roughly Symmetric: Upper and lower halves have approximately equal spread.

    • Mound-Shaped: Single central peak sloping symmetrically down on both sides.

    • Right-Skewed: Tail extends toward higher values; majority of observations are lower.

    • Left-Skewed: Tail extends toward lower values; majority of observations are higher.

    • Modality: Unimodal (one peak), Bimodal (two peaks), or Uniform (flat distribution).

    • Center (C):

    • Mean (xˉ\bar{x} or μ\mu): Arithmetic average calculated as:       xˉ=∑xin\bar{x} = \frac{\sum x_i}{n}

    • Median: Midpoint of a dataset where 50%50\% of observations are smaller and 50%50\% are larger.

    • Variability (V):

    • Range: Maximum minus minimum value (Range=Max−Min\text{Range} = \text{Max} - \text{Min}).

    • Interquartile Range (IQRIQR): Range of the middle 50%50\% of observations:       IQR=Q3−Q1IQR = Q3 - Q1

    • Five-Number Summary: Minimum, Q1Q1 (25th25\text{th} percentile), Median (50th50\text{th} percentile), Q3Q3 (75th75\text{th} percentile), Maximum.

    • Standard Deviation (sxs_x): Typical distance of values from the mean:       sx=∑(xi−xˉ)2n−1s_x = \sqrt{\frac{\sum (x_i - \bar{x})^2}{n - 1}}

    • Unusual Features (U):

    • Outlier Rule: Observations falling beyond upper or lower fences:       Upper Fence=Q3+1.5×IQR\text{Upper Fence} = Q3 + 1.5 \times IQR       Lower Fence=Q1−1.5×IQR\text{Lower Fence} = Q1 - 1.5 \times IQR

    • Gaps and Clusters: Note empty intervals or distinct concentrations of data.

Resistance and Distribution Shapes

  • Resistance to Extreme Values:

    • Resistant Measures: Statistics that remain unchanged or barely change when extreme values or outliers are added (e.g., Median and IQRIQR).

    • Non-Resistant Measures: Statistics significantly affected by extreme values or outliers (e.g., Mean, Range, and Standard Deviation sxs_x).

  • Effect of Skewness on Center:

    • Symmetric Distribution: Mean is approximately equal to median (xˉ≈Median\bar{x} \approx \text{Median}).

    • Right-Skewed Distribution: Mean is pulled toward the higher tail, making xˉ>Median\bar{x} > \text{Median}.

    • Left-Skewed Distribution: Mean is pulled toward the lower tail, making xˉ<Median\bar{x} < \text{Median}.

    • Selection Rule: Use Median and IQRIQR for skewed data or distributions with outliers; use Mean and Standard Deviation for roughly symmetric distributions without outliers.

Quantitative Graph Types

  • Dotplots:

    • Plots individual data values as dots above a number line.

    • Preferred measures: Median for center, Range/IQRIQR/clusters for variability.

  • Stemplots (Stem-and-Leaf Displays):

    • Separates observations into a stem (leading digits) and leaf (final single digit).

    • Requires a key specifying units (e.g., 1∣2=12 hours1 \mid 2 = 12\,\text{hours}).

    • Do not skip stem values even if empty to clearly demonstrate gaps.

  • Histograms:

    • Groups continuous quantitative data into adjacent bins/classes of equal width.

    • Values on class boundaries belong to the bin on the right.

    • Median class is found by accumulating bin counts from left to right until reaching the middle position.

  • Boxplots:

    • Visual display of the five-number summary and outliers.

    • Central box spans Q1Q1 to Q3Q3 with a line at the median; whiskers extend to the extreme non-outlier values, and outliers are plotted as isolated points.

    • Cannot display distribution modality (peaks) or sample size nn.