Quantitative Data Distributions and Histograms

Quantitative Data Requirements
  • Order: Discrete quantitative variables must be listed in strict numerical order from smallest to largest.

  • Completeness: Include all possible values between the smallest and largest, even those with a frequency of 00.

Histograms vs. Bar Graphs
  • Histogram: Used for quantitative data; vertical bars must touch (contiguous, no gaps).

  • Bar Graph: Used for qualitative or categorical data; bars stand alone with spaces between them.

Sibling Dataset Summary
  • Total Sample Size: n=78n = 78

  • Formula:
    Relative Frequency=FrequencyTotal Sample Size\text{Relative Frequency} = \frac{\text{Frequency}}{\text{Total Sample Size}}

  • Relative Frequency Values:

    • 00 siblings: 578≈0.064\frac{5}{78} \approx 0.064

    • 11 sibling: 2778≈0.346\frac{27}{78} \approx 0.346

    • 22 siblings: 1878≈0.231\frac{18}{78} \approx 0.231

    • 33 siblings: 978≈0.115\frac{9}{78} \approx 0.115

    • 44 siblings: 878≈0.103\frac{8}{78} \approx 0.103

    • 55 siblings: 778≈0.090\frac{7}{78} \approx 0.090

    • 66 siblings: 278≈0.026\frac{2}{78} \approx 0.026

    • 77 siblings: 278≈0.026\frac{2}{78} \approx 0.026

  • Total Sum: 1.0011.001 (differs slightly from 1.0001.000 due to rounding).

Constructing Histograms
  • Title: Clear descriptive title.

  • Horizontal Axis (x-axis): Sequential values from smallest to largest (00 to 77).

  • Vertical Axis (y-axis): Uniform relative frequency increments (e.g., 0.0500.050).

  • Bars: Connected rectangular bars centered over each discrete value.

Common Histogram Shapes
  • Bell-Shaped: Symmetric with one central peak.

  • Right-Skewed: Tail extends to the right (larger values); peak is on the left.

  • Left-Skewed: Tail extends to the left (smaller values); peak is on the right.

  • Uniform: Flat profile with approximately equal bar heights.

  • Bimodal: Features two distinct peaks.

Core Components of Distribution Descriptions
  • Shape: Profile category (e.g., right-skewed for the sibling dataset).

  • Spread: Variability and overall range of values.

  • Outliers: Unusually large jumps or values far outside the main data pattern.