stats3

Chapter 2: Exploring Data with Tables and Graphs

2.1 Frequency Distributions for Organizing and Summarizing Data

  • Frequency distribution (or frequency table) helps in organizing large data sets.

  • A frequency distribution clarifies the nature of the data distribution.

2.2 Histograms

  • Histograms represent frequency distributions graphically, allowing for easier visualization of data.

  • Each bar in a histogram represents the frequency of data within defined ranges (classes).

2.3 Graphs that Enlighten and Graphs that Deceive

  • Understanding how to read graphs is essential, as they can sometimes misrepresent the data.

  • Awareness of graphical misleading is crucial for accurate data interpretation.

2.4 Scatterplots, Correlation, and Regression

  • Scatterplots show relationships between two quantitative variables.

  • Correlation measures the strength and direction of the relationship.

  • Regression is used for predicting values based on correlations.

Key Concepts

  • Frequency Distribution: A table that outlines how data are divided into categories or classes, showing the count (frequency) of occurrences in each category.

Definitions

  • Lower Class Limits: The smallest values in each class.

  • Upper Class Limits: The largest values in each class.

  • Class Boundaries: Numbers separating classes without gaps in limits.

  • Class Midpoints: Values in the middle of each class, calculated by averaging the lower and upper limits.

  • Class Width: Difference between consecutive lower class limits or boundaries.

Procedure for Constructing a Frequency Distribution

  1. Select the number of classes (5 to 20 recommended).

  2. Calculate class width: [ Class~width \approx \frac{(max~-~min)}{number~of~classes} ]

  3. Choose the first lower class limit (minimum value or a convenient number below it).

  4. List other lower class limits using the first limit and computed width.

  5. Identify upper class limits according to the lower class limits.

  6. Tally each data value in the appropriate class, then count to find frequency.

Example: Commute Time in Los Angeles

Steps to Construct Frequency Distribution

  • Step 1: Choose 7 classes.

  • Step 2: Calculate class width: Class width ≈ 15 (rounded from 12.1).

  • Step 3: First lower limit selected as 0.

  • Step 4: Lower class limits identified as 0, 15, 30, 45, 60, 75, 90.

  • Step 5: Upper class limits identified as 14, 29, 44, 59, 74, 89, 104.

  • Step 6: Tally marks counted for each class, resulting in:

    • 0-14: 6

    • 15-29: 18

    • 30-44: 14

    • 45-59: 5

    • 60-74: 5

    • 75-89: 1

    • 90-104: 1

Relative Frequency Distribution

  • Each class frequency can be represented as a relative frequency (percentage).

  • Relative frequency formula: [ Relative~frequency \approx \frac{frequency~for~class}{sum~of~all~frequencies} ]

  • Cumulative percentages should total close to 100%.

    • Example frequencies:

      • 0-14: 12%

      • 15-29: 36%

      • 30-44: 28%

      • 45-59: 10%

      • 60-74: 10%

      • 75-89: 2%

      • 90-104: 2%

Comparisons

  • Comparing two or more relative frequency distributions aids in understanding differences in data.

Example: Comparing Daily Commute Times in NY and Boise

  • Observed differences in relative frequencies reflect variations in urban commute climates.

  • Higher percentages of shorter commute times found in Boise than in New York.

Cumulative Frequency Distribution

  • Cumulative frequency adds the frequencies for each class and all previous classes, used for summarization.

Critical Thinking: Analyzing Data Distributions

  1. A normal distribution shows frequencies rising to a peak and then decreasing symmetrically.

  2. Gaps in distributions may indicate multiple populations.

Example: Weight Distribution of Pennies

  • Frequency distribution shows weights of pennies with a notable gap indicating two different sources:

    • Pennies made before 1983: 95% copper.

    • Pennies made after 1983: 2.5% copper.