Statistics Study Guide: Daily Number of Emails and Sampling Techniques

Statistics Exam 1: Study Guide for Chapters 1-3

Examination Overview

  • Calculator Policy: Permitted for all sections, but specific problems require calculations/work shown "by hand."

  • Scoring Criteria:

    • Correct answers with inconsistent work may not receive full credit.

    • Incorrect answers with appropriate work may receive partial credit.

    • Final answers must be circled and labeled with the correct variables and correct units where applicable.

  • Total Points Possible: 150150

Chapter 1: Basic Statistical Concepts and Sampling

1. Data Classification and Calculations

Case Study: Daily Emails at OCC

  • Data Set: The study observes the daily number of emails received by OCC students in their OCC email account. A sample of 5 results consists of: 1,0,5,6, and 21, 0, 5, 6, \text{ and } 2 emails.

  • Mean Calculation (By Hand):     xˉ=1+0+5+7+25=3emails\bar{x} = \frac{1 + 0 + 5 + 7 + 2}{5} = 3\,\text{emails}     (Note: Per the transcript key, the calculation uses the value 7 instead of 6 to arrive at 3 emails).

  • Statistic vs. Parameter:

    • The mean calculated above is a statistic.

    • Explanation: It is derived from a sample rather than a full population.

  • Median Calculation (By Hand):

    • Ordered Data: 0,1,2,5,60, 1, 2, 5, 6

    • Median value: 2emails2\,\text{emails}

  • Standard Deviation Calculation (By Hand):     s=(13)2+(03)2+(53)2+(73)2+(23)251s = \sqrt{\frac{(1 - 3)^2 + (0 - 3)^2 + (5 - 3)^2 + (7 - 3)^2 + (2 - 3)^2}{5 - 1}}     s=4+9+4+16+14s = \sqrt{\frac{4 + 9 + 4 + 16 + 1}{4}}     s=344s = \sqrt{\frac{34}{4}}     s=8.5s = \sqrt{8.5}     s2.915emailss \approx 2.915\,\text{emails}

2. Sampling Techniques and Design

Non-Random Sampling Techniques

  • Stratified Sampling: Surveying exactly 5050 fans from specific age groups (1-201\text{-}20, 21-4021\text{-}40, 41-6041\text{-}60, 61-8061\text{-}80, and 81-10081\text{-}100 years old).

  • Cluster Sampling: Surveying all participants within a specific group, such as surveying 200200 fans from only the 21-4021\text{-}40 age group.

Random Sampling Evaluation

  • Scenario: Randomly picking 5050 male and 5050 female fans for a survey.

  • Is it a random sample? No.

  • Reasoning: This is only a random sample if there are exactly the same number of male and female fans in the population. Since this is highly unlikely, it does not meet the definition of a random sample.

Simple Random Samples (SRS)

  • Scenario: A teacher puts stickers on the bottom of 55 chairs out of 2525 in a room. Students enter and sit; those on sticker chairs win prizes.

  • Is it an SRS? Yes.

  • Reasoning: Every possible group of 55 children is equally likely to be selected.

3. Critiquing Data Collection

Scenario: A student asks his 33 sisters (all attending OCC) which course students perform best in at OCC.

Flaws in Sampling Process:

  1. Small Sample: The sample size is too limited to be representative.

  2. Does Not Reflect Population: Asking only siblings (and likely people with similar backgrounds or majors) introduces bias and fails to represent the wider student population.

  3. Reported vs. Collected: The data relies on reported opinions rather than measured academic performance data.

Chapter 2: Describing, Exploring, and Comparing Data

4. TikTok Usage Case Study

Data Set (15 Days of Minutes): 24,36,22,31,41,22,4,47,0,35,27,24,24,14,3524, 36, 22, 31, 41, 22, 4, 47, 0, 35, 27, 24, 24, 14, 35

Sorted Data: 0,7,14,22,22,24,24,24,27,31,33,35,36,41,470, 7, 14, 22, 22, 24, 24, 24, 27, 31, 33, 35, 36, 41, 47

  • Mode: 24min24\,\text{min}

  • Range: [0,47][0, 47], which is 47minutes47\,\text{minutes}.

  • Calculated Statistics (via STAT function):

    • Mean: xˉ=25.6min\bar{x} = 25.6\,\text{min}

    • Standard Deviation: s=12.7mins = 12.7\,\text{min}

  • Range Rule of Thumb (Usual/Typical Values):     xˉ±2s25.6±(12.7×2)\bar{x} \pm 2s \Rightarrow 25.6 \pm (12.7 \times 2)     The typical range is between 22 and 51minutes51\,\text{minutes}.

  • Percentile Calculations:

    • Finding Percentile of a Value (32 minutes): L=1015=0.6=66.5%L = \frac{10}{15} = 0.6 = 66.5\%. Calculated as P67=32minP_{67} = 32\,\text{min}.

    • Finding P15 (15th Percentile):         L=15100×15=2.25L = \frac{15}{100} \times 15 = 2.25         Round up to the 3rd3\text{rd} value.         P15=14minP_{15} = 14\,\text{min}

5. Graphical Representations

Frequency Distribution Chart (TikTok Data)

Daily Minutes

# Days (Frequency)

Relative Frequency (%)

0-90\text{-}9

22

13.3%13.3\%

10-1910\text{-}19

11

6.6%6.6\%

20-2920\text{-}29

66

40%40\%

30-3930\text{-}39

44

26.5%26.5\%

40-4940\text{-}49

22

13.3%13.3\%

Relative Frequency Histogram Details:

  • X-axis: Boundaries set at 9.5,19.5,29.5,39.5,49.59.5, 19.5, 29.5, 39.5, 49.5.

  • Y-axis: Percentage of days (10%10\%, 20%20\%, 30%30\%, 40%40\%).

  • Distribution Description: The distribution is described as "a little normal/symmetric" but ultimately "skewed Right."

Frequency Polygon:

  • The graph uses midpoints for plotting: 4.5,14.5,24.5,34.5,44.54.5, 14.5, 24.5, 34.5, 44.5.

  • Anchored at endpoints: 5.5-5.5 and 54.554.5 to show 00 frequency.

Ineffective Graphs:

  • Types of graphs that would not work well for this TikTok data:

    • Dot Plot

    • Time Series

    • Pareto Chart

6. Misleading Graphs

Example Analysis: A study of 10,00010,000 high school seniors asked about graduation plans (61%61\% College, 32%32\% Work).

  • Construction Errors: The y-axis did not start at 00, and the scale was not consistent.

  • Deception Result: It exaggerates differences. Viewers might think college attendance is 44 to 55 times higher than joining the workforce, when it is actually less than double.

Chapter 3: Probability and Normal Distributions

7. Normal Distribution Problems

Scenario: Time spent on a math test is normally distributed with μ=88minutes\mu = 88\,\text{minutes} and σ=14minutes\sigma = 14\,\text{minutes}.

  • Middle 67% Range: Using the empirical rule approximation (11 standard deviation):     88±1474to102minutes88 \pm 14 \Rightarrow 74\,\text{to}\,102\,\text{minutes}.

  • Z-score Interpretation (Case: Tom):

    • Tom's z-score (zz) = 1.81.8.

    • Calculation: 88+(1.8×14)=113.2minutes88 + (1.8 \times 14) = 113.2\,\text{minutes}.

  • Probability/Percentage (Case: More than 116 minutes):

    • 116minutes116\,\text{minutes} is exactly 22 standard deviations above the mean.

    • Using the empirical rule, 95%95\% is within 2σ2\sigma. The remainder in the tails is 5%5\%. The upper tail (more than 2σ2\sigma) is 2.5%2.5\%.

  • Z-score Calculation (Case: Kirsten):

    • Kirsten’s time (xx) = 75minutes75\,\text{minutes}.

    • z=758814=0.93z = \frac{75 - 88}{14} = -0.93.

8. Identification of Data Types
  • Discrete: Number of hairs (countable values).

  • Continuous: Length of hair (measurable values along a continuum).

  • Nominal: Ethnicity (categories with no inherent order).

  • Interval: Year in history (ordered with meaningful differences but no true zero).

  • Ratio: Height of people (ordered, meaningful differences, and a true zero point).