Statistics Study Guide: Daily Number of Emails and Sampling Techniques
Statistics Exam 1: Study Guide for Chapters 1-3
Examination Overview
Calculator Policy: Permitted for all sections, but specific problems require calculations/work shown "by hand."
Scoring Criteria:
Correct answers with inconsistent work may not receive full credit.
Incorrect answers with appropriate work may receive partial credit.
Final answers must be circled and labeled with the correct variables and correct units where applicable.
Total Points Possible:
Chapter 1: Basic Statistical Concepts and Sampling
1. Data Classification and Calculations
Case Study: Daily Emails at OCC
Data Set: The study observes the daily number of emails received by OCC students in their OCC email account. A sample of 5 results consists of: emails.
Mean Calculation (By Hand): (Note: Per the transcript key, the calculation uses the value 7 instead of 6 to arrive at 3 emails).
Statistic vs. Parameter:
The mean calculated above is a statistic.
Explanation: It is derived from a sample rather than a full population.
Median Calculation (By Hand):
Ordered Data:
Median value:
Standard Deviation Calculation (By Hand):
2. Sampling Techniques and Design
Non-Random Sampling Techniques
Stratified Sampling: Surveying exactly fans from specific age groups (, , , , and years old).
Cluster Sampling: Surveying all participants within a specific group, such as surveying fans from only the age group.
Random Sampling Evaluation
Scenario: Randomly picking male and female fans for a survey.
Is it a random sample? No.
Reasoning: This is only a random sample if there are exactly the same number of male and female fans in the population. Since this is highly unlikely, it does not meet the definition of a random sample.
Simple Random Samples (SRS)
Scenario: A teacher puts stickers on the bottom of chairs out of in a room. Students enter and sit; those on sticker chairs win prizes.
Is it an SRS? Yes.
Reasoning: Every possible group of children is equally likely to be selected.
3. Critiquing Data Collection
Scenario: A student asks his sisters (all attending OCC) which course students perform best in at OCC.
Flaws in Sampling Process:
Small Sample: The sample size is too limited to be representative.
Does Not Reflect Population: Asking only siblings (and likely people with similar backgrounds or majors) introduces bias and fails to represent the wider student population.
Reported vs. Collected: The data relies on reported opinions rather than measured academic performance data.
Chapter 2: Describing, Exploring, and Comparing Data
4. TikTok Usage Case Study
Data Set (15 Days of Minutes):
Sorted Data:
Mode:
Range: , which is .
Calculated Statistics (via STAT function):
Mean:
Standard Deviation:
Range Rule of Thumb (Usual/Typical Values): The typical range is between and .
Percentile Calculations:
Finding Percentile of a Value (32 minutes): . Calculated as .
Finding P15 (15th Percentile): Round up to the value.
5. Graphical Representations
Frequency Distribution Chart (TikTok Data)
Daily Minutes | # Days (Frequency) | Relative Frequency (%) |
|---|---|---|
Relative Frequency Histogram Details:
X-axis: Boundaries set at .
Y-axis: Percentage of days (, , , ).
Distribution Description: The distribution is described as "a little normal/symmetric" but ultimately "skewed Right."
Frequency Polygon:
The graph uses midpoints for plotting: .
Anchored at endpoints: and to show frequency.
Ineffective Graphs:
Types of graphs that would not work well for this TikTok data:
Dot Plot
Time Series
Pareto Chart
6. Misleading Graphs
Example Analysis: A study of high school seniors asked about graduation plans ( College, Work).
Construction Errors: The y-axis did not start at , and the scale was not consistent.
Deception Result: It exaggerates differences. Viewers might think college attendance is to times higher than joining the workforce, when it is actually less than double.
Chapter 3: Probability and Normal Distributions
7. Normal Distribution Problems
Scenario: Time spent on a math test is normally distributed with and .
Middle 67% Range: Using the empirical rule approximation ( standard deviation): .
Z-score Interpretation (Case: Tom):
Tom's z-score () = .
Calculation: .
Probability/Percentage (Case: More than 116 minutes):
is exactly standard deviations above the mean.
Using the empirical rule, is within . The remainder in the tails is . The upper tail (more than ) is .
Z-score Calculation (Case: Kirsten):
Kirsten’s time () = .
.
8. Identification of Data Types
Discrete: Number of hairs (countable values).
Continuous: Length of hair (measurable values along a continuum).
Nominal: Ethnicity (categories with no inherent order).
Interval: Year in history (ordered with meaningful differences but no true zero).
Ratio: Height of people (ordered, meaningful differences, and a true zero point).