Quantitative Data Distributions, Frequency Tables, Histograms, and Statistical Calculations
Course Logistics, Administrative Policies, and Exam Schedule
Chapter 2 Roadmap:
Section 2.1: Qualitative Data (Categorical organization, frequency/relative frequency tables, bar charts, pie charts).
Section 2.2: Quantitative Data (Numerical data organization, grouped frequency tables, histograms).
Section 2.3: Quantitative Data continued (Additional graphical displays and methods).
Section 2.4: Misleading and Deceptive Graphs (How visual representations can intentionally or accidentally cause misinterpretation).
Upcoming Schedule & Key Dates:
Monday: Section 2.2 completion.
Tuesday: Section 2.3 completion and potential start/completion of Section 2.4.
Wednesday: Dedicated working day in class for the Chapter 2 Written Assignment.
Thursday: Mandatory Exam Review Day. Covers exam structure, question types, and includes an mandatory in-class activity that provides direct support for the exam.
Following Tuesday (Tuesday after Labor Day): Exam covering Chapters 1 and 2.
Attendance Policy:
Syllabus Requirement: To receive full attendance credit for a class session, a student must be physically present for at least half of the total scheduled class time.
For an -minute class period, the minimum threshold is minutes.
Participation Credit Policy:
Participation credit is evaluated using digital timestamps within Pearson MyLab and physical assignment submissions.
To earn full participation credit when work days or early releases are granted, students must satisfy all of the following criteria:
Complete all assigned online homework through the current section (e.g., Section 2.1).
Submit original written assignments or complete required corrections on returned written assignments.
Remain in class working for the required minimum time if assignments are incomplete.
Alternative Participation Option ("Ask the Instructor"):
Students who are uncomfortable raising their hands in class can utilize the Ask the Instructor feature in Pearson MyLab.
This tool generates an email directly to the instructor with the exact problem being worked on.
The instructor presents the problem in class anonymously and grants full participation credit for that day to the student who submitted it.
Fundamentals of Quantitative Data and Frequency Distributions
Definition of Quantitative Data:
Refers to data expressed in true numerical form representing objective scalar values.
Valid arithmetic operations (addition, subtraction, multiplication, division) can be performed on quantitative data.
Distinction from Qualitative Data displaying numbers: Categorical scale responses (such as a survey scale from to representing "least satisfied" to "most satisfied") display numbers but remain qualitative data because they represent ordinal categories rather than true numerical counts or measurements.
Terminology Comparison Between Qualitative and Quantitative Tables:
Qualitative Frequency Tables: Divide data into distinct categories.
Quantitative Frequency Tables: Divide data into distinct classes (or bins).
Discrete Quantitative Frequency Distributions
Ungrouped Frequency Distributions for Small Ranges:
When discrete quantitative data contains a small set of distinct integer values (e.g., values ranging only from to ), each unique number serves as its own class row.
Frequencies are calculated by counting the exact number of occurrences for each individual data value.
Data Counting & Verification Example ( total survey responses):
Frequency of :
Frequency of :
Frequency of :
Frequency of :
Frequency of :
Frequency of :
Sum of Frequencies Verification Formula:
Verifying that the sum of all class frequencies equals the sample size () ensures no data points were omitted or double-counted.
Relative Frequency Calculation:
Relative frequency represents the proportional occurrence of each class relative to the total dataset:
Expressible as a fraction, decimal, or percentage.
Calculation for Class (, ):
Calculation for Class (, ):
Precision Rule: When directions require a percentage rounded to decimal place, retain at least decimal places in the intermediate decimal value.
Grouped Frequency Distributions for Continuous and Wide-Ranging Data
Rationale for Grouping (Bins/Classes):
Continuous data (containing decimal fractions) or discrete data spanning a wide numerical range (e.g., discrete integers from to generating rows) makes individual value tables impractical and uninformative.
Grouping groups the range of numbers into contiguous, non-overlapping intervals called classes (or bins).
Key Definitions of Class Components:
Lower Class Limit (LCL): The smallest objective data value that can belong to a specific class interval.
Upper Class Limit (UCL): The largest objective data value that can belong to a specific class interval.
Class Width: The linear distance from the lower limit of one class to the lower limit of the immediately succeeding class (or upper limit to upper limit).
Critical Distinction: Class width is measured vertically down table rows, never horizontally across a single class interval.
Class Midpoint: The exact central numerical value of a class interval.
Required Formula (Textbook / Formula Sheet Standard):
Note on Alternative Sources: Non-textbook sources sometimes define midpoint as . Students must use the textbook/formula sheet method to ensure correct homework and exam grading.
Class Boundaries: Continuous threshold values located precisely halfway in the gap separating the upper limit of one class and the lower limit of the next class.
Comprehensive Case Study 1: Volume of Stock Traded (Continuous Data)
Dataset Context: Stock volume measured in millions of shares over random trading days (). Data recorded to decimal places (e.g., represents million shares).
Construction Parameters:
Target number of classes:
Initial Lower Class Limit ():
Fixed Class Width:
Step-by-Step Table Construction Process:
Determine Lower Class Limits (Add Class Width vertically):
Class 1 LCL =
Class 2 LCL =
Class 3 LCL =
Class 4 LCL =
Class 5 LCL =
Class 6 LCL =
Determine Upper Class Limits (Step back by dataset precision from the subsequent LCL to avoid overlap):
Class 1 UCL =
Class 2 UCL =
Class 3 UCL =
Class 4 UCL =
Class 5 UCL =
Class 6 UCL =
Note: Class width () also applies vertically across upper limits ().
Grouped Frequency Distribution Summary Table:
| Class Interval (Millions of Shares) | Frequency () | | :--- | :--- | | | | | | | | | | | | | | | | | | | | Total () | |
Handling High Values / Outliers:
If extreme values exceed initial limits (e.g., a data point of ), additional classes of width must be added sequentially until all data points fit.
Intermediate classes with no data points receive a frequency of . Class width must remain strictly consistent across all intervals.
Comprehensive Case Study 2: Six-Year College Graduation Rates (Table Interpretation)
Dataset Context: Sample of U.S. four-year colleges and universities measuring their six-year graduation percentage rates.
Frequency Distribution Table:
| Interval (Graduation Rate \%) | Frequency () | | :--- | :--- | | | | | | | | | | | | |
Quantitative Analytical Questions & Solutions:
Total Sample Size ():
Class Width: (Common Misconception Warning: Subtracting horizontally is incorrect).
Lower Class Limit of the 4th Interval:
Upper Class Limit of the 2nd Interval:
Class Boundaries for the 3rd Interval ():
Lower Boundary = Midpoint between and :
Upper Boundary = Midpoint between and :
Midpoint of the 1st Interval () (Textbook Formula):
Relative Frequency of the 4th Interval (, ):
Cumulative Frequency Below Graduation Rate:
Sum of frequencies for all classes strictly below (Interval 1 + Interval 2):
Visualizing Quantitative Data: Histograms
Definition & Mechanics:
A histogram is a bar graph designed specifically for quantitative data.
The horizontal axis displays continuous class ranges, and the vertical axis displays frequency or relative frequency.
Bar Contiguity: Unlike qualitative bar charts (which maintain spaces between bars), histogram bars must touch one another to represent continuous underlying numerical intervals.
Horizontal Axis Alignment Options:
Option A (Lower Limits): Left boundary of each bar aligns with the Lower Class Limit (e.g., ).
Option B (Class Boundaries): Bar edges align with exact Class Boundaries (e.g., ), centering each bar over the Class Midpoint.
Classifying Shapes of Data Distributions
Uniform Distribution:
Frequencies across all class intervals are approximately equal.
The histogram displays a flat, rectangular horizontal shape across the top.
Bell-Shaped / Normal / Symmetric Distribution:
Frequencies start low, rise to a central maximum peak, and taper off symmetrically on both sides.
In statistics, perfect or approximate bell-shaped symmetric distributions are referred to as normally distributed data.
Skewed Right (Positively Skewed):
The majority of high-frequency data is concentrated on the left side (lower values).
The tail of the graph extends outward to the right (higher values).
Rule: Skewness follows the direction of the tail.
Associated with potential high-side outliers.
Skewed Left (Negatively Skewed):
The majority of high-frequency data is concentrated on the right side (higher values).
The tail of the graph extends outward to the left (lower values).
Associated with potential low-side outliers.
Statistical Software Integration: StatCrunch Overview
Platform Context:
StatCrunch is a web-based statistical software package hosted by Pearson at
statcrunch.com.Sign-in utilizes standard Pearson MyLab credentials.
Functional Capabilities:
Automates frequency distributions, relative frequency tables, midpoints, and histograms for quantitative datasets.
Replaces manual counting and drawing for complex decimal datasets.
Permitted for use during in-class exams via dedicated laboratory desktop computers.
Questions & Discussion
Question: Can students leave early if homework is complete?
Answer: Early departure is only permitted if online homework through the current section (Section 2.1) is completely finished AND written assignment submissions/corrections are complete. Students must meet the mandatory -minute attendance requirement to be marked present.
Question: How do you extend a frequency table if a dataset has an outlier far larger than the planned classes?
Answer: Append additional class rows using the fixed class width sequentially until the maximum value is encompassed. Intermediate empty classes are assigned a frequency of .
Question: Why is the class width for an interval of equal to rather than ?
Answer: Class width is computed vertically by measuring the distance between consecutive lower limits (), not horizontally across a single row ().