Quantitative Data Distributions, Frequency Tables, Histograms, and Statistical Calculations

Course Logistics, Administrative Policies, and Exam Schedule

  • Chapter 2 Roadmap:

    • Section 2.1: Qualitative Data (Categorical organization, frequency/relative frequency tables, bar charts, pie charts).

    • Section 2.2: Quantitative Data (Numerical data organization, grouped frequency tables, histograms).

    • Section 2.3: Quantitative Data continued (Additional graphical displays and methods).

    • Section 2.4: Misleading and Deceptive Graphs (How visual representations can intentionally or accidentally cause misinterpretation).

  • Upcoming Schedule & Key Dates:

    • Monday: Section 2.2 completion.

    • Tuesday: Section 2.3 completion and potential start/completion of Section 2.4.

    • Wednesday: Dedicated working day in class for the Chapter 2 Written Assignment.

    • Thursday: Mandatory Exam Review Day. Covers exam structure, question types, and includes an mandatory in-class activity that provides direct support for the exam.

    • Following Tuesday (Tuesday after Labor Day): Exam covering Chapters 1 and 2.

  • Attendance Policy:

    • Syllabus Requirement: To receive full attendance credit for a class session, a student must be physically present for at least half of the total scheduled class time.

    • For an 8080-minute class period, the minimum threshold is 4040 minutes.

  • Participation Credit Policy:

    • Participation credit is evaluated using digital timestamps within Pearson MyLab and physical assignment submissions.

    • To earn full participation credit when work days or early releases are granted, students must satisfy all of the following criteria:

    • Complete all assigned online homework through the current section (e.g., Section 2.1).

    • Submit original written assignments or complete required corrections on returned written assignments.

    • Remain in class working for the required minimum time if assignments are incomplete.

  • Alternative Participation Option ("Ask the Instructor"):

    • Students who are uncomfortable raising their hands in class can utilize the Ask the Instructor feature in Pearson MyLab.

    • This tool generates an email directly to the instructor with the exact problem being worked on.

    • The instructor presents the problem in class anonymously and grants full participation credit for that day to the student who submitted it.

Fundamentals of Quantitative Data and Frequency Distributions

  • Definition of Quantitative Data:

    • Refers to data expressed in true numerical form representing objective scalar values.

    • Valid arithmetic operations (addition, subtraction, multiplication, division) can be performed on quantitative data.

    • Distinction from Qualitative Data displaying numbers: Categorical scale responses (such as a survey scale from 11 to 55 representing "least satisfied" to "most satisfied") display numbers but remain qualitative data because they represent ordinal categories rather than true numerical counts or measurements.

  • Terminology Comparison Between Qualitative and Quantitative Tables:

    • Qualitative Frequency Tables: Divide data into distinct categories.

    • Quantitative Frequency Tables: Divide data into distinct classes (or bins).

Discrete Quantitative Frequency Distributions

  • Ungrouped Frequency Distributions for Small Ranges:

    • When discrete quantitative data contains a small set of distinct integer values (e.g., values ranging only from 00 to 55), each unique number serves as its own class row.

    • Frequencies are calculated by counting the exact number of occurrences for each individual data value.

  • Data Counting & Verification Example (n=40n = 40 total survey responses):

    • Frequency of 00: 11

    • Frequency of 11: 1414

    • Frequency of 22: 1414

    • Frequency of 33: 88

    • Frequency of 44: 22

    • Frequency of 55: 11

    • Sum of Frequencies Verification Formula:     ∑f=f1+f2+f3+f4+f5+f6=n\sum f = f_1 + f_2 + f_3 + f_4 + f_5 + f_6 = n     1+14+14+8+2+1=401 + 14 + 14 + 8 + 2 + 1 = 40

    • Verifying that the sum of all class frequencies equals the sample size (n=40n = 40) ensures no data points were omitted or double-counted.

  • Relative Frequency Calculation:

    • Relative frequency represents the proportional occurrence of each class relative to the total dataset:     Relative Frequency=Class FrequencyTotal Sample Size=fn\text{Relative Frequency} = \frac{\text{Class Frequency}}{\text{Total Sample Size}} = \frac{f}{n}

    • Expressible as a fraction, decimal, or percentage.

    • Calculation for Class 00 (f=1f = 1, n=40n = 40):     Relative Frequency=140=0.025=2.5%\text{Relative Frequency} = \frac{1}{40} = 0.025 = 2.5\%

    • Calculation for Class 33 (f=8f = 8, n=40n = 40):     Relative Frequency=840=15=0.200=20.0%\text{Relative Frequency} = \frac{8}{40} = \frac{1}{5} = 0.200 = 20.0\%

    • Precision Rule: When directions require a percentage rounded to 11 decimal place, retain at least 33 decimal places in the intermediate decimal value.

Grouped Frequency Distributions for Continuous and Wide-Ranging Data

  • Rationale for Grouping (Bins/Classes):

    • Continuous data (containing decimal fractions) or discrete data spanning a wide numerical range (e.g., discrete integers from 00 to 2020 generating 2121 rows) makes individual value tables impractical and uninformative.

    • Grouping groups the range of numbers into contiguous, non-overlapping intervals called classes (or bins).

  • Key Definitions of Class Components:

    • Lower Class Limit (LCL): The smallest objective data value that can belong to a specific class interval.

    • Upper Class Limit (UCL): The largest objective data value that can belong to a specific class interval.

    • Class Width: The linear distance from the lower limit of one class to the lower limit of the immediately succeeding class (or upper limit to upper limit).     Class Width=LCLi+1−LCLi\text{Class Width} = \text{LCL}_{i+1} - \text{LCL}_i

    • Critical Distinction: Class width is measured vertically down table rows, never horizontally across a single class interval.

    • Class Midpoint: The exact central numerical value of a class interval.

    • Required Formula (Textbook / Formula Sheet Standard):       Midpointi=LCLi+LCLi+12\text{Midpoint}_i = \frac{\text{LCL}_i + \text{LCL}_{i+1}}{2}

    • Note on Alternative Sources: Non-textbook sources sometimes define midpoint as LCLi+UCLi2\frac{\text{LCL}_i + \text{UCL}_i}{2}. Students must use the textbook/formula sheet method to ensure correct homework and exam grading.

    • Class Boundaries: Continuous threshold values located precisely halfway in the gap separating the upper limit of one class and the lower limit of the next class.     Lower Boundaryi=UCLi−1+LCLi2\text{Lower Boundary}_i = \frac{\text{UCL}_{i-1} + \text{LCL}_i}{2}     Upper Boundaryi=UCLi+LCLi+12\text{Upper Boundary}_i = \frac{\text{UCL}_i + \text{LCL}_{i+1}}{2}

Comprehensive Case Study 1: Volume of Stock Traded (Continuous Data)

  • Dataset Context: Stock volume measured in millions of shares over 3535 random trading days (n=35n = 35). Data recorded to 22 decimal places (e.g., 6.426.42 represents 6.426.42 million shares).

  • Construction Parameters:

    • Target number of classes: 66

    • Initial Lower Class Limit (LCL1\text{LCL}_1): 6.006.00

    • Fixed Class Width: 3.003.00

  • Step-by-Step Table Construction Process:

    1. Determine Lower Class Limits (Add Class Width 3.003.00 vertically):

    • Class 1 LCL = 6.006.00

    • Class 2 LCL = 6.00+3.00=9.006.00 + 3.00 = 9.00

    • Class 3 LCL = 9.00+3.00=12.009.00 + 3.00 = 12.00

    • Class 4 LCL = 12.00+3.00=15.0012.00 + 3.00 = 15.00

    • Class 5 LCL = 15.00+3.00=18.0015.00 + 3.00 = 18.00

    • Class 6 LCL = 18.00+3.00=21.0018.00 + 3.00 = 21.00

    1. Determine Upper Class Limits (Step back by dataset precision 0.010.01 from the subsequent LCL to avoid overlap):

    • Class 1 UCL = 8.998.99

    • Class 2 UCL = 11.9911.99

    • Class 3 UCL = 14.9914.99

    • Class 4 UCL = 17.9917.99

    • Class 5 UCL = 20.9920.99

    • Class 6 UCL = 23.9923.99

    • Note: Class width (3.003.00) also applies vertically across upper limits (8.99+3.00=11.998.99 + 3.00 = 11.99).

    1. Grouped Frequency Distribution Summary Table:

     | Class Interval (Millions of Shares) | Frequency (ff) |      | :--- | :--- |      | 6.00−8.996.00 - 8.99 | 1515 |      | 9.00−11.999.00 - 11.99 | 99 |      | 12.00−14.9912.00 - 14.99 | 44 |      | 15.00−17.9915.00 - 17.99 | 44 |      | 18.00−20.9918.00 - 20.99 | 22 |      | 21.00−23.9921.00 - 23.99 | 11 |      | Total (∑f\sum f) | 3535 |

  • Handling High Values / Outliers:

    • If extreme values exceed initial limits (e.g., a data point of 34.0034.00), additional classes of width 3.003.00 must be added sequentially until all data points fit.

    • Intermediate classes with no data points receive a frequency of 00. Class width must remain strictly consistent across all intervals.

Comprehensive Case Study 2: Six-Year College Graduation Rates (Table Interpretation)

  • Dataset Context: Sample of U.S. four-year colleges and universities measuring their six-year graduation percentage rates.

  • Frequency Distribution Table:

  | Interval (Graduation Rate \%) | Frequency (ff) |   | :--- | :--- |   | 10−2410 - 24 | 11 |   | 25−3925 - 39 | 1010 |   | 40−5440 - 54 | 1313 |   | 55−6955 - 69 | 1717 |

  • Quantitative Analytical Questions & Solutions:

    • Total Sample Size (nn):     ∑f=1+10+13+17+⋯=50 schools\sum f = 1 + 10 + 13 + 17 + \dots = 50\text{ schools}

    • Class Width:     Class Width=25−10=15\text{Class Width} = 25 - 10 = 15     (Common Misconception Warning: Subtracting horizontally 24−10=1424 - 10 = 14 is incorrect).

    • Lower Class Limit of the 4th Interval: 5555

    • Upper Class Limit of the 2nd Interval: 3939

    • Class Boundaries for the 3rd Interval (40−5440 - 54):

    • Lower Boundary = Midpoint between 3939 and 4040:       Lower Boundary=39+402=39.5\text{Lower Boundary} = \frac{39 + 40}{2} = 39.5

    • Upper Boundary = Midpoint between 5454 and 5555:       Upper Boundary=54+552=54.5\text{Upper Boundary} = \frac{54 + 55}{2} = 54.5

    • Midpoint of the 1st Interval (10−2410 - 24) (Textbook Formula):     Midpoint1=10+252=17.5\text{Midpoint}_1 = \frac{10 + 25}{2} = 17.5

    • Relative Frequency of the 4th Interval (f=17f = 17, n=50n = 50):     Relative Frequency=1750=0.340=34.0%\text{Relative Frequency} = \frac{17}{50} = 0.340 = 34.0\%

    • Cumulative Frequency Below 40%40\% Graduation Rate:

    • Sum of frequencies for all classes strictly below 4040 (Interval 1 + Interval 2):       Cumulative Frequency=1+10=11 schools\text{Cumulative Frequency} = 1 + 10 = 11\text{ schools}

Visualizing Quantitative Data: Histograms

  • Definition & Mechanics:

    • A histogram is a bar graph designed specifically for quantitative data.

    • The horizontal axis displays continuous class ranges, and the vertical axis displays frequency or relative frequency.

    • Bar Contiguity: Unlike qualitative bar charts (which maintain spaces between bars), histogram bars must touch one another to represent continuous underlying numerical intervals.

  • Horizontal Axis Alignment Options:

    • Option A (Lower Limits): Left boundary of each bar aligns with the Lower Class Limit (e.g., 6.00,9.00,12.006.00, 9.00, 12.00).

    • Option B (Class Boundaries): Bar edges align with exact Class Boundaries (e.g., 9.5,24.5,39.5,54.59.5, 24.5, 39.5, 54.5), centering each bar over the Class Midpoint.

Classifying Shapes of Data Distributions

  • Uniform Distribution:

    • Frequencies across all class intervals are approximately equal.

    • The histogram displays a flat, rectangular horizontal shape across the top.

  • Bell-Shaped / Normal / Symmetric Distribution:

    • Frequencies start low, rise to a central maximum peak, and taper off symmetrically on both sides.

    • In statistics, perfect or approximate bell-shaped symmetric distributions are referred to as normally distributed data.

  • Skewed Right (Positively Skewed):

    • The majority of high-frequency data is concentrated on the left side (lower values).

    • The tail of the graph extends outward to the right (higher values).

    • Rule: Skewness follows the direction of the tail.

    • Associated with potential high-side outliers.

  • Skewed Left (Negatively Skewed):

    • The majority of high-frequency data is concentrated on the right side (higher values).

    • The tail of the graph extends outward to the left (lower values).

    • Associated with potential low-side outliers.

Statistical Software Integration: StatCrunch Overview

  • Platform Context:

    • StatCrunch is a web-based statistical software package hosted by Pearson at statcrunch.com.

    • Sign-in utilizes standard Pearson MyLab credentials.

  • Functional Capabilities:

    • Automates frequency distributions, relative frequency tables, midpoints, and histograms for quantitative datasets.

    • Replaces manual counting and drawing for complex decimal datasets.

    • Permitted for use during in-class exams via dedicated laboratory desktop computers.

Questions & Discussion

  • Question: Can students leave early if homework is complete?

    • Answer: Early departure is only permitted if online homework through the current section (Section 2.1) is completely finished AND written assignment submissions/corrections are complete. Students must meet the mandatory 4040-minute attendance requirement to be marked present.

  • Question: How do you extend a frequency table if a dataset has an outlier far larger than the planned classes?

    • Answer: Append additional class rows using the fixed class width sequentially until the maximum value is encompassed. Intermediate empty classes are assigned a frequency of 00.

  • Question: Why is the class width for an interval of 10−2410 - 24 equal to 1515 rather than 1414?

    • Answer: Class width is computed vertically by measuring the distance between consecutive lower limits (25−10=1525 - 10 = 15), not horizontally across a single row (24−10=1424 - 10 = 14).