Comprehensive Introduction to Elementary Statistics and Methodology

Course Overview and Goals of Statistics

  • Purpose of Studying Statistics: Statistical information is encountered continuously throughout daily life, including advertisements, news reports, and workplace data. Understanding statistical techniques allows for thoughtful, critical analysis of this information rather than blind acceptance.

  • Definition of Statistics: Statistics is the science of performing four interconnected tasks with data:

    • Collecting data

    • Organizing data

    • Analyzing data

    • Interpreting data

    • Ultimate Goal: To gain a thorough understanding of data in order to make informed decisions.

  • Two Primary Branches of Statistics:

    • Descriptive Statistics: The branch of statistics focused on organizing, summarizing, and presenting data in a meaningful way.

    • Inferential Statistics: The branch of statistics focused on drawing formal conclusions, making predictions, or inferring properties about a larger group based on sample data.

  • Role of Technology vs. Human Analysis: Calculators and computer software programs are tools used to execute calculations and generate statistics. However, the critical understanding of conclusions, contextual evaluation, and decision-making relies entirely on the human researcher.

  • Case Study — "Your Baby Can Read":

    • Program Claims: Advertisements claim that using flashcards, pop-up books, and focused parent-child quality time enables almost any preschooler to learn to read before entering kindergarten.

    • Critical Analysis: Parents should not blindly purchase this program because it lacks rigorous, peer-reviewed statistical studies to validate its efficacy. Furthermore, multiple confounding variables affect early childhood literacy.

    • Illustrative Personal Examples of Confounding Factors in Learning:

    • Family Literacy Factors: A family homeschooled five children through the sixth grade. The oldest daughter taught herself to read fluently during kindergarten after sitting down for just a couple of basic book sessions without formal reading lessons. In contrast, her twin sisters suffered from toddler ear infections that hindered their speech development and subsequent reading acquisition; they did not read well until the third grade.

    • Individual Disposition Factors: Among five grandchildren, a four-year-old granddaughter reads fluently simply due to natural predisposition and parental engagement without specialized commercial programs. Conversely, a two-year-old grandson demonstrates high intelligence but possesses an active personality, leaning naturally toward hands-on physical activities rather than reading.

    • Takeaway: Researchers and consumers must critically evaluate data gathering, acknowledge confounding factors, and refrain from uncritically accepting claims.

Fundamentals of Probability

  • Definition of Probability: Probability is a mathematical tool used to quantify and study randomness, dealing directly with the likelihood of a specific event occurring.

  • Origins and Applications: The study of probability originated from games of chance (e.g., card games, poker, dice, roulette). Today, probability is applied broadly across fields such as weather forecasting (e.g., chance of rain) and sports predictions.

  • Core Probability Formula: The probability of an event occurring is defined as the ratio of the number of favorable outcomes to the total number of possible outcomes:   Probability=Number of ways an event can occurTotal number of possible outcomes\text{Probability} = \frac{\text{Number of ways an event can occur}}{\text{Total number of possible outcomes}}

  • Concrete Examples:

    • Drawing a Specific Card: The probability of drawing a 22 of Hearts from a standard deck of 5252 cards is 152\frac{1}{52}, as there is only 11 two of Hearts out of 5252 possible cards.

    • Rolling a Die: The probability of rolling a 55 on a standard six-sided die is 16\frac{1}{6}, as there is only 11 five out of 66 possible outcomes.

  • Mathematical Representations: Probability predictions are expressed interchangeably as fractions, decimals, or percentages.

Mathematical Conversions: Percents, Decimals, and Fractions

  • Interrelationships: Percents, decimals, fractions, and division problems are different mathematical representations of the exact same values:

    • A percentage is a decimal expressed relative to 100100

    • A decimal is a fractional value in base-1010 notation

    • A fraction is a literal representation of a division problem

  • Computational Rules: When performing statistical calculations or plugging values into formulas, percentages must be converted into decimals first. When asked to report a final answer as a percentage, the resulting decimal must be converted back to a percent.

  • Percent Definition: The word percent originates from per (meaning part) and cent (meaning century or 100100). It represents a part out of 100100.

  • Converting Percent to Decimal:

    • Method 1: Divide the percentage value by 100100.

    • Method 2: Shift the decimal point two places to the left.

    • Example: 50%50÷100=0.550\% \rightarrow 50 \div 100 = 0.5

  • Converting Decimal to Percent:

    • Method 1: Multiply the decimal value by 100100

    • Method 2: Shift the decimal point two places to the right.

    • Example: 0.200.20×100=20%0.20 \rightarrow 0.20 \times 100 = 20\%

  • Converting Fractions, Decimals, and Percents:

    • Structure of a Fraction: Composed of a Numerator\text{Numerator} (the part) over a Denominator\text{Denominator} (the whole):     Fraction=NumeratorDenominator=PartWhole\text{Fraction} = \frac{\text{Numerator}}{\text{Denominator}} = \frac{\text{Part}}{\text{Whole}}

    • Reducing Fractions: Simplify by dividing out common factors. For numbers ending in zeros, equal numbers of trailing zeros can be canceled from the numerator and denominator:     50100=510=12\frac{50}{100} = \frac{5}{10} = \frac{1}{2}

    • Converting Decimal to Fraction:

    1. Place all numbers to the right of the decimal point into the numerator.

    2. Place a 11 followed by as many zeros as there are decimal digits into the denominator.

    3. Simplify/reduce the resulting fraction.

    • Example 1: 0.2210=150.2 \rightarrow \frac{2}{10} = \frac{1}{5}

    • Example 2: 0.33331000.33 \rightarrow \frac{33}{100} (cannot be further reduced)

    • Converting Fraction to Decimal to Percent:

    • Divide the numerator by the denominator to obtain a decimal, then move the decimal point two places to the right to obtain the percentage.

    • Example: For 23\frac{2}{3}, perform 2÷3=0.666666...2 \div 3 = 0.666666...

      • Rounded to three decimal places: 0.6670.667

      • Converted to percentage: 66.7%66.7\%

    • Rounding Terms: Rounding to the "hundredths place" means rounding to two decimal places (since 100100 has two zeros).

Key Statistical Terminology: Populations, Samples, Parameters, and Statistics

  • Population: The complete, total collection of all individuals, items, objects, or measurements under study (e.g., all students enrolled at UTC).

  • Parameter: A numerical value that describes a specific characteristic or property of an entire Population.

    • Mnemonic: Parameter aligns with Population.

    • Examples: The exact percentage of all UTC students who are full-time, commuters, involved in extracurriculars, or employed.

  • Sample: A smaller subset or subgroup selected from the population, studied to gain insights and infer conclusions about the total population.

  • Statistic: A numerical value that describes a specific characteristic or property of a Sample.

    • Mnemonic: Statistic aligns with Sample.

  • Variable: A characteristic or attribute of interest for each subject or object in a population, traditionally denoted by a capital letter (such as XX).

  • Practice Scenarios for Parameter vs. Statistic:

    • Scenario A: It is reported that 84.9%84.9\% of all students on campus have a job.

    • Classification: Population = All students on campus; Parameter = 84.9%84.9\%

    • Scenario B: A sample of 250250 students is surveyed, revealing that 86.4%86.4\% have a job.

    • Classification: Sample = 250250 students; Statistic = 86.4%86.4\%

Data Classifications: Qualitative vs. Quantitative

  • Qualitative (Categorical) Data: Data consisting of attributes, labels, quality descriptions, or categories.

    • Examples: Vehicle make and model, hair colors, birth order, nationality, level of education (e.g., Associate degree, Bachelor's degree), and zip codes.

    • Note on Zip Codes: Although zip codes consist of numbers (e.g., 3732137321, 3700337003), they are categorical data because they function as regional labels rather than mathematical quantities.

  • Quantitative (Numerical) Data: Data consisting of numerical measurements or counts.

    • Examples: Number of siblings, annual salary, household income, daily whole grain intake measured in grams per day.

  • Data: The actual observed values collected for a variable (e.g., specific dollar amounts like $40,000\$40,000, $35,000\$35,000, $55,000\$55,000).

    • Grammar Note: Data is a plural noun. The singular form is datum.

  • Detailed Case Example — San Jacinto College Study:

    • Objective: Determine the mean (average) amount of money first-year college students at San Jacinto College spend on school supplies, excluding textbooks.

    • Study Details: A random survey of 100100 first-year students at the college is conducted. Three individual student responses recorded were $150\$150, $200\$200, and $225\$225.

    • Component Breakdown:

    • Population: All first-year college students at San Jacinto College.

    • Sample: The 100100 first-year students surveyed.

    • Variable (XX): Total dollar cost of school supplies (excluding textbooks) spent by a first-year student.

    • Variable Type: Quantitative (Numerical).

    • Data Points: $150\$150, $200\$200, $225\$225

Sub-Classifications of Quantitative Data: Discrete vs. Continuous

  • Quantitative Discrete Variables: Quantitative data resulting from values that can be counted individually (0,1,2,3...0, 1, 2, 3...). Discrete data cannot contain fractional parts between countable units in real-world contexts.

    • Examples: Number of children in a family, number of cars owned by a household, number of students present in a classroom, household income from the prior year (typically reported in discrete dollar increments such as $50,000\$50,000 or $45,000\$45,000).

  • Quantitative Continuous Variables: Quantitative data resulting from an infinite continuum of possible values obtained through measurement rather than counting.

    • Key Identifier: Presence of measurement units (e.g., weight, height, distance, time).

    • Example: Daily intake of whole grains measured in grams per day.

  • Comprehensive Grocery Store Classification Exercise:

    • Purchases Made: Three cans of soup (19oz19\,oz tomato bisque, 14.1oz14.1\,oz lentil, 19oz19\,oz Italian wedding); two packages of nuts (walnuts, peanuts); four types of vegetables (broccoli, cauliflower, spinach, carrots); two desserts (16oz16\,oz pistachio ice cream, 32oz32\,oz chocolate chip cookies).

    • Categorization of Data Sets from Shopping Trip:

    • Quantitative Discrete: Count of soup cans (33), count of vegetable types (44), count of dessert packages (22), count of nut packages (22).

    • Quantitative Continuous: Measured weights of items, such as soup weight (19oz19\,oz, 14.1oz14.1\,oz) and dessert weight (16oz16\,oz, 32oz32\,oz).

    • Qualitative (Categorical): Descriptive identities/names of the products purchased (e.g., tomato bisque, lentil, walnuts, peanuts, broccoli, spinach, pistachio ice cream).

Data Visualization and Presentation

  • Tables: Effective for organizing raw data, but often inadequate for visually conveying proportions or comparisons.

  • Pie Charts: Circular charts divided into triangular wedges representing categorical data.

    • Proportionality: Each wedge area is strictly proportional to the percentage of total items in that category.

    • Required Visual Elements: Every pie chart must include a clear descriptive title, explicit category labels/percentages, and a color-coded key/legend.

    • Example: A comparison between De Anza College and Foothill College student enrollment status. At Foothill College, 71.4%71.4\% of students attend part-time; the corresponding yellow slice represents exactly 71.4%71.4\% of the total circle area.

  • Bar Graphs: Displays rectangular bars where the length or height of each bar corresponds directly to the numerical frequency or percentage of categorical data, facilitating direct comparison between distinct groups.

Sampling Techniques and Methodologies

  • Notation Conventions:

    • NN = Population size (capital letter)

    • nn = Sample size (lowercase letter)

  • Random Sampling: The overarching selection process using chance to choose individuals from a population to participate in a sample.

  • Simple Random Sample (SRS): A sampling procedure structured such that every possible sample of size nn has an equal probability of being chosen from a population of size NN.

  • Steps to Conduct Simple Random Sampling:

    1. Assign a unique sequential integer label (11 to NN) to every individual item in the target population.

    2. Use a random number generator (calculator, computer algorithm, or random digit table) to select nn distinct numbers.

  • SRS Demonstration Example (5-Student Study Group):

    • Group Members: Bob (11), Patricia (22), Mike (33), Jan (44), Maria (55).

    • Task: Randomly select 22 students (n=2n = 2) to solve homework problems on the board.

    • Execution: Generating random pairs from 11 to 55:

    • First iteration produced 55 and 44 \rightarrow Sample consists of Maria and Jan.

    • Second iteration produced 44 and 33 \rightarrow Sample consists of Jan and Mike.

  • Four Major Alternative Sampling Methods:

    • Stratified Sampling: The population is partitioned into non-overlapping subgroups (strata) based on specific characteristics. A random sample of proportional size is drawn from every individual subgroup.

    • Example: Selecting a sample of 5050 students from each distinct grade level (Freshmen, Sophomores, Juniors, Seniors) in a high school.

    • Cluster Sampling: The population is partitioned into distinct subgroups (clusters), often geographically based. Entire clusters are randomly selected, and all members within those chosen clusters are surveyed.

    • Distinction: Stratified samples take a few individuals from all groups; Cluster samples take all individuals from a few selected groups.

    • Example: Partitioning a municipality into city blocks, randomly selecting 1010 city blocks, and surveying every voter living within those 1010 blocks.

    • Systematic Sampling: Starting from a random point, every kthk^{\text{th}} individual in an ordered population list is chosen.

    • Examples:

      • Calling every 100th100^{\text{th}} person listed in a town phone book (residential white pages or business yellow pages).

      • Surveying every 20th20^{\text{th}} individual on an official library card registry.

      • Selecting every 4th4^{\text{th}} student who enters a classroom.

    • Convenience Sampling: Data is gathered from individuals or sources that are readily accessible and easily reached.

    • Characteristics: Highly prone to severe bias, though fast and inexpensive.

    • Example: An administrative assistant standing outside the campus library on a Wednesday morning surveying the first 100100 undergraduate students encountered regarding fall textbook costs.

TI-84 Calculator Procedure for Simple Random Sampling

  • Command Location:

    • Press the MATH key.

    • Navigate right using arrow keys to the PRB (Probability) menu header.

    • Select Option 5: randInt( (Random Integer generator).

  • Syntax Structure:   randInt(lower_bound,upper_bound,sample_size)\text{randInt}(\text{lower\_bound}, \text{upper\_bound}, \text{sample\_size})

  • Applied Class Demonstration:

    • To draw a random sample of 1010 items (n=10n = 10) from a population of 10001000 (N=1000N = 1000):

    • Enter: randInt(1, 1000, 10)

    • Hit ENTER. The calculator outputs an array of integers (e.g., 696,702,883...696, 702, 883...), identifying the specific numbered population subjects to be sampled.

Sources of Variation, Errors, and Sampling Biases

  • Sampling Variability: Natural, expected variation observed among sample responses drawn repeatedly from the same overall population.

    • Example (Quality Control at Soda Bottling Plant): Inspecting nominally filled 16oz16\,oz soda cans yields minor variations: 15.8oz15.8\,oz, 16.1oz16.1\,oz, 15.2oz15.2\,oz, 14.8oz14.8\,oz, 15.8oz15.8\,oz, 15.9oz15.9\,oz.

    • Causes: Operator measurement differences, mechanical filling variations, or equipment calibration issues.

    • Political Sampling Example: Repeated samples of 10001000 voters in Washington State will naturally fluctuate in Republican/Democrat proportions depending on whether a local sample disproportionately hits specific geographic sub-regions (e.g., retirement communities).

  • Sampling Errors vs. Non-Sampling Errors:

    • Sampling Error: Errors arising directly from the inherent limitations or flaws of the chosen sampling design (e.g., sample size being far too small or using convenience sampling).

    • Non-Sampling Error: Errors arising from external factors unrelated to the sampling method itself (e.g., defective measuring instruments, data entry errors, or uncalibrated scales).

  • Five Major Types of Sampling Bias:

    • Voluntary Response Bias: Bias resulting when participation is entirely voluntary. People with strong negative or positive opinions disproportionately respond.

    • Example: A magazine requesting readers to voluntarily fill out and mail back an advertising survey.

    • Self-Interest Study: Bias occurring when the entity conducting or funding the research stands to gain financially or operationally from specific results.

    • Example: A commercial potato company conducting a public poll regarding favorite vegetables and publishing results claiming potatoes are #1.

    • Response Bias: Bias occurring when respondents intentionally or unintentionally provide inaccurate, false, or misleading answers.

    • Loaded Question Bias: Bias resulting from phrasing survey questions in a leading manner that sways or forces the respondent toward a specific answer.

    • Example: "Do you prefer the delicious taste of Brand X or the taste of Brand Y?"

    • Non-Response Bias: Bias occurring when selected sample participants refuse to participate or fail to return survey responses.

  • Unbiased Study Example: A hospital study evaluating a new heart monitoring software collects data from all patients currently using the software. Because every subject in the target population is measured, no sampling bias exists.

Misleading Statistics and Data Interpretation Pitfalls

  • Sample Size Distortions:

    • Samples that are too small fail to adequately represent populations, resulting in unreliable statistics.

    • While larger samples generally reduce sampling error and bias, large samples can still be severely biased if selected improperly (e.g., large voluntary response samples).

  • Outlier Distortions:

    • Outlier Definition: A data value that lies an extreme distance away from the vast majority of other observations in a data set.

    • Impact on Averages: Outliers severely distort summary measures like the mean (average).

    • Example: A store advertises an average merchandise price of $7.00\$7.00. However, individual item shelf prices are $8.00\$8.00, $9.00\$9.00, $8.50\$8.50, and $10.00\$10.00. A single clearanced outlier priced at $1.00\$1.00 artificially pulls the mean down to $7.00\$7.00, giving consumers a completely misleading expectation of shelf prices.

  • Inaccurate Context and Comparisons:

    • Presenting raw counts without adjusting for population scale produces false claims.

    • Example: An article claims that California has more school teachers than Missouri, proving California cares more about teachers' rights and benefits. This claim is fundamentally misleading because California's total population is vastly larger than Missouri's. On a proportional per-capita basis, teacher densities may be identical.

Questions & Discussion

  • Question Regarding Vocabulary Assessment Format:

    • Student Question: Will the vocabulary quiz consist of matching definitions to terms, or will students be required to write out definitions by hand?

    • Instructor Response: The vocabulary assessment will primarily consist of multiple-choice questions, accompanied by brief short-answer questions. Homework target dates on Canvas are guidelines designed to keep students on schedule and avoid last-minute backlogs without penalty.