Comprehensive Study Notes on Introductory Statistics: Variables, Visualizations, Sampling Designs, and Errors

Fundamental Statistical Concepts and Variable Classifications

  • Definitions of Core Statistical Variables:

    • Data: The actual values or pieces of information collected from individuals or objects.

    • Variable: A characteristic or attribute under study that can assume different values for different individuals in a population or sample.

    • Categorical (Qualitative) Variables: Variables that express a quality, attribute, label, or category rather than a numerical value. Responses consist of words, names, or code labels.

    • Quantitative Variables: Variables that express a numerical quantity or amount. Responses consist of meaningful numbers that permit arithmetic operations.

  • Subtypes of Quantitative Variables:

    • Discrete Quantitative Variables:

    • Data resulting from a counting process.

    • Take on a finite or countably infinite set of distinct values.

    • Characterized by numerical gaps between permissible values.

    • While often represented by whole non-negative integers (0,1,2,3,…0, 1, 2, 3, \dots), discrete variables can also consist of finite decimal values if those values are isolated, fixed, and exact rather than falling anywhere along a continuous numerical interval.

    • Continuous Quantitative Variables:

    • Data resulting from a measurement process.

    • Take on any real numerical value within a given interval along a continuum.

    • Possess no gaps between values; between any two values, an infinite number of intermediate values can exist.

    • Precision is limited only by the accuracy and sensitivity of the measuring instrument used, generating values that can extend to arbitrary decimal places.

Deep-Dive Case Studies in Variable Classification

  • Case Study: Tuition Costs Analysis

    • Context: Determination of whether the total tuition cost for a student's classes at an institution (such as Austin Peay State University) is quantitative discrete or quantitative continuous.

    • Institutional Cost Structure:

    • For credit hours 11 through 1212: The fee is fixed at $331\$331 per credit hour.

    • For credit hours 1313 and above: The fee is $67\$67 per additional credit hour beyond 1212.

    • Calculated Exact Tuition Values:

    • 1 credit hour=$3311\,\text{credit hour} = \$331

    • 2\,\text{credit hours} = \331 \times 2 = \662662

    • 3\,\text{credit hours} = \331 \times 3 = \993993

    • 12\,\text{credit hours} = \331 \times 12 = \3,9723,972

    • 13 credit hours=$3,972+$67=$4,03913\,\text{credit hours} = \$3,972 + \$67 = \$4,039

    • Classification Determination:

    • Because tuition amounts are computed via multiplication and addition of fixed institutional rates, the potential outputs form a set of exact, isolated numerical values (\$331, \$662, \993, \dots, \3972,$40393972, \$4039).

    • Even if a rate includes fractional dollar amounts (e.g., $331.50\$331.50 per credit hour yielding $663.00\$663.00 for 22 hours), the resulting monetary outputs are exact and non-approximated.

    • Conclusion: In this specific transactional context, class tuition is classified as a discrete quantitative variable.

    • Macroeconomic vs. Contextual Rules for Currency:

    • Money in general macroeconomic models is often treated as continuous due to complex division, fractional interest accrual, and continuous mathematical modeling.

    • In real-world reporting, currency is bounded by cents (two decimal places, e.g., $0.01\$0.01), but continuous theoretical models accommodate infinite decimal division.

    • Distinction in phrasing: Questions asking "how many" strictly denote counting (discrete), whereas questions asking "how much" generally denote continuous measurement or evaluation of quantity. Context dictates whether currency functions as discrete or continuous.

  • Analysis of Common Operational Variables:

    • Calculator Brand/Type (e.g., TI-84, Casio): Categorical / Qualitative variable (label/name).

    • Political Party Preference: Categorical / Qualitative variable (political affiliation label).

    • Weight of Sumo Wrestlers: Continuous quantitative variable (measured mass falling along an unbroken continuum with no intermediate gaps).

    • Movie Ratings:

    • Maturity/Content Ratings (e.g., PG, PG-13, R): Qualitative variable (age appropriateness categories).

    • Ordinal Sentiment Ratings (e.g., Very Bad, Bad, Fair, Good, Excellent): Qualitative variable. Even if assigned arbitrary integer codes 11 through 55 (where 1=Very Bad1 = \text{Very Bad} and 5=Excellent5 = \text{Excellent}), the numbers represent underlying categorical ranks rather than measured quantities.

    • Numeric User Ratings (e.g., 1 to 10 scale): Discrete quantitative if restricted strictly to whole numbers (1,2,3,…,101, 2, 3, \dots, 10); continuous quantitative if reported as aggregate, unrounded decimal averages (e.g., 7.867427.86742).

    • Number of Correct Quiz Answers: Discrete quantitative variable (countable whole numbers).

    • Attitude Toward Government: Qualitative variable (categorical responses describing opinions or positions).

    • Intelligence Quotient (IQ) Scores: Continuous quantitative variable. Although IQ scores are conventionally reported as rounded whole integers (e.g., 171171), the score reflects an underlying measurement of cognitive ability derived from testing instruments. Intermediate continuous calculations (e.g., 170.5170.5) undergo rounding to produce final reported integer metrics.

Summarizing and Graphical Representations of Data

  • Mathematical Summarization vs. Underlying Data Type:

    • Collecting qualitative responses (e.g., student standing: Freshman, Sophomore, Junior, Senior) yields non-numerical categorical data.

    • Summarizing qualitative data into aggregated counts or frequencies (e.g., 6767 Freshmen, 1313 Sophomores out of 167167 total students) produces summary numbers, but does not convert the underlying variable into quantitative data.

    • Similarly, categorizing college enrollment as Full-Time versus Part-Time yields qualitative responses from individuals, even when presented in institutional summary tables containing counts and percentages (e.g., De Anza College reporting 40.9%40.9\% full-time and 59.1%59.1\% part-time enrollment).

  • Visualizations for Qualitative Data:

    • Pie Charts:

    • Graphically display the relative proportions or percentages of a complete whole (100%100\%).

    • Requirement 1: Every individual data point or observation must belong to exactly one category (mutually exclusive categories).

    • Requirement 2: The categories represented must sum to the complete total/whole (100%100\%).

    • Bar Graphs:

    • Display categories along one axis and counts/frequencies or percentages along the opposite axis using rectangular bars.

    • Flexibility: Unlike pie charts, bar graphs can be utilized when individual units fall into multiple categories simultaneously (e.g., students double-majoring in Mathematics and Physics).

    • Pareto Charts:

    • A specialized form of a bar graph designed for qualitative data.

    • Category bars are arranged strictly in descending order of frequency or percentage, from the largest category on the far left to the smallest category on the right.

Probability Sampling Methods and Design

  • Core Principles of Sampling:

    • Sampling Process: The operational method used to select a subset (sample) from an overall target population.

    • Representativeness: The primary objective of sampling is to obtain a sample that reflects the structural characteristics and properties of the parent population.

  • Primary Sampling Methodologies:

    • Simple Random Sampling (SRS):

    • Every member of the population, and every possible combination of sample size nn, has an equal probability of being selected.

    • Best executed using an indexed sampling frame and an algorithmic random number generator to eliminate conscious or unconscious human selection bias.

    • Stratified Random Sampling:

    • The population is divided into non-overlapping, mutually exclusive subgroups called strata (singular: stratum).

    • A random sample is drawn from every single stratum in proportion to each stratum's relative representation in the overall population.

    • Mathematical Proportions Example:

      • Population contains Strata 11 (1010 individuals), Strata 22 (1515 individuals), and Strata 33 (2020 individuals).

      • If a 20%20\% sampling fraction (210\frac{2}{10}) is selected from Strata 11 (22 individuals), proportionate selection requires selecting 20%20\% from Strata 22 (0.20×15=30.20 \times 15 = 3 individuals) and 20%20\% from Strata 33 (0.20×20=40.20 \times 20 = 4 individuals).

      • If a stratum contains 1212 individuals, exact proportion dictates 0.20×12=2.40.20 \times 12 = 2.4 individuals, requiring standard mathematical rounding to select 22 individuals.

    • Cluster Sampling:

    • The population is divided into non-overlapping subgroups called clusters.

    • A random selection of entire clusters is drawn.

    • All individuals contained within the chosen clusters are surveyed; no individuals from unselected clusters are included in the sample.

    • Systematic Sampling:

    • A random starting point is selected within the sampling frame.

    • Every kthk^{\text{th}} (or nthn^{\text{th}}) individual in the ordered population frame is selected sequentially thereafter (e.g., every 4th4^{\text{th}}, 5th5^{\text{th}}, or 10th10^{\text{th}} person).

    • Convenience Sampling:

    • A non-probability sampling design that selects readily accessible or convenient individuals from the population.

  • Operational Replacement Protocols:

    • Sampling with Replacement: An individual or unit selected from the population is returned to the pool prior to drawing the next unit. Consequently, a single individual can be selected more than once.

    • Sampling without Replacement: An individual or unit selected from the population is permanently removed from the pool for all subsequent draws. Consequently, an individual can be selected at most once.

Statistical Errors and Sampling Bias

  • Statistical Errors Defined:

    • Sampling Error:

    • The inevitable numerical discrepancy between a sample statistic and the true population parameter.

    • Caused strictly by natural sampling variability inherent in selecting a random subset rather than examining the entire population.

    • Non-Sampling Error:

    • Errors stemming from operational, technological, or human factors unrelated to the natural sampling process.

    • Examples include defective diagnostic equipment (e.g., faulty thermometers), data recording errors, uncalibrated instruments, or poorly constructed questions.

    • Sampling Bias:

    • Systematic favoring of specific groups or characteristics over others during sample selection.

    • Occurs when certain population members have zero or significantly reduced probabilities of selection, rendering the resulting sample unrepresentative and invalidating conclusions.

Applied Sampling Scenarios and Identification

  • Scenario A: A soccer coach categorizes players into age groups (88 to 1010, 1111 to 1212, and 1313 to 1414) and selects 66 players from the 88 to 1010 group, 77 players from the 1111 to 1212 group, and 33 players from the 1313 to 1414 group.

    • Classification: Stratified Sampling (Population partitioned into age strata; samples drawn from every stratum).

  • Scenario B: A researcher randomly selects 55 high-tech companies and surveys all human resource personnel working within those 55 companies.

    • Classification: Cluster Sampling (Companies act as clusters; clusters randomly drawn; all members within selected clusters surveyed).

  • Scenario C: A high school researcher surveys 5050 female high school teachers and 5050 male high school teachers.

    • Classification: Stratified Sampling (Population partitioned into gender strata; samples drawn from both strata).

  • Scenario D: A medical study selects and tests every 4th4^{\text{th}} cancer patient from an official patient registry.

    • Classification: Systematic Sampling (Selection follows a fixed periodic interval k=4k = 4).

  • Scenario E: A researcher assigns unique identification numbers to a population frame and uses a random number generator to select participants where every individual has an equal selection probability.

    • Classification: Simple Random Sampling.

  • Scenario F: A student researcher conducts a survey by interviewing classmates sitting in the immediate classroom environment.

    • Classification: Convenience Sampling.