Comprehensive Study Notes on Statistical Sampling Methods and Simple Random Sampling Techniques
Fundamentals of Sampling and Simple Random Sampling (SRS)
Definition of Simple Random Sampling (SRS): Simple random sampling is a fair probability sampling method where every possible group of size from a population of size has the exact same chance of being picked, and every individual has an equal chance of being selected.
Key Notation:
represents the total population size (uppercase letter).
represents the sample size (lowercase letter).
Strict Rule: The sample size must always be strictly less than the population size ().
Advanced Probability and Combinatorics in SRS
Combinatorial Enumeration: In SRS without replacement, the order of selection does not matter. To calculate the total number of distinct possible samples, use the combination formula:
Case Study (Sofia's Concert Tickets):
Scenario: Sofia has concert tickets. She keeps for herself and randomly picks companions from interested friends (Yolanda, Michael, Kevin, Marissa, Annie, and Katie).
Sample Space: There are exactly unique, equally likely combinations of size :
Probability Calculation: Because each trio is equally likely, the probability of selecting any specific trio (e.g., Michael, Kevin, and Marissa) is:
Methods for Obtaining Simple Random Samples
Population Frame: A frame is a complete list of all individuals in the population. Every individual must be assigned a unique numerical label from to .
Digit Consistency Rule: When using random digit tables, all labels must have the same number of digits as .
Example: If , labels are to .
Example: If , labels are to .
Sampling Protocols:
Sampling Without Replacement: Once selected, an individual is removed and cannot be picked again. The probability of selection for remaining individuals increases with each draw. This is the standard procedure in statistics.
Sampling With Replacement: Selected individuals are placed back into the pool and can be chosen multiple times. Probability remains constant at .
Implementation Tools:
Table of Random Digits:
Pick an initial starting spot.
Read numbers in blocks matching the label digit length.
Ignore numbers exceeding or equal to , and skip duplicates when sampling without replacement.
Continue until distinct items are selected.
Sennesee and Associates Case Study:
Parameters: Population clients labeled to ; sample size n = 5$.
Trace (Starting Row 13, Col 4-5):
Row 13, Col 4-5: 01\rightarrow52\rightarrow52 > 3007\rightarrow44\rightarrow44 > 3046\rightarrow46 > 3037\rightarrow37 > 3023\rightarrow Selected (Client 23: Travel Zone)
Graphing Calculators & Software:
Menu:
MATH\rightarrownn5, 11, 4, 20, 29nN = 21\,909$.Three Non-Overlapping Strata:
Resident students:
Nonresident / commuting students:
Staff and faculty:
Proportional Sample Weighting ():
Resident students: students
Nonresident students: students
Staff / faculty: staff members
Execution: Draw an SRS of staff, resident students, and nonresident students.
Systematic Sampling
Definition: A systematic sample is obtained by selecting every individual from the population. The starting individual is randomly chosen between and k$.
Procedure with Known Population Frame (NNn$.
Calculate interval size . ALWAYS round DOWN to the nearest whole integer.
Randomly choose starting integer between and inclusive.
Sequence of selected labels: p, p+k, p+2k, p+3k, \dots, p+(n-1)k$.
Procedure Without Frame (Unknown Population Size Nk7^{\text{th}}p1k$.
Sample every individual until sample size is reached.
Detailed Case Studies:
Known Frame Example (, ):
Interval , start p = 23$.
Sequence (1023, 27, 31, 35, 39, 43, 47, 51, 55, 59$.
Kroger Customer Satisfaction Example (Unknown , ):
Interval (every customer), random start p = 5$.
Selected sequence: Customer 5, 12, 19, 26, \dotsn^{\text{th}}p + (n-1)k5 + (40-1) \times 7 = 5 + 39 \times 7 = 5 + 273 = 278278^{\text{th}}N = 4\,408n = 40k = \frac{4\,408}{40} = 110.2 \rightarrow 110p = 10$.
First three individuals: , , 10+2(110)=230$.
Final (40^{\text{th}}10 + (40-1) \times 110 = 10 + 39 \times 110 = 10 + 4\,290 = 4\,300^{\text{th}}10\,493$.
Procedure: Number blocks to . Draw an SRS of block numbers. Survey every household on those selected blocks.
Comparison of Sampling Techniques
Stratified vs. Cluster Sampling Distinction:
Stratified Sampling: Divide population into strata Take SRS from EACH stratum Survey selected individuals. (NOT everyone in a stratum is surveyed).
Cluster Sampling: Divide population into clusters Take SRS of the CLUSTERS Survey ALL individuals in chosen clusters. (EVERYONE in chosen clusters is surveyed).
Comparative Overview:
Simple Random Sampling: Pure chance selection from full list. Lowest bias, but high cost for widespread groups.
Stratified Sampling: Guarantees representation across key subgroups; best when opinions differ by subgroup.
Systematic Sampling: Interval selection (). Best for continuous lines or assembly processes where a full frame is unknown.
Cluster Sampling: Whole group selection. Highly cost-efficient for geographically scattered populations.
Non-Probability Sampling: Convenience and Voluntary Response
Convenience Sampling:
Mechanism: Individuals are selected based on easy accessibility without chance or randomization.
Example: Surveying the first students leaving a library to estimate total student study hours.
Flaws: Excludes online students and non-library visitors. Results are unrepresentative, heavily biased, and scientifically invalid.
Voluntary Response Sampling:
Mechanism: Participants self-select whether to respond (e.g., call-in polls, text surveys, online forms).
Flaws: Strongly over-represents individuals with extreme or negative opinions. Highly biased and unreliable.
Multistage Sampling
Definition: Multistage sampling combines two or more sampling techniques in successive steps. Used by professional research organizations for national data collection.
Nielsen Media Research Case Study:
Objective: Track national TV viewing habits using electronic "people meter" devices.
Two-Stage Sampling Process:
Stage 1 (Stratified Sampling): Uses Census data to divide the U.S. into roughly geographic strata (city blocks in urban areas, regional zones in rural areas).
Stage 2 (Simple Random Sampling): Lists households within selected strata and picks final households via simple random sampling.