Comprehensive Study Notes on Statistical Sampling Methods and Simple Random Sampling Techniques

Fundamentals of Sampling and Simple Random Sampling (SRS)

  • Definition of Simple Random Sampling (SRS): Simple random sampling is a fair probability sampling method where every possible group of size nn from a population of size NN has the exact same chance of being picked, and every individual has an equal chance of being selected.

  • Key Notation:

    • NN represents the total population size (uppercase letter).

    • nn represents the sample size (lowercase letter).

    • Strict Rule: The sample size nn must always be strictly less than the population size NN (n<Nn < N).

Advanced Probability and Combinatorics in SRS

  • Combinatorial Enumeration: In SRS without replacement, the order of selection does not matter. To calculate the total number of distinct possible samples, use the combination formula:
    (Nn)\binom{N}{n}

  • Case Study (Sofia's Concert Tickets):

    • Scenario: Sofia has 44 concert tickets. She keeps 11 for herself and randomly picks n=3n = 3 companions from N=6N = 6 interested friends (Yolanda, Michael, Kevin, Marissa, Annie, and Katie).

    • Sample Space: There are exactly (63)=20\binom{6}{3} = 20 unique, equally likely combinations of size n=3n = 3:

    1. Yolanda,Michael,Kevin{\text{Yolanda}, \text{Michael}, \text{Kevin}}

    2. Yolanda,Michael,Marissa{\text{Yolanda}, \text{Michael}, \text{Marissa}}

    3. Yolanda,Michael,Annie{\text{Yolanda}, \text{Michael}, \text{Annie}}

    4. Yolanda,Michael,Katie{\text{Yolanda}, \text{Michael}, \text{Katie}}

    5. Yolanda,Kevin,Marissa{\text{Yolanda}, \text{Kevin}, \text{Marissa}}

    6. Yolanda,Kevin,Annie{\text{Yolanda}, \text{Kevin}, \text{Annie}}

    7. Yolanda,Kevin,Katie{\text{Yolanda}, \text{Kevin}, \text{Katie}}

    8. Yolanda,Marissa,Annie{\text{Yolanda}, \text{Marissa}, \text{Annie}}

    9. Yolanda,Marissa,Katie{\text{Yolanda}, \text{Marissa}, \text{Katie}}

    10. Yolanda,Annie,Katie{\text{Yolanda}, \text{Annie}, \text{Katie}}

    11. Michael,Kevin,Marissa{\text{Michael}, \text{Kevin}, \text{Marissa}}

    12. Michael,Kevin,Annie{\text{Michael}, \text{Kevin}, \text{Annie}}

    13. Michael,Kevin,Katie{\text{Michael}, \text{Kevin}, \text{Katie}}

    14. Michael,Marissa,Annie{\text{Michael}, \text{Marissa}, \text{Annie}}

    15. Michael,Marissa,Katie{\text{Michael}, \text{Marissa}, \text{Katie}}

    16. Michael,Annie,Katie{\text{Michael}, \text{Annie}, \text{Katie}}

    17. Kevin,Marissa,Annie{\text{Kevin}, \text{Marissa}, \text{Annie}}

    18. Kevin,Marissa,Katie{\text{Kevin}, \text{Marissa}, \text{Katie}}

    19. Kevin,Annie,Katie{\text{Kevin}, \text{Annie}, \text{Katie}}

    20. Marissa,Annie,Katie{\text{Marissa}, \text{Annie}, \text{Katie}}

    • Probability Calculation: Because each trio is equally likely, the probability of selecting any specific trio (e.g., Michael, Kevin, and Marissa) is:
      P(Specific Trio)=120=0.05=5%P(\text{Specific Trio}) = \frac{1}{20} = 0.05 = 5\%

Methods for Obtaining Simple Random Samples

  • Population Frame: A frame is a complete list of all individuals in the population. Every individual must be assigned a unique numerical label from 11 to NN.

  • Digit Consistency Rule: When using random digit tables, all labels must have the same number of digits as NN.

    • Example: If N=30N = 30, labels are 0101 to 3030.

    • Example: If N=500N = 500, labels are 001001 to 500500.

  • Sampling Protocols:

    • Sampling Without Replacement: Once selected, an individual is removed and cannot be picked again. The probability of selection for remaining individuals increases with each draw. This is the standard procedure in statistics.

    • Sampling With Replacement: Selected individuals are placed back into the pool and can be chosen multiple times. Probability remains constant at 1N\frac{1}{N}.

  • Implementation Tools:

    • Table of Random Digits:

    1. Pick an initial starting spot.

    2. Read numbers in blocks matching the label digit length.

    3. Ignore numbers exceeding NN or equal to 0000, and skip duplicates when sampling without replacement.

    4. Continue until nn distinct items are selected.

    • Sennesee and Associates Case Study:

    • Parameters: Population N=30N = 30 clients labeled 0101 to 3030; sample size n = 5$.

    • Trace (Starting Row 13, Col 4-5):

      • Row 13, Col 4-5: 01\rightarrow<strong>Selected</strong>(Client01:ABCElectric)</p></li><li><p>Row14,Col4−5:<strong>Selected</strong> (Client 01: ABC Electric)</p></li><li><p>Row 14, Col 4-5:52\rightarrow<strong>Ignore</strong>(<strong>Ignore</strong> (52 > 30)</p></li><li><p>Row15,Col4−5:)</p></li><li><p>Row 15, Col 4-5:07\rightarrow<strong>Selected</strong>(Client07:DynoJump)</p></li><li><p>Row16,Col4−5:<strong>Selected</strong> (Client 07: Dyno Jump)</p></li><li><p>Row 16, Col 4-5:44\rightarrow<strong>Ignore</strong>(<strong>Ignore</strong> (44 > 30)</p></li><li><p>Row17,Col4−5:)</p></li><li><p>Row 17, Col 4-5:46\rightarrow<strong>Ignore</strong>(<strong>Ignore</strong> (46 > 30)</p></li><li><p>Row18,Col4−5:)</p></li><li><p>Row 18, Col 4-5:37\rightarrow<strong>Ignore</strong>(<strong>Ignore</strong> (37 > 30)</p></li><li><p>TransitiontoCol6:)</p></li><li><p>Transition to Col 6:23\rightarrow Selected (Client 23: Travel Zone)

    • Graphing Calculators & Software:

    • Menu: MATH \rightarrow<code>PROB</code></p></li><li><p><code>randInt(lower,upper,n)</code>:Generates<code>PROB</code></p></li><li><p><code>randInt(lower, upper, n)</code>: Generatesnrandomintegers(allowsduplicates).</p></li><li><p><code>randIntNoRep(lower,upper,n)</code>:Generatesrandom integers (allows duplicates).</p></li><li><p><code>randIntNoRep(lower, upper, n)</code>: Generatesn<strong>unique</strong>randomintegers(<strong>noduplicates</strong>).</p></li><li><p>SyntaxExample:<code>randIntNoRep(1,30,5)</code>yieldsfiveuniqueselectionslike<strong>unique</strong> random integers (<strong>no duplicates</strong>).</p></li><li><p>Syntax Example: <code>randIntNoRep(1, 30, 5)</code> yields five unique selections like5, 11, 4, 20, 29.</p></li><li><p><strong>PhysicalRandomization:</strong>Writinglabelsonidenticalslipsofpaper,mixingthoroughlyinacontainer,anddrawing.</p></li><li><p><strong>Physical Randomization:</strong> Writing labels on identical slips of paper, mixing thoroughly in a container, and drawingnslips.</p></li></ul></li></ul><h4>StratifiedSampling</h4><ul><li><p><strong>Definition:</strong>A<strong>stratifiedsample</strong>isobtainedbydividingthepopulationintoseparate,non−overlappingsubgroupscalled<strong>strata</strong>,andthentakinga<strong>simplerandomsamplefromEACHstratum</strong>.</p></li><li><p><strong>StratumHomogeneity:</strong>Individualswithineachstratummustbe<strong>homogeneous(similar)</strong>regardingaspecificcharacteristic(e.g.,biologicalsex,agebracket,location).</p></li><li><p><strong>DePaulUniversityCaseStudy:</strong></p><ul><li><p><strong>Objective:</strong>Surveycampussafetyopinionsacrosspopulationslips.</p></li></ul></li></ul><h4>Stratified Sampling</h4><ul><li><p><strong>Definition:</strong> A <strong>stratified sample</strong> is obtained by dividing the population into separate, non-overlapping subgroups called <strong>strata</strong>, and then taking a <strong>simple random sample from EACH stratum</strong>.</p></li><li><p><strong>Stratum Homogeneity:</strong> Individuals within each stratum must be <strong>homogeneous (similar)</strong> regarding a specific characteristic (e.g., biological sex, age bracket, location).</p></li><li><p><strong>DePaul University Case Study:</strong></p><ul><li><p><strong>Objective:</strong> Survey campus safety opinions across populationN = 21\,909$.

    • Three Non-Overlapping Strata:

    1. Resident students: N1=6 204N_1 = 6\,204

    2. Nonresident / commuting students: N2=13 304N_2 = 13\,304

    3. Staff and faculty: N3=2 401N_3 = 2\,401

    • Proportional Sample Weighting (n=100n = 100):

    • Resident students: 6 20421 909≈28.31%→28\frac{6\,204}{21\,909} \approx 28.31\% \rightarrow 28 students

    • Nonresident students: 13 30421 909≈60.72%→61\frac{13\,304}{21\,909} \approx 60.72\% \rightarrow 61 students

    • Staff / faculty: 2 40121 909≈10.96%→11\frac{2\,401}{21\,909} \approx 10.96\% \rightarrow 11 staff members

    • Execution: Draw an SRS of 1111 staff, 2828 resident students, and 6161 nonresident students.

Systematic Sampling

  • Definition: A systematic sample is obtained by selecting every kthk^{\text{th}} individual from the population. The starting individual pp is randomly chosen between 11 and k$.

  • Procedure with Known Population Frame (N):</strong></p><ol><li><p>Determinepopulationsize):</strong></p><ol><li><p>Determine population sizeNandsamplesizeand sample sizen$.

  • Calculate interval size k=⌊Nn⌋k = \left\lfloor \frac{N}{n} \right\rfloor. ALWAYS round DOWN to the nearest whole integer.

  • Randomly choose starting integer pp between 11 and kk inclusive.

  • Sequence of selected labels: p, p+k, p+2k, p+3k, \dots, p+(n-1)k$.

  • Procedure Without Frame (Unknown Population Size N):</strong></p><ol><li><p>Pickaninterval):</strong></p><ol><li><p>Pick an intervalk(e.g.,every(e.g., every7^{\text{th}}customer).</p></li><li><p>Choosearandomstartcustomer).</p></li><li><p>Choose a random startpbetweenbetween1andandk$.

  • Sample every kthk^{\text{th}} individual until sample size nn is reached.

  • Detailed Case Studies:

    • Known Frame Example (N=100N = 100, n=10n = 10):

    • Interval k=4k = 4, start p = 23$.

    • Sequence (10individuals):individuals):23, 27, 31, 35, 39, 43, 47, 51, 55, 59$.

    • Kroger Customer Satisfaction Example (Unknown NN, n=40n = 40):

    • Interval k=7k = 7 (every 7th7^{\text{th}} customer), random start p = 5$.

    • Selected sequence: Customer 5, 12, 19, 26, \dots</p></li><li><p>Formulaforfinal(</p></li><li><p>Formula for final (n^{\text{th}})customer:<br>) customer:<br>p + (n-1)k</p></li><li><p>Calculation:<br></p></li><li><p>Calculation:<br>5 + (40-1) \times 7 = 5 + 39 \times 7 = 5 + 273 = 278</p></li><li><p>Result:The<strong></p></li><li><p>Result: The <strong>278^{\text{th}}customer</strong>exitingthestoreisthefinalpersonsurveyed.</p></li><li><p><strong>HRBenefitsSurveyExample(customer</strong> exiting the store is the final person surveyed.</p></li><li><p><strong>HR Benefits Survey Example (N = 4\,408,,n = 40):</strong></p></li><li><p>Interval:):</strong></p></li><li><p>Interval:k = \frac{4\,408}{40} = 110.2 \rightarrow 110(<strong>roundeddown</strong>).</p></li><li><p>Randomstart:(<strong>rounded down</strong>).</p></li><li><p>Random start:p = 10$.

    • First three individuals: 1010, 10+110=12010+110=120, 10+2(110)=230$.

    • Final (40^{\text{th}})individual:<br>) individual:<br>10 + (40-1) \times 110 = 10 + 39 \times 110 = 10 + 4\,290 = 4\,300^{\text{th}}employee.</p></li></ul></li></ul><h4>ClusterSampling</h4><ul><li><p><strong>Definition:</strong>A<strong>clustersample</strong>isobtainedbyselecting<strong>ALLindividuals</strong>withinarandomlyselectedgrouporcollectionofgroups.</p></li><li><p><strong>Mechanism:</strong></p><ol><li><p>Dividethepopulationintogroups(<strong>clusters</strong>).</p></li><li><p>Takeasimplerandomsampleofthe<strong>clusters</strong>.</p></li><li><p>Survey<strong>EVERYindividual</strong>insidetheselectedclusters.</p></li></ol></li><li><p><strong>PrimaryAdvantages:</strong>Savessignificanttimeandmoney;<strong>doesnotrequirealistofeveryindividual</strong>(onlyaframeofclusters).</p></li><li><p><strong>BostonHouseholdIncomeCaseStudy:</strong></p><ul><li><p>Target:HouseholdincomeacrossBostoncityblocks.</p></li><li><p>Totalcityblocks(clusters):employee.</p></li></ul></li></ul><h4>Cluster Sampling</h4><ul><li><p><strong>Definition:</strong> A <strong>cluster sample</strong> is obtained by selecting <strong>ALL individuals</strong> within a randomly selected group or collection of groups.</p></li><li><p><strong>Mechanism:</strong></p><ol><li><p>Divide the population into groups (<strong>clusters</strong>).</p></li><li><p>Take a simple random sample of the <strong>clusters</strong>.</p></li><li><p>Survey <strong>EVERY individual</strong> inside the selected clusters.</p></li></ol></li><li><p><strong>Primary Advantages:</strong> Saves significant time and money; <strong>does not require a list of every individual</strong> (only a frame of clusters).</p></li><li><p><strong>Boston Household Income Case Study:</strong></p><ul><li><p>Target: Household income across Boston city blocks.</p></li><li><p>Total city blocks (clusters):10\,493$.

    • Procedure: Number blocks 11 to 10 49310\,493. Draw an SRS of 2020 block numbers. Survey every household on those 2020 selected blocks.

Comparison of Sampling Techniques

  • Stratified vs. Cluster Sampling Distinction:

    • Stratified Sampling: Divide population into strata →\rightarrow Take SRS from EACH stratum →\rightarrow Survey selected individuals. (NOT everyone in a stratum is surveyed).

    • Cluster Sampling: Divide population into clusters →\rightarrow Take SRS of the CLUSTERS →\rightarrow Survey ALL individuals in chosen clusters. (EVERYONE in chosen clusters is surveyed).

  • Comparative Overview:

    • Simple Random Sampling: Pure chance selection from full list. Lowest bias, but high cost for widespread groups.

    • Stratified Sampling: Guarantees representation across key subgroups; best when opinions differ by subgroup.

    • Systematic Sampling: Interval selection (kthk^{\text{th}}). Best for continuous lines or assembly processes where a full frame is unknown.

    • Cluster Sampling: Whole group selection. Highly cost-efficient for geographically scattered populations.

Non-Probability Sampling: Convenience and Voluntary Response

  • Convenience Sampling:

    • Mechanism: Individuals are selected based on easy accessibility without chance or randomization.

    • Example: Surveying the first 5050 students leaving a library to estimate total student study hours.

    • Flaws: Excludes online students and non-library visitors. Results are unrepresentative, heavily biased, and scientifically invalid.

  • Voluntary Response Sampling:

    • Mechanism: Participants self-select whether to respond (e.g., call-in polls, text surveys, online forms).

    • Flaws: Strongly over-represents individuals with extreme or negative opinions. Highly biased and unreliable.

Multistage Sampling

  • Definition: Multistage sampling combines two or more sampling techniques in successive steps. Used by professional research organizations for national data collection.

  • Nielsen Media Research Case Study:

    • Objective: Track national TV viewing habits using electronic "people meter" devices.

    • Two-Stage Sampling Process:

    • Stage 1 (Stratified Sampling): Uses Census data to divide the U.S. into roughly 6 0006\,000 geographic strata (city blocks in urban areas, regional zones in rural areas).

    • Stage 2 (Simple Random Sampling): Lists households within selected strata and picks final households via simple random sampling.