Sampling Methods and Sample Size Determination

Fundamental Concepts in Sampling

  • Sampling Definition: The process of selecting a subset (sample) of a population to make inferences, estimate population parameters, and conduct statistical hypothesis tests.

  • Reasons for Sampling: Reduced cost, greater speed/timeliness, greater efficiency and accuracy, greater scope, convenience, necessity (for destructive sampling), and ethical considerations (such as testing new drugs).

  • Core Population & Frame Concepts:

    • Target Population: The population about which information is desired.

    • Sampled / Sampling Population: The population from which a sample is actually taken as determined by the sampling frame.

    • Sampling Frame: A complete list or mechanism providing observational access to sampling units in the population.

    • Under-coverage Error: Occurs when the sampling frame excludes units that are part of the target population.

    • Over-coverage Error: Occurs when the sampling frame includes units that are not part of the target population.

Classification of Sampling Techniques

  • Sample Selection Criteria: Samples must be representative, provide precise and measurably reliable estimates, and minimize selection costs.

  • Non-Probability Sampling:

    • Units are selected non-randomly; inclusion probabilities are unknown or zero for some units.

    • Restricted strictly to descriptive statements; cannot be used to make generalizations about the population.

    • Convenience / Accidental Sampling: Uses whichever units are readily available (e.g., selecting the first 100100 customers entering a store).

    • Judgment / Purposive Sampling: Selected according to the sampler's subjective judgment or intuition.

    • Quota Sampling: Arbitrarily selecting a specified number of units possessing given characteristics.

  • Probability Sampling Methods:

    • Each population unit has a known, non-zero probability of inclusion, enabling valid statistical inference.

    • Simple Random Sampling (SRS): Every unit has an equal chance of inclusion. Conducted with or without replacement using random number tables, calculator RAN functions, chips-in-a-box, or statistical software.

    • Stratified Random Sampling: Population is divided into LL non-overlapping, homogeneous sub-populations (strata), and simple random samples are drawn independently from each stratum.

    • Equal Allocation: n1=n2=⋯=nh=nLn_1 = n_2 = \dots = n_h = \frac{n}{L}

    • Proportional Allocation: nh=n(NhN)n_h = n \left( \frac{N_h}{N} \right)

    • Systematic Random Sampling: Selecting every kkth unit from an ordered frame starting from a random start rr (1≤r≤k1 \le r \le k), where kk is the whole number ratio Nn\frac{N}{n} and 1k\frac{1}{k} is the sampling fraction.

    • Cluster Sampling: Population is partitioned into non-overlapping clusters containing heterogeneous elements. Sampling units are entire clusters, and all elements within selected clusters are enumerated. Sample size is not fixed.

Sample Size Determination

  • Key Determinants:

    • Population Size (NN): Total number of individuals in the demographic.

    • Margin of Error (ee): Acceptable variation between the sample statistic and population parameter.

    • Confidence Level & Z-Score:

    • 90%90\% Confidence Level: Z=1.645Z = 1.645

    • 95%95\% Confidence Level: Z=1.96Z = 1.96

    • 99%99\% Confidence Level: Z=2.326Z = 2.326

    • Standard Deviation / Variance (σ\sigma or pp): Assumed as p=0.5p = 0.5 for maximum variance when prior data is unavailable.

  • Calculating Sample Size for Infinite or Very Large Populations:

    • Formula for estimating mean or proportion:     n0=Z2σ2e2andn0=Z2p(1−p)e2n_0 = \frac{Z^2 \sigma^2}{e^2} \quad \text{and} \quad n_0 = \frac{Z^2 p(1 - p)}{e^2}

    • Example (95%95\% confidence level, p=0.5p = 0.5, e=0.05e = 0.05):     n0=(1.96)2(0.5)(1−0.5)(0.05)2=(3.8416)(0.25)0.0025=384.16  ⟹  385 respondentsn_0 = \frac{(1.96)^2 (0.5)(1 - 0.5)}{(0.05)^2} = \frac{(3.8416)(0.25)}{0.0025} = 384.16 \implies 385 \text{ respondents}

  • Finite Population Correction (FPC):

    • Sampling error formulas incorporating FPC:     e=ZσnN−nN−1e = Z \frac{\sigma}{\sqrt{n}} \sqrt{\frac{N - n}{N - 1}}     e=Zp(1−p)nN−nN−1e = Z \sqrt{\frac{p(1 - p)}{n}} \sqrt{\frac{N - n}{N - 1}}

    • Adjusted sample size formula:     n=n0Nn0+(N−1)n = \frac{n_0 N}{n_0 + (N - 1)}

    • Saxon Home Improvement Company Examples (N=5000N = 5000):

    • Estimating Mean (S=$25S = \$25, e=$5e = \$5, Z=1.96Z = 1.96):       n=(96.04)(5000)96.04+(5000−1)=94.24  ⟹  95n = \frac{(96.04)(5000)}{96.04 + (5000 - 1)} = 94.24 \implies 95

    • Estimating Proportion (p=0.15p = 0.15, e=0.07e = 0.07, Z=1.96Z = 1.96):       n=(99.96)(5000)99.96+(5000−1)=98.02  ⟹  99n = \frac{(99.96)(5000)}{99.96 + (5000 - 1)} = 98.02 \implies 99

    • To satisfy both estimations simultaneously in a single sample, the larger sample size of 9999 is chosen.

Sample size for estimating the mean with finite population correction