Chapter 4 Notes: Sample Surveys in the Real World

Overview of Sample Surveys and Response Rates

  • Survey Participation and Response Dynamics

    • Reaching human subjects and obtaining completed survey responses presents substantial practical challenges.

    • Incentives, monetary compensation, or brief survey durations are frequently employed to encourage participation.

    • Below is an example of a corporate survey invitation sent via email to collect customer feedback:

Netflix survey invitation email example
  • Mathematical Formulation of Survey Metrics

    • Response Rate Formula:     Response Rate=Total Number of ResponsesTotal Number Contacted×100%\text{Response Rate} = \frac{\text{Total Number of Responses}}{\text{Total Number Contacted}} \times 100\%

    • Nonresponse Rate Formula:     Nonresponse Rate=100%Response Rate\text{Nonresponse Rate} = 100\% - \text{Response Rate}

  • Empirical Telephone Survey Outcome Breakdown

    • The table below illustrates contact outcomes from a real-world dual-frame survey (landline vs. cell phone):

Survey outcome table for landline and cell phone samples
  • Landline Sample Call Metrics:

    • Noncontacts: 24642464

    • Not eligible: 18,42718,427

    • Other: 104104

    • Unknown eligibility: 33053305

    • Refusals, breakoffs, and partials: 47194719

    • Complete interviews: 902902

    • Total landline numbers called: 29,92129,921

  • Cell Phone Sample Call Metrics:

    • Noncontacts: 31143114

    • Not eligible: 60846084

    • Other: 5656

    • Unknown eligibility: 361361

    • Refusals, breakoffs, and partials: 48364836

    • Complete interviews: 605605

    • Total cell phone numbers called: 15,05615,056

  • Combined Response Calculation:

    • Total completes across both frames: 902+605=1507902 + 605 = 1507

    • Total phone numbers called across both frames: 29,921+15,056=44,97729,921 + 15,056 = 44,977

    • Overall Response Rate:       Response Rate=902+60529,921+15,056=150744,9773.4%\text{Response Rate} = \frac{902 + 605}{29,921 + 15,056} = \frac{1507}{44,977} \approx 3.4\%

    • Overall Nonresponse Rate:       Nonresponse Rate=100%3.4%=96.6%\text{Nonresponse Rate} = 100\% - 3.4\% = 96.6\%

  • Extreme nonresponse rates (such as 96.6\%$) call into question whether the participating sample remains representative of the overall target population.\n\n# Classification of Survey Errors\n\n* **Sampling Errors**\n * Errors that occur directly as a result of the act of sampling.\n * Includes deliberate or structural sampling bias (e.g., convenience or voluntary response sampling) as well as natural random sampling error.\n\n* **Non-Sampling Errors**\n * Errors that arise from factors unrelated to the act of sampling.\n * Caused by structural flaws in the experimental design, execution, data collection, or participant interactions.\n\n# Sampling Errors and Mitigation Strategies\n\n* **Types of Sampling Errors**\n * **Biased Sampling Design:** Non-probability designs (e.g., convenience sampling, voluntary response sampling) systematically favor certain outcomes over others.\n * **Random Sampling Error:** Natural variation resulting from surveying a sample subset rather than observing the entire population.\n\n* **Characteristics of Random Sampling Error**\n * It is inherent and unavoidable in any sample-based study.\n * Margin of Error (MOE) in formal confidence statements exclusively accounts for random sampling error.\n * MOE does **not** account for nonresponse, bad wording, interviewer bias, or sampling bias.\n * **Reduction:** Increasing the sample size n reduces random sampling error.\n * **Elimination:** Conducting a full population census completely eliminates random sampling error.\n\n# Non-Sampling Errors\n\n* **Primary Sources of Non-Sampling Errors**\n * Nonresponse bias.\n * Misleading or poorly worded survey items.\n * Interviewer bias (influencing responses intentionally or unintentionally).\n * Data entry errors by researchers.\n * False responses supplied by subjects (either accidentally or intentionally).\n\n* **Strategies for Mitigating Nonresponse**\n * **Alternative Contact Methods:** Follow up with nonresponders using different modes of communication (e.g., switching from phone calls to email outreach).\n * **Household Substitution:** Replace nonresponding households with demographically similar households using historical demographic data (Note: this practice can introduce secondary errors).\n * **Probability Weighting:** Adjust final estimates by mathematically weighting responses based on known population distributions.\n\n# Probability Weighting and Population Adjustment\n\n* **Mechanics of Probability Weighting**\n * If a demographic subgroup's response rate is below its population proportion, responses from that subgroup are assigned a higher statistical weight.\n * If a demographic subgroup's response rate is above its population proportion (e.g., overrepresentation of rural areas), responses from that subgroup are assigned a lower statistical weight.\n * Real-World Weighting Example (New York Times Polls): Survey results are weighted to adjust for household size, total residential phone lines, geographic region, sex, race, age, and educational attainment.\n\n* **Worked Mathematical Example of Probability Weighting**\n\n![Probability weighting calculation table by gender](https://assets.knowt.com/pdf-flow-prod/c8ada8ff-6f2f-4e73-9965-ff7d823e800a-figures/6.png)\n\n * **Population Parameters:** Target population is University of South Carolina (USC) students, known to be 60\%female(female (0.60)and) and40\%male(male (0.40).\n * **Sample Collected:** n = 100totalstudents,consistingoftotal students, consisting of50malesandmales and50 females.\n * **Subgroup Survey Results:**\n * Males: n_{\text{male}} = 50,numbersupportingbill=, number supporting bill =20.\n      \hat{p}{\text{male}} = \frac{20}{50} = 0.40\n * Females: n{\text{female}} = 50,numbersupportingbill=, number supporting bill =35.\n      \hat{p}{\text{female}} = \frac{35}{50} = 0.70\n * **Unadjusted Sample Proportion:**\n    \hat{p}{\text{unadjusted}} = \frac{20 + 35}{50 + 50} = \frac{55}{100} = 0.55\n * **Weighted (Adjusted) Population Estimate:**\n    \hat{p}_{\text{weighted}} = (0.40)(0.40) + (0.70)(0.60) = 0.16 + 0.42 = 0.58\n * **Conclusion:** Weighting gives females a higher weight (0.60insteadoftheirinstead of their0.50 sample proportion) to accurately reflect the overall population composition.\n\n# Misleading Questions and Question Phrasing Effects\n\n* **Psychology of Wording**\n * Minor changes in sentence structure, term choice, or framing heavily alter participant responses.\n\n* **Government Spending Survey Comparison**\n * **Framing A:** "Is our government spending too much money for welfare programs?"\n * Result: 44\% of respondents answered "yes".\n * **Framing B:** "Is our government spending too much money for assistance to the poor?"\n * Result: 13\% of respondents answered "yes".\n\n* **Gallup Tax Survey Study (April 2018)**\n\n![Gallup Poll question wording on tax assessment](https://assets.knowt.com/pdf-flow-prod/c8ada8ff-6f2f-4e73-9965-ff7d823e800a-figures/2.png)\n\n * **Question 1:** "Do you consider the amount of federal income tax you have to pay as too high, about right, or too low?"\n * Results: 48\%stated"aboutright",stated "about right",45\% stated "too high".\n * **Question 2:** "Do you regard the income tax which you will have to pay this year as fair?"\n * Results: 61\% stated that their taxes were "fair".\n * **Analysis:** Asking respondents about the specific dollar amount yields significantly more negative sentiment than asking whether their tax burden is fair.\n\n# Response Errors and the Randomized Response Technique\n\n* **Classification of Response Errors**\n * **Accidental Response Error:** Occurs when subjects genuinely cannot recall precise details and resort to guessing or estimating.\n * Example Prompt: "How many times did you shop at Publix during the month of July?"\n * Behavioral Reality: Humans are poor estimators and consistently over- or underestimate historical behavior.\n * **Purposeful Response Error:** Occurs when subjects knowingly provide false answers to avoid embarrassment, legal consequences, or social stigma.\n * Example Prompt: "As an athlete, have you ever taken performance enhancing drugs that are prohibited?"\n * Behavioral Reality: Direct questioning regarding illicit behavior leads to severe underestimation of the population proportion p$.

    • Randomized Response Technique (RRT)

  • Designed to ensure complete anonymity when asking sensitive or potentially incriminating questions.

  • Procedure:

    • Participants are presented with two questions:

      • Q1Q_1 (Sensitive): "As an athlete, have you ever taken prohibited performance enhancing drugs?"

      • Q2Q_2 (Baseline/Neutral): "Are there seven days in a week?"

    • Participants randomly select a number between 11 and 1010 (e.g., rolling a die or using a random number generator).

    • If the random number selected is odd, the participant answers Q1Q_1.

    • If the random number selected is even, the participant answers Q2Q_2.

    • The interviewer records only a "Yes" or "No" without knowing which question the participant was assigned.

  • Analytical Benefit: Preserves complete respondent privacy, encouraging truthful answers and enabling researchers to calculate unbiased population proportion estimates using known probability distributions.

Probability Sampling Designs

  • Simple Random Sample (SRS)

    • Every distinct sample of size nn has an equal probability of selection from the population.

  • Stratified Random Sample

    • Step I: Divide the total population into mutually exclusive subgroups called strata, where individuals within each stratum share specific characteristics.

    • Step II: Draw an independent Simple Random Sample (SRS) from every stratum.

    • Types: Proportionate sampling (stratum sample size is proportional to population size) vs. Disproportionate sampling (stratum sample size is intentional to oversample rare groups).

Stratified random sampling process flowchartComparison of proportionate and disproportionate stratified sampling
  • Empirical Example: 1990 Houston Intravenous Drug Study

    • Objective: Estimate HIV prevalence in heterosexual males who use intravenous drugs across racial groups.

HIV prevalence stratified by race in Houston intravenous drug study
* Data breakdown:
  * Hispanic Stratum: n1=107n_1 = 107, Infected = 44, Sample proportion p^1=3.7%\hat{p}_1 = 3.7\%
  * White Stratum: n2=214n_2 = 214, Infected = 1414, Sample proportion p^2=6.5%\hat{p}_2 = 6.5\%
  * Black Stratum: n3=600n_3 = 600, Infected = 5959, Sample proportion p^3=9.8%\hat{p}_3 = 9.8\%
  * Combined Sample: Total n=921n = 921, Total Infected = 7777, Combined Proportion p^=8.4%\hat{p} = 8.4\%
* Mathematical Warning: Individual subgroup proportions cannot simply be summed or averaged directly without taking overall sample weights into account.
  • Cluster Random Sample

    • Step I: Divide the target population into naturally occurring sub-units or groups called clusters.

    • Step II: Randomly select a subset of full clusters, and sample individuals within the selected clusters.

Cluster sampling process diagram
  • Applied Example: Studying teenager eating habits in Columbia, SC.

    • Define 33 high schools as distinct clusters.

    • Randomly select 6 high schools out of the 33, and collect data from students within those 6 selected schools.

Comparative Analysis: Stratified vs. Cluster Sampling

  • Stratified Sampling Characteristics

    • Subdivides population based on shared homogeneous characteristics (e.g., sex, race, age, income bracket).

    • Primary Purpose: Ensures adequate representation of key demographic subgroups and enables robust cross-group comparative analysis.

    • Data Collection: Samples are drawn from every established stratum.

  • Cluster Sampling Characteristics

    • Subdivides population into miniature heterogeneous representations of the target population.

    • Primary Purpose: Administrative efficiency, logistical convenience, and cost reduction.

    • Data Collection: Only a subset of selected clusters is sampled; non-selected clusters are ignored.

Systematic and Multistage Combination Sampling

  • Systematic Random Sample

    • Selection occurs at fixed, pre-determined numerical increments from an ordered sampling frame.

    • Examples: Selecting every 10th10\text{th} retail customer, every 50th50\text{th} registered student, or every 100th100\text{th} unique website visitor.

    • Benefit: Removes subject selection bias from field researchers.

  • Multistage Combination Sampling

    • Combines multiple probability sampling methodologies sequentially to solve complex field research challenges.

    • Applied Example: Sampling 4 million4\text{ million} adult residents of South Carolina:

    • Stage 1 (Cluster Sampling): Use the 4646 counties of South Carolina as clusters; randomly select a sample of counties.

    • Stage 2 (Stratified Sampling): Within each randomly selected county, perform stratified random sampling based on demographic variables (e.g., sex, race, political affiliation).