Understanding Populations, Sampling & Data Collection

Terminology & Key Definitions

  • Population: Entire set of observations or measurements under study; often very large or even infinite.
    • Examples
    • Australian Census counts ≈23.64 million\approx 23.64\text{ million} residents (finite).
    • Number of fish in Brisbane River (conceptually infinite).
  • Sample: A subset selected from the population; always smaller.
    • Examples
    • 100100 Australians drawn from the full population.
    • Today’s catch of a single fisherman.
  • Census: Investigation of every element of a population.
  • Parameter (population characteristic)
    • Symbols & meaning
    • μ\mu – population mean
    • σ\sigma – population standard deviation
    • pp – population proportion
    • Examples
    • Average annual income of all Australians.
    • Proportion of defective shoes produced in a factory.
  • Statistic (sample descriptive measure)
    • Symbols & meaning
    • Xˉ\bar X – sample mean
    • ss – sample standard deviation
    • p^\hat p – sample proportion
    • Examples
    • Mean income of 100100 randomly selected Australians.
    • Defect rate in 200200 sampled shoes.
  • Relationship
    • Parameters ↔ population; Statistics ↔ sample.

Representation, Inclusion & Exclusion

  • Researchers must describe the sample and justify why/how it was chosen.
  • Core concern: Representativeness—degree to which the sample mirrors the population.
  • Inclusion criteria: Conditions participants must satisfy to enter the study.
  • Exclusion criteria: Conditions leading to participant omission.
  • Clear criteria sharpen focus and data quality.

Why Sample Instead of Census?

  • Large populations often impossible to study fully due to:
    • Cost constraints
    • Time constraints
    • Analytical practicality / accuracy limits
    • General impracticality
  • Well-designed samples allow inference about population parameters from sample statistics.

The Sampling Frame

  • Probability sampling requires that each population member has equal selection probability.
  • Necessitates a complete list/map/chart of population—called the sampling frame or working population.
  • Samples are randomly drawn from this frame.

Probability Sampling

  • Goal: Statistically representative, findings generalisable to entire population.
  • Grounded in probability theory.
  • Main types
    1. Simple Random Sampling (SRS)
    2. Stratified Random Sampling
    3. Cluster Sampling
  • Proper use yields high precision with small sample fractions.
1. Simple Random Sampling (SRS)
  • Every population member equally likely to be chosen.
  • Procedure
    • Assign unique numbers to each element.
    • Select numbers using random-number table, computer generator, or “names-in-hat.”
  • When to use
    • Only a basic list of names exists, no additional grouping info (e.g., alphabetical list of CDU students).
    • No compelling reason for stratification or clustering.
2. Stratified Random Sampling
  • Population divided into non-overlapping strata; SRS performed within each stratum.
  • Enables estimates for:
    • Whole population
    • Each stratum
    • Inter-stratum relationships
  • Example (tourist gender):
    • Population composition – 40%40\% male, 60%60\% female.
    • Desired n=200n=200 ⇒ select 0.4×200=800.4\times200=80 males & 0.6×200=1200.6\times200=120 females via SRS inside each gender.
  • Typical strata bases: gender, age bands, occupation categories, etc.
3. Cluster Sampling
  • Random sample of groups (clusters) rather than individuals.
  • Useful when
    • Complete population list difficult/costly.
    • Population geographically dispersed.
  • Example (Darwin household income)
    1. Identify 5050 suburbs → each acts as a cluster.
    2. Randomly pick, say, 33 suburbs.
    3. Survey all households within chosen suburbs.

Non-Probability Sampling

  • Selected to illustrate phenomenon; not statistically representative.
  • Employed when no full sampling frame exists (e.g., Cosmopolitan magazine readers).
  • Key techniques
    • Judgemental/Purposive: Researcher hand-picks key informants able to contribute rich insight.
    • Quota: Fill pre-set category quotas (age, gender, etc.).
    • Convenience: Engage easiest accessible participants (shoppers, passers-by).
    • Snowball: Each participant recommends the next; continues until sample complete, meeting inclusion criteria.

Errors in Data Collection

Two broad sources:

  1. Sampling Error
  2. Non-Sampling Error (Systematic)
Sampling Error
  • Difference between sample statistic and population parameter.
  • Formula: Sampling Error=Xˉ−μorμ−Xˉ\text{Sampling Error}=\bar X-\mu \quad\text{or}\quad \mu-\bar X
  • Depends on sampling method; reducible by increasing sample size or better design.
  • Numerical illustration
    • Population incomes 10,20,30,40,50{10, 20, 30, 40, 50} ($000\$000)
    • μ=30\mu=30
    • Random sample 10,30{10,30} ⇒ Xˉ=20\bar X=20
    • Error =20−30=−10=20-30=-10 ($10 000\$10\,000 under-estimate).
Non-Sampling Errors
  • Mistakes in design, execution, measurement.
  • Not mitigated by larger sample size.
  • Categories & examples
    1. Data acquisition errors – mis-recording, misinterpretation.
    2. Non-response error – sampled units unavailable/refuse.
    3. Response bias – respondents slant answers consciously/unconsciously.
    4. Selection error – some population members never reach sampling frame (e.g., no telephone).
  • Generally more serious than sampling error because persistence unaffected by sample size.

Data Collection Methods Overview

  • Data = backbone of statistical analysis; reliability & accuracy hinge on collection method.
  • Two broad approaches
    • Census – whole population.
    • Sample survey – subset.
  • Most common: Survey using a questionnaire.
Survey Administration Modes
  1. Interviewer-Administered
    • Personal interview
    • Telephone interview
  2. Self-Administered
    • Postal (mail)
    • Delivery & collection (drop-off/pick-up)
    • Online / electronic
  3. Direct Observation
Response Rates
  • Definition: Count of valid, properly completed responses.
  • Higher rates ⇒ more complete data, stronger generalisability.
  • Researchers actively pursue strategies (follow-ups, incentives) to maximise responses.

Personal Interview (Face-to-Face)

  • Advantages
    • Highest potential response rate (often 100%100\%).
    • Interviewer clarifies questions ⇒ fewer errors; high data quality.
  • Disadvantages
    • Expensive (travel, personnel).
    • Poorly trained interviewer may bias or mis-record data.
    • Hard to access certain populations (e.g., street kids).
  • Remedy: Rigorous interviewer training; standardised protocols.

Telephone Interview

  • Advantages
    • Cheaper & faster than face-to-face.
    • Reaches wide geography; typical response ≈80%\approx80\%.
  • Disadvantages
    • Slightly lower response rate than personal.
    • Coverage limited to phone owners.
    • Same interviewer bias risk.
  • Remedies
    • Train interviewers, offer incentives, call at optimal times.

Self-Administered Surveys

  • Types: Electronic/online, postal, delivery & collection.
  • Advantages
    • Least expensive per respondent; feasible for large, dispersed samples.
    • Rapid turnaround (especially online).
  • Disadvantages
    • Generally low response rates.
    • Higher mis-understanding risk causing incorrect answers.
    • Cost escalates if multiple follow-ups needed.
  • Remedies
    • Drop-off & pick-up strategy to boost engagement & clarify queries.
    • Incentives; prepaid return envelopes; reminder contacts.

Direct Observation

  • Collect data by observing individuals/units directly (e.g., health conditions of Darwin residents).
  • No control over influencing factors; relatively low cost.

Ethical & Practical Implications

  • Transparency in sampling explanation safeguards study credibility.
  • Inclusion/exclusion criteria prevent unethical participant selection and enhance focus.
  • Non-probability sampling demands honesty about limited generalisability.
  • Minimising non-sampling errors (through training, pilot tests, robust instrument design) is ethically critical to avoid misinformation.

Mathematical & Statistical Connections

  • Sampling theory underpins inferential statistics (estimating μ,σ,p\mu, \sigma, p via Xˉ,s,p^\bar X, s, \hat p).
  • Stratified design improves estimator precision, analogous to weighted means where stratum weights correspond to population proportions.
  • Cluster sampling relates to multi-stage designs and intraclass correlation considerations in variance estimation.
  • Error taxonomy aligns with total survey error framework used in large-scale official statistics (e.g., ABS Census quality reports).