Understanding Populations, Sampling & Data Collection
Terminology & Key Definitions
- Population: Entire set of observations or measurements under study; often very large or even infinite.
- Examples
- Australian Census counts residents (finite).
- Number of fish in Brisbane River (conceptually infinite).
- Sample: A subset selected from the population; always smaller.
- Examples
- Australians drawn from the full population.
- Today’s catch of a single fisherman.
- Census: Investigation of every element of a population.
- Parameter (population characteristic)
- Symbols & meaning
- – population mean
- – population standard deviation
- – population proportion
- Examples
- Average annual income of all Australians.
- Proportion of defective shoes produced in a factory.
- Statistic (sample descriptive measure)
- Symbols & meaning
- – sample mean
- – sample standard deviation
- – sample proportion
- Examples
- Mean income of randomly selected Australians.
- Defect rate in sampled shoes.
- Relationship
- Parameters ↔ population; Statistics ↔ sample.
Representation, Inclusion & Exclusion
- Researchers must describe the sample and justify why/how it was chosen.
- Core concern: Representativeness—degree to which the sample mirrors the population.
- Inclusion criteria: Conditions participants must satisfy to enter the study.
- Exclusion criteria: Conditions leading to participant omission.
- Clear criteria sharpen focus and data quality.
Why Sample Instead of Census?
- Large populations often impossible to study fully due to:
- Cost constraints
- Time constraints
- Analytical practicality / accuracy limits
- General impracticality
- Well-designed samples allow inference about population parameters from sample statistics.
The Sampling Frame
- Probability sampling requires that each population member has equal selection probability.
- Necessitates a complete list/map/chart of population—called the sampling frame or working population.
- Samples are randomly drawn from this frame.
Probability Sampling
- Goal: Statistically representative, findings generalisable to entire population.
- Grounded in probability theory.
- Main types
- Simple Random Sampling (SRS)
- Stratified Random Sampling
- Cluster Sampling
- Proper use yields high precision with small sample fractions.
1. Simple Random Sampling (SRS)
- Every population member equally likely to be chosen.
- Procedure
- Assign unique numbers to each element.
- Select numbers using random-number table, computer generator, or “names-in-hat.”
- When to use
- Only a basic list of names exists, no additional grouping info (e.g., alphabetical list of CDU students).
- No compelling reason for stratification or clustering.
2. Stratified Random Sampling
- Population divided into non-overlapping strata; SRS performed within each stratum.
- Enables estimates for:
- Whole population
- Each stratum
- Inter-stratum relationships
- Example (tourist gender):
- Population composition – male, female.
- Desired ⇒ select males & females via SRS inside each gender.
- Typical strata bases: gender, age bands, occupation categories, etc.
3. Cluster Sampling
- Random sample of groups (clusters) rather than individuals.
- Useful when
- Complete population list difficult/costly.
- Population geographically dispersed.
- Example (Darwin household income)
- Identify suburbs → each acts as a cluster.
- Randomly pick, say, suburbs.
- Survey all households within chosen suburbs.
Non-Probability Sampling
- Selected to illustrate phenomenon; not statistically representative.
- Employed when no full sampling frame exists (e.g., Cosmopolitan magazine readers).
- Key techniques
- Judgemental/Purposive: Researcher hand-picks key informants able to contribute rich insight.
- Quota: Fill pre-set category quotas (age, gender, etc.).
- Convenience: Engage easiest accessible participants (shoppers, passers-by).
- Snowball: Each participant recommends the next; continues until sample complete, meeting inclusion criteria.
Errors in Data Collection
Two broad sources:
- Sampling Error
- Non-Sampling Error (Systematic)
Sampling Error
- Difference between sample statistic and population parameter.
- Formula:
- Depends on sampling method; reducible by increasing sample size or better design.
- Numerical illustration
- Population incomes ()
- Random sample ⇒
- Error ( under-estimate).
Non-Sampling Errors
- Mistakes in design, execution, measurement.
- Not mitigated by larger sample size.
- Categories & examples
- Data acquisition errors – mis-recording, misinterpretation.
- Non-response error – sampled units unavailable/refuse.
- Response bias – respondents slant answers consciously/unconsciously.
- Selection error – some population members never reach sampling frame (e.g., no telephone).
- Generally more serious than sampling error because persistence unaffected by sample size.
Data Collection Methods Overview
- Data = backbone of statistical analysis; reliability & accuracy hinge on collection method.
- Two broad approaches
- Census – whole population.
- Sample survey – subset.
- Most common: Survey using a questionnaire.
Survey Administration Modes
- Interviewer-Administered
- Personal interview
- Telephone interview
- Self-Administered
- Postal (mail)
- Delivery & collection (drop-off/pick-up)
- Online / electronic
- Direct Observation
Response Rates
- Definition: Count of valid, properly completed responses.
- Higher rates ⇒ more complete data, stronger generalisability.
- Researchers actively pursue strategies (follow-ups, incentives) to maximise responses.
Personal Interview (Face-to-Face)
- Advantages
- Highest potential response rate (often ).
- Interviewer clarifies questions ⇒ fewer errors; high data quality.
- Disadvantages
- Expensive (travel, personnel).
- Poorly trained interviewer may bias or mis-record data.
- Hard to access certain populations (e.g., street kids).
- Remedy: Rigorous interviewer training; standardised protocols.
Telephone Interview
- Advantages
- Cheaper & faster than face-to-face.
- Reaches wide geography; typical response .
- Disadvantages
- Slightly lower response rate than personal.
- Coverage limited to phone owners.
- Same interviewer bias risk.
- Remedies
- Train interviewers, offer incentives, call at optimal times.
Self-Administered Surveys
- Types: Electronic/online, postal, delivery & collection.
- Advantages
- Least expensive per respondent; feasible for large, dispersed samples.
- Rapid turnaround (especially online).
- Disadvantages
- Generally low response rates.
- Higher mis-understanding risk causing incorrect answers.
- Cost escalates if multiple follow-ups needed.
- Remedies
- Drop-off & pick-up strategy to boost engagement & clarify queries.
- Incentives; prepaid return envelopes; reminder contacts.
Direct Observation
- Collect data by observing individuals/units directly (e.g., health conditions of Darwin residents).
- No control over influencing factors; relatively low cost.
Ethical & Practical Implications
- Transparency in sampling explanation safeguards study credibility.
- Inclusion/exclusion criteria prevent unethical participant selection and enhance focus.
- Non-probability sampling demands honesty about limited generalisability.
- Minimising non-sampling errors (through training, pilot tests, robust instrument design) is ethically critical to avoid misinformation.
Mathematical & Statistical Connections
- Sampling theory underpins inferential statistics (estimating via ).
- Stratified design improves estimator precision, analogous to weighted means where stratum weights correspond to population proportions.
- Cluster sampling relates to multi-stage designs and intraclass correlation considerations in variance estimation.
- Error taxonomy aligns with total survey error framework used in large-scale official statistics (e.g., ABS Census quality reports).