1.2 Data Collection: Population, Sampling, and Bias

1.2 Data Collection: Key Concepts and Methods

Objectives

  • Define Population and Sample

    • Understand the statistical meaning of a population.

    • Recognize the relationship between a sample and the population.

    • Identify both the population and sample in various scenarios.

  • Identify Sampling Methods

    • Distinguish between common sampling procedures: simple random, stratified, cluster, systematic, and convenience sampling.

    • Recognize the advantages and limitations specific to each method.

    • Match given data collection descriptions to the appropriate sampling method.

  • Recognize Bias in Data Collection

    • Understand the definition of bias in statistics.

    • Identify potential sources of bias, including selection, measurement, and nonresponse bias.

    • Evaluate whether a data collection method is likely to produce reliable and representative results.

Research Begins with a Question

Research always starts with a question, guiding the data collection process.

  • Newborn Scenario Examples:

    • What is the average birth weight of newborns in a specific hospital?

    • How many newborns exhibit a normal heart rate during their first week of life?

    • How many newborns sleep well in their first week?

Population vs. Sample

  • Population (NN):

    • The full set of observations of interest.

    • Newborn Scenario Example: All newborns born in that week at the hospital.

  • Sample (nn):

    • A subset of the population.

    • Newborn Scenario Example: The newborns a doctor is able to record data for during a particular week.

  • Goal: To examine variables within the entire population by studying a manageable sample.

  • Inference: If a sample accurately represents the population, statistical techniques use information from the sample to make inferences or generalizations about the population.

    • Newborn Scenario Implication: If the characteristics observed in the sample of newborns (e.g., low average birthweight) are representative of all newborns at the hospital, then this characteristic can be generalized to the larger population, suggesting a hospital-wide trend of lower birthweights.

Importance of Variable Classification (Review from 1.1)

Correctly classifying variables (by type and level of measurement) is crucial because:

  • Statistical inference tests often require specific types of variables.

  • Accurate classification ensures that appropriate statistical methods are applied.

  • Misclassification can lead to invalid or misleading results and conclusions.

Data Collection and Sampling Techniques

How data is collected is as important as the types of variables being measured. Poor data collection can introduce bias, even if variables are correctly classified, thereby reducing the reliability of conclusions.

  • Minimizing Bias: Understanding data collection methods is essential to minimize bias and ensure the data accurately represents the population.

  • Improving Reliability and Validity: Proper data collection enhances the reliability and validity of results, facilitating meaningful conclusions and informed decision-making.

Random Sampling Methods

These methods aim to give every unit in the population a known, often equal, chance of selection, thus increasing the likelihood of a representative sample.

  1. Simple Random Sample (SRS)

    • Definition: Every possible sample of size nn has an equal chance of being selected.

    • Example: Randomly selecting newborns from a comprehensive list of all births that week.

    • Advantages: Often considered the best method as it provides every individual an equal chance, making it most likely to produce a truly representative sample.

    • Limitations: Can be difficult and impractical to carry out a true SRS in real-world scenarios.

  2. Stratified Sampling

    • Definition: The population is divided into distinct, non-overlapping groups (strata) based on a shared characteristic, and then a sample is drawn from each stratum.

    • Example: Dividing newborns into strata of males and females, and then sampling an equal number from each group.

  3. Cluster Sampling

    • Definition: The population is divided into clusters (usually naturally occurring groups), a selection of clusters is made, and all members within the selected clusters are included in the sample.

    • Example: Choosing all newborns from a particular hospital ward to form the sample.

  4. Systematic Sampling

    • Definition: Selecting every kthk^{th} observation from an ordered list of the entire population (NN), starting from a random point.

    • Example: Recording every 5th5^{th} baby born that week.

Non-Random Sampling Methods

These methods do not rely on random selection and are often prone to bias, as some individuals may have a higher chance of being selected than others.

  1. Convenience Sampling

    • Definition: Choosing observations that are easiest to access or readily available to the researcher.

    • Example: Recording only newborns in the NICU who are closest to the nurse's station.

  2. Voluntary Sampling

    • Definition: Observations self-select or volunteer to participate in the study (less common for newborns but applicable to parents/guardians).

    • Example: Parents agreeing to have their baby measured for a study at a hospital booth.

Special Sampling Methods
  1. Multi-Stage Sampling

    • Definition: Always begins with cluster sampling in the initial stage, but subsequent stages can utilize any other sampling method (e.g., stratified, systematic, or simple random sampling).

    • Example: Randomly selecting provinces (first stage - cluster), then within those provinces, randomly selecting school districts (second stage - cluster), and finally, within each chosen school, stratifying students by grade level before sampling from each grade (third stage - stratified).

  2. Mixed (Hybrid) Sampling

    • Definition: Combines different sampling techniques without necessarily starting with cluster sampling.

    • Example: Dividing newborns into groups of preemies and full-term babies (stratification), then selecting every 4th4^{th} preemie and every 6th6^{th} full-term baby (systematic sampling within strata).

General Principle for Sampling

The closer the selection process is to being truly random, the more likely the resulting sample will accurately represent the population and minimize bias. Therefore, the