Data Collection Principles

Lesson 1: Introduction

  • This lesson introduces the fundamental principles of data collection and different study types.
  • It emphasizes understanding the strengths and weaknesses of various data collection methods to evaluate the credibility of research findings.
  • Learning Objective: Identify methods of data collection.
  • Essential Question: How do we define and differentiate the fundamental terms and concepts in data collection, and why are these distinctions important for understanding and interpreting data?

Populations and Samples

  • The lesson explores how samples are chosen and their impact on the conclusions drawn from the data.
  • It is crucial to distinguish between a population and a sample to evaluate data-driven claims effectively.
  • Population: The entire group of individuals or objects of interest in a study.
    • Example: All high school students in the United States.
  • Sample: A smaller, manageable group selected from the population to represent the larger group.
    • Example: A group of 500 high school students from different states across the U.S.
  • A researcher studying the sleeping habits of university students in the U.S. provides another example.
    • Population: All university students enrolled in universities across the United States.
    • Sample: 1,000 university students selected from various universities across the United States.
  • The "Sample Considerations" video from LinkedIn Learning (up to 2:58) explores challenges in selecting a representative sample, including:
    • Sample size
    • Selection methods
    • Potential biases
    • Importance of random sampling

Describing Data Characteristics

  • Numerical values describing populations and samples are known as parameters and statistics.
  • Parameter: A numerical value describing an entire population's characteristics.
    • It is a fixed but often unknown value.
    • Example: The average height of all adult women in the United States.
  • Statistic: A numerical value that describes a characteristic of a sample.
    • It is calculated from sample data and used to estimate the population parameter.
    • Example: The average height of 100 adult women randomly selected from the United States.
  • The accuracy of a statistic as an estimate of a population parameter depends on how well the sample represents the population.
  • Biased or non-representative samples may not accurately reflect the true population parameter; using appropriate sampling techniques is crucial.

The Raw Materials of Statistical Analysis

  • Understanding the basic elements that make up a dataset is important:
    • Individuals
    • Variables
    • Data
  • Individuals: The objects described by a set of data (people, animals, or things).
  • Variables: Characteristics or measurements of interest, which can be quantitative or categorical.
    • Quantitative variables: Numerical measurements (age, height, weight, income).
    • Categorical variables: Categories or labels (gender, race, occupation, favorite color).
  • Data: The actual values of the variables; "datum" is the singular form.
  • Example: A medical study
    • Each row represents an individual (a patient)
    • Each column represents a variable (age, gender, medication dosage, side effects)
    • The cells contain the data (quantitative and categorical)
    • For instance, Patient 1 is a 35-year-old male who received a 100mg dose and experienced nausea and headaches.

Lesson 1.2: Common Sampling Methods

  • Essential Question: How do different sampling methods impact the representativeness and reliability of data collected from a population?
  • Sampling is a fundamental process that allows researchers to draw conclusions about a population without examining every individual.
  • The key to effective sampling is choosing a method that ensures the sample is representative of the population.

Sampling and Simple Random Sampling

  • Simple Random Sample: All individuals are put into a single list, and we randomly select from that list until we reach the desired sample size.
  • For a particular sample size, nn, any combination of nn individuals is equally likely to be selected.
  • This helps ensure the sample is not skewed towards any particular group or characteristic within the population.
  • Example: For a high school survey, simple random sampling ensures that every combination of 100 students at City High School has an equal chance of being selected.
  • Using a simple random sample is one of the most straightforward and effective ways to produce data that is likely to be representative of the population.

Other Random Sampling Methods

  • Stratified Sampling:
    • Involves dividing the population into distinct subgroups called strata, based on specific characteristics (age, gender, socioeconomic status).
    • A random sample is taken from each stratum in proportion to its representation in the overall population.
    • Example: A high school population divided into strata based on grade level (freshmen, sophomores, juniors, seniors). If 25% of the high school population is freshmen, then 25% of the sample should also be freshmen.
  • Cluster Sampling:
    • Involves dividing the population into clusters, which are naturally occurring groups (schools, neighborhoods, or cities).
    • Researchers randomly select a few clusters and include all individuals within those selected clusters in the sample.
    • Example: A researcher studying the opinions of high school students in a state might randomly select five schools and survey all students within those schools.
  • Systematic Sampling:
    • Involves selecting every nthn^{th} individual from a list of the population, starting from a randomly chosen point.
    • Example: In a list of 1000 people, if you want a sample of 100, the sampling interval would be 10 (every 10th person is selected).
  • Stratified sampling takes some members from all groups, while cluster sampling takes all members from some of the groups and none from the others.

Lesson 1.3: Types of Studies

  • Essential Question: What are the key characteristics of observational studies, sample surveys, and experiments?
  • Effective study design involves considering how we will collect information about our sample.
  • Key terms:
    • Explanatory variable (independent variable): A variable that may cause a change in another variable.
    • Response variable (dependent variable): The affected variable.

Types of Studies

  • Observational Studies:
    • Researchers observe and record data on variables as they naturally occur, without any intervention or manipulation.
    • Example: A researcher observes and records the eating habits and weight of a group of individuals over a year to see if there is a relationship between diet and weight gain.
  • Sample Surveys:
    • A specific type of observational study where individuals self-report the values of variables, often by providing their opinions or answering questions.
    • Example: Researchers select a random sample of 1,750 U.S. eligible voters and collect data on their opinions regarding their political preferences.
  • Experiments:
    • Researchers intentionally manipulate one or more variables (the explanatory variables) to observe their effect on another variable (the response variable).
    • Participants are randomly assigned to different groups, with each group receiving a different treatment or level of the explanatory variable.
    • Example: A researcher randomly assigns participants to two groups: one group receives a new drug for high blood pressure, while the other group receives an inactive treatment (placebo). The researcher then compares the blood pressure readings of the two groups.

Identifying Study Designs

  • Being able to distinguish between observational studies, sample surveys, and experiments is a key component of evaluating the validity and reliability of research findings.
  • Example: Researchers want to investigate how various exercises affect heart rates; they assemble a group of 100 members of a local gym as participants.
    • Experiment: The researchers assign each participant to one of three exercise groups (running, weightlifting, yoga) and measure heart rates before and after.
    • Observational Study: The researchers give each participant a heart rate monitor and observe their normal exercise behavior and the effect on heart rate.
    • Sample Survey: The researchers ask people leaving the gym to rate how much they think various exercises elevate their heart rate on a scale of 1-10.
  • Experiments provide a more controlled environment for investigating potential cause-and-effect relationships.

Lesson 1.4: Experimental Design

  • Essential Question: What are the key principles involved in designing an effective experiment?
  • The purpose of an experiment is to investigate the relationship between two variables in a controlled environment to prevent other factors from influencing the variables.
  • In a randomized experiment, researchers manipulate the explanatory variable and measure the resulting changes in the response variable.
    • Treatments: The different values of the explanatory variable.
    • Experimental unit: A single object or individual being measured.

Key Principles of Experimental Design

  • Randomization:
    • Participants are randomly assigned to different groups to minimize bias and ensure that any observed differences between groups are due to the treatment, not other factors (confounding variables).
    • Example: Participants are randomly assigned to groups where one group receives a new drug, and the other does not.
  • Replication:
    • Repeating the experiment with a sufficiently large sample size or reproducing the entire study to confirm previous findings.
    • Example: A researcher conducts a study on the effects of a new teaching method on student achievement and repeats the study with multiple classrooms.
  • Control Groups:
    • A group of participants in an experiment who do not receive the experimental treatment.
    • They might receive no treatment or a placebo.
    • Example: In a study testing the effectiveness of a new fertilizer, one group of plants receives the fertilizer, while another group (the control group) does not.

Minimizing Bias in Experiments

  • Control groups are essential for establishing a baseline for comparison and isolating the actual effect of the treatment.
  • The Placebo Effect:
    • Sometimes patients improve solely because they believe they are receiving treatment.
    • Researchers often administer a placebo to the control group.
  • Blinding and Double-Blinding:
    • Blinding is a technique used to prevent participants and/or researchers from knowing who is receiving the experimental treatment and who is receiving the placebo.
    • Single-Blind: Only the participants are unaware of their group assignment.
    • Double-Blind: Both the participants and the researchers interacting with them are unaware of the group assignments.
  • The most reliable way to determine whether the explanatory variable is actually causing changes in the response variable is to carry out a randomized, controlled, double-blind experiment.