Data Collection Principles
Lesson 1: Introduction
- This lesson introduces the fundamental principles of data collection and different study types.
- It emphasizes understanding the strengths and weaknesses of various data collection methods to evaluate the credibility of research findings.
- Learning Objective: Identify methods of data collection.
- Essential Question: How do we define and differentiate the fundamental terms and concepts in data collection, and why are these distinctions important for understanding and interpreting data?
Populations and Samples
- The lesson explores how samples are chosen and their impact on the conclusions drawn from the data.
- It is crucial to distinguish between a population and a sample to evaluate data-driven claims effectively.
- Population: The entire group of individuals or objects of interest in a study.
- Example: All high school students in the United States.
- Sample: A smaller, manageable group selected from the population to represent the larger group.
- Example: A group of 500 high school students from different states across the U.S.
- A researcher studying the sleeping habits of university students in the U.S. provides another example.
- Population: All university students enrolled in universities across the United States.
- Sample: 1,000 university students selected from various universities across the United States.
- The "Sample Considerations" video from LinkedIn Learning (up to 2:58) explores challenges in selecting a representative sample, including:
- Sample size
- Selection methods
- Potential biases
- Importance of random sampling
Describing Data Characteristics
- Numerical values describing populations and samples are known as parameters and statistics.
- Parameter: A numerical value describing an entire population's characteristics.
- It is a fixed but often unknown value.
- Example: The average height of all adult women in the United States.
- Statistic: A numerical value that describes a characteristic of a sample.
- It is calculated from sample data and used to estimate the population parameter.
- Example: The average height of 100 adult women randomly selected from the United States.
- The accuracy of a statistic as an estimate of a population parameter depends on how well the sample represents the population.
- Biased or non-representative samples may not accurately reflect the true population parameter; using appropriate sampling techniques is crucial.
The Raw Materials of Statistical Analysis
- Understanding the basic elements that make up a dataset is important:
- Individuals: The objects described by a set of data (people, animals, or things).
- Variables: Characteristics or measurements of interest, which can be quantitative or categorical.
- Quantitative variables: Numerical measurements (age, height, weight, income).
- Categorical variables: Categories or labels (gender, race, occupation, favorite color).
- Data: The actual values of the variables; "datum" is the singular form.
- Example: A medical study
- Each row represents an individual (a patient)
- Each column represents a variable (age, gender, medication dosage, side effects)
- The cells contain the data (quantitative and categorical)
- For instance, Patient 1 is a 35-year-old male who received a 100mg dose and experienced nausea and headaches.
Lesson 1.2: Common Sampling Methods
- Essential Question: How do different sampling methods impact the representativeness and reliability of data collected from a population?
- Sampling is a fundamental process that allows researchers to draw conclusions about a population without examining every individual.
- The key to effective sampling is choosing a method that ensures the sample is representative of the population.
Sampling and Simple Random Sampling
- Simple Random Sample: All individuals are put into a single list, and we randomly select from that list until we reach the desired sample size.
- For a particular sample size, n, any combination of n individuals is equally likely to be selected.
- This helps ensure the sample is not skewed towards any particular group or characteristic within the population.
- Example: For a high school survey, simple random sampling ensures that every combination of 100 students at City High School has an equal chance of being selected.
- Using a simple random sample is one of the most straightforward and effective ways to produce data that is likely to be representative of the population.
Other Random Sampling Methods
- Stratified Sampling:
- Involves dividing the population into distinct subgroups called strata, based on specific characteristics (age, gender, socioeconomic status).
- A random sample is taken from each stratum in proportion to its representation in the overall population.
- Example: A high school population divided into strata based on grade level (freshmen, sophomores, juniors, seniors). If 25% of the high school population is freshmen, then 25% of the sample should also be freshmen.
- Cluster Sampling:
- Involves dividing the population into clusters, which are naturally occurring groups (schools, neighborhoods, or cities).
- Researchers randomly select a few clusters and include all individuals within those selected clusters in the sample.
- Example: A researcher studying the opinions of high school students in a state might randomly select five schools and survey all students within those schools.
- Systematic Sampling:
- Involves selecting every nth individual from a list of the population, starting from a randomly chosen point.
- Example: In a list of 1000 people, if you want a sample of 100, the sampling interval would be 10 (every 10th person is selected).
- Stratified sampling takes some members from all groups, while cluster sampling takes all members from some of the groups and none from the others.
Lesson 1.3: Types of Studies
- Essential Question: What are the key characteristics of observational studies, sample surveys, and experiments?
- Effective study design involves considering how we will collect information about our sample.
- Key terms:
- Explanatory variable (independent variable): A variable that may cause a change in another variable.
- Response variable (dependent variable): The affected variable.
Types of Studies
- Observational Studies:
- Researchers observe and record data on variables as they naturally occur, without any intervention or manipulation.
- Example: A researcher observes and records the eating habits and weight of a group of individuals over a year to see if there is a relationship between diet and weight gain.
- Sample Surveys:
- A specific type of observational study where individuals self-report the values of variables, often by providing their opinions or answering questions.
- Example: Researchers select a random sample of 1,750 U.S. eligible voters and collect data on their opinions regarding their political preferences.
- Experiments:
- Researchers intentionally manipulate one or more variables (the explanatory variables) to observe their effect on another variable (the response variable).
- Participants are randomly assigned to different groups, with each group receiving a different treatment or level of the explanatory variable.
- Example: A researcher randomly assigns participants to two groups: one group receives a new drug for high blood pressure, while the other group receives an inactive treatment (placebo). The researcher then compares the blood pressure readings of the two groups.
Identifying Study Designs
- Being able to distinguish between observational studies, sample surveys, and experiments is a key component of evaluating the validity and reliability of research findings.
- Example: Researchers want to investigate how various exercises affect heart rates; they assemble a group of 100 members of a local gym as participants.
- Experiment: The researchers assign each participant to one of three exercise groups (running, weightlifting, yoga) and measure heart rates before and after.
- Observational Study: The researchers give each participant a heart rate monitor and observe their normal exercise behavior and the effect on heart rate.
- Sample Survey: The researchers ask people leaving the gym to rate how much they think various exercises elevate their heart rate on a scale of 1-10.
- Experiments provide a more controlled environment for investigating potential cause-and-effect relationships.
Lesson 1.4: Experimental Design
- Essential Question: What are the key principles involved in designing an effective experiment?
- The purpose of an experiment is to investigate the relationship between two variables in a controlled environment to prevent other factors from influencing the variables.
- In a randomized experiment, researchers manipulate the explanatory variable and measure the resulting changes in the response variable.
- Treatments: The different values of the explanatory variable.
- Experimental unit: A single object or individual being measured.
Key Principles of Experimental Design
- Randomization:
- Participants are randomly assigned to different groups to minimize bias and ensure that any observed differences between groups are due to the treatment, not other factors (confounding variables).
- Example: Participants are randomly assigned to groups where one group receives a new drug, and the other does not.
- Replication:
- Repeating the experiment with a sufficiently large sample size or reproducing the entire study to confirm previous findings.
- Example: A researcher conducts a study on the effects of a new teaching method on student achievement and repeats the study with multiple classrooms.
- Control Groups:
- A group of participants in an experiment who do not receive the experimental treatment.
- They might receive no treatment or a placebo.
- Example: In a study testing the effectiveness of a new fertilizer, one group of plants receives the fertilizer, while another group (the control group) does not.
Minimizing Bias in Experiments
- Control groups are essential for establishing a baseline for comparison and isolating the actual effect of the treatment.
- The Placebo Effect:
- Sometimes patients improve solely because they believe they are receiving treatment.
- Researchers often administer a placebo to the control group.
- Blinding and Double-Blinding:
- Blinding is a technique used to prevent participants and/or researchers from knowing who is receiving the experimental treatment and who is receiving the placebo.
- Single-Blind: Only the participants are unaware of their group assignment.
- Double-Blind: Both the participants and the researchers interacting with them are unaware of the group assignments.
- The most reliable way to determine whether the explanatory variable is actually causing changes in the response variable is to carry out a randomized, controlled, double-blind experiment.