Experimental Design and Sampling Bias Study Notes
Fundamentals of Experimental Design - Observational Studies vs. Experiments:
An observational study measures variables without manipulating the subjects or changing their environment.
An experiment is a controlled study where researchers intentionally change one or more explanatory variables (factors) to see how they affect a response variable.
Key Experimental Terminology:
Factor (Explanatory Variable): The main independent variable () changed by the researcher.
Response Variable: The outcome or dependent variable () measured to check the effect of the factor.
Treatment: A specific setting or combination of factor levels applied to test subjects.
Experimental Units / Subjects: The entities tested in the study. Humans are called subjects. Non-human units can be animals (like lab rats used in early safety tests) or objects.
Control Group: A baseline group given no active treatment, a placebo, or a standard treatment to serve as a benchmark.
Blinding Techniques and Placebo Controls
Placebos and Baseline Neutralization:
A placebo is a fake or inactive treatment (such as a plain sugar tablet made to look, feel, and taste identical to an active medicine pill) given to match the active treatment.
A placebo control group sets a baseline score and cancels out psychological changes (the placebo effect) that could distort findings.
Blinding Classifications:
Single-Blind Experiment: An experiment where the subject does not know whether they are getting the active treatment or the placebo.
Double-Blind Experiment: An experiment where neither the subject nor the researcher running the test knows who gets which treatment. This prevents researchers from accidentally swaying behavior or results.
Feasibility of Blinding:
Blinding is not always possible. For example, testing exercise on a treadmill versus no exercise cannot be blinded because subjects know if they are on a treadmill.
Blinding works in physical tests if items are unmarked and identical in look, such as testing unmarked treadmills from Brand A versus Brand B while researchers record data blindly.
Case Study: The CARDS Lipitor Experiment
Overview:
Study Title: Collaborative Atorvastatin Diabetes Study (CARDS).
Sponsor/Manufacturer: Pfizer (maker of Lipitor).
Target Population: Adults aged to years with type 2 diabetes and no past history of heart disease.
Sample Size: subjects.
Experimental Design:
Design Type: Randomized, double-blind, placebo-controlled experiment.
Treatment Allocation:
Active Group: subjects took of Lipitor daily.
Control Group: subjects took a daily placebo pill.
Duration: Subjects were tracked for years.
Response Variable: Whether a major heart event (such as a stroke) occurred (measured as a simple Yes/No qualitative outcome).
Quantitative Results:
Major Cardiovascular Events:
Lipitor Group: events.
Placebo Group: events.
Total Deaths:
Lipitor Group: deaths.
Placebo Group: deaths.
Conclusion: Taking of Lipitor daily significantly reduced major heart events and total deaths compared to the placebo.
Six-Step Process for Designing Experiments
Step 1: Identify the Problem to be Solved:
State a clear goal or research question.
Identify the response variable (the outcome to measure) and the target population.
Example: Testing the claim that the average GPA of all Wright College students is . Population: All Wright College students; Response variable: Average GPA.
Step 2: Determine Factors Affecting the Response Variable:
List all variables that could influence the response variable.
Example Factors for GPA: Class load, failing a test, study time, time management, sleep duration, home environment, attendance, tutoring use, motivation, health, and learning disabilities.
Group factors into controlled factors and uncontrolled factors (such as health conditions or pre-existing disabilities that researchers cannot alter).
Step 3: Determine the Number of Experimental Units:
Pick a sample size (). Larger samples give higher statistical reliability, but budget and time limits cap how many subjects can be included.
Step 4: Determine the Level of Each Factor:
Control:
Fixed Level: Keep a factor at a single fixed value if you are not testing its impact (for example, limiting subjects to a specific area, like living between Fullerton Avenue and Foster Avenue).
Varied Levels: Change a factor across specific levels if you want to evaluate its impact on the response variable.
Study Time: Set groups to hours/week, hours/week, hours/week, or hours/week.
Credit Hours: Set groups to , , , or credit hours.
Specific combinations of factor levels form the treatments.
Randomization:
Randomly assign subjects to treatment groups to balance out hidden or uncontrolled extra variables.
Step 5: Conduct the Experiment:
Replication: Repeat the experiment on many subjects to check that outcomes are reliable and repeatable.
Equal Group Sizes: Put equal numbers of subjects into each treatment group to make comparisons valid.
Collect raw data, prepare summary statistics, and show results in simple charts and tables.
Step 6: Test the Claim (Inferential Statistics):
Use statistical hypothesis testing to figure out if the collected evidence supports or rejects the original claim.
Overview of Bias in Sampling
Sampling Bias Defined:
Sampling bias occurs when the method used to pick a sample favors one group over another, making the sample unrepresentative of the overall population.
Example: Estimating the average GPA of all City Colleges of Chicago students by surveying only Wright College students creates sampling bias by leaving out the other six colleges.
Undercoverage:
Undercoverage occurs when a specific sub-group in the sample is smaller than its actual proportion in the whole population.
Extreme undercoverage leaves out whole sub-groups entirely ( representation).
Common in political polls that draw samples from unrepresentative geographic or demographic groups.
Types and Sources of Response Bias
Response Bias Defined:
Response bias occurs when survey answers do not reflect the true opinions, beliefs, or actions of respondents.
Sources of Response Bias:
Interviewer Error:
Happens when an interviewer's tone or body language pushes respondents toward a specific answer, or when an untrained interviewer fails to build trust.
Includes basic mistakes like giving wrong survey quotas to target groups.
Misrepresented Answers:
Happens when respondents intentionally lie or give incorrect info.
Very common when people are forced to take surveys against their will.
Example: UNLV requires students to complete faculty evaluations within a -week window or locks them out of Canvas. Forced participation leads students to pick random ratings (such as checking all s) without reading.
Wording of Questions:
Questions must be posed in a balanced form so they do not coax or steer respondents toward a specific answer. Changing a single word can alter results.
Ordering of Questions or Words:
Long surveys cause survey fatigue, making respondents answer later questions carelessly.
Early questions can prime respondents and influence their later answers.
Type of Question (Open vs. Closed):
Open Question: Lets respondents answer in their own words (short answer format). Example: "What is your college major?" with a blank text field.
Closed Question: Restricts respondents to a set list of choices (multiple-choice or checklist). Example: "What is your major? (a) Business, (b) History, (c) Sociology, (d) Nursing, (e) Psychology, (f) Other."
Using open questions in preliminary surveys helps identify common answers to design balanced closed questions later.
Data Entry Error:
Typing mistakes during data entry skew the statistics.
Example: Weekly driving miles for people are , , , , and miles. True average:
If a typist enters instead of , the data set becomes , , , , and , skewing the average to:Swapping digits (entering as ) distorts statistical results similarly.
Nonresponse Bias and Mitigation Strategies
Nonresponse Bias Defined:
Nonresponse bias happens when chosen people choose not to participate, and non-responders have systematically different opinions than those who do respond.
Common in voluntary email, web, or mail surveys.
Voluntary surveys usually attract people with extreme views (strongly in favor or strongly opposed), rendering results unrepresentative.
Strategies to Reduce Nonresponse Bias:
Rewards and Incentives: Offering small guaranteed prizes (such as a Amazon gift card) or entry into a raffle for large prizes (such as a gift card) increases completion rates.
Callbacks and Follow-ups: Using repeated phone callbacks, follow-up mailings, or automated text links prompts non-responders to reply over time.
Practice Examples and Conceptual Applications
Shopping Mall Food Court Study:
Scenario: A mall manager surveys the first customers entering the food court on weekday afternoons to evaluate food court expansion options.
Bias Type: Sampling bias (convenience sample leading to undercoverage of evening and weekend shoppers).
Fix: Survey shoppers randomly at varied times across both weekdays and weekends.
Homeschooling Survey:
Scenario: A polling group mails surveys to random households nationwide about homeschooling; only households respond.
Bias Type: Nonresponse bias (response rate is under ).
Fix: Conduct phone callbacks or in-person visits to non-responding households.
Email Survey with Expected Rate:
Scenario: A researcher emails random addresses expecting a response rate ( replies) to get a target sample size of .
Flaw: Suffers from undercoverage (excluding people without email/internet access or elderly populations) and nonresponse/voluntary response bias.
Subjective vs. Objective Survey Questions (Question Order Evaluation):
Scenario: Testing if question order creates bias between Question A ("What is the most important problem facing the nation?") and Question B ("Would you be willing to pay for social media services if your user data was your personal property?").
Analysis: In a poll of students, voted "Yes" (order affects results) and voted "No" (order does not affect results). Because the problem is subjective and open to interpretation, ambiguous questions are thrown out and credit is granted.
Pop-up Survey Incentive:
Scenario: A pop-up offers to complete a short health survey.
Tactic: Incentive/reward strategy used to raise response rates.
Movie Phone Survey:
Scenario: People selected for a phone survey about a new movie refuse to participate.
Bias Type: Nonresponse bias, because decliners may hold different opinions about the movie than participants.
Accuracy of Sampling Frames:
Reason for Inaccuracy: Sampling lists are created periodically, whereas populations change continuously as people move (such as Illinois losing population while Texas gains population).
Note: Sampling lists are set before surveys are conducted, so sampling frame errors are distinct from nonresponse errors.