Research Methods in Psychology: Scientific Reasoning, Research Designs, and Statistical Safeguards
Introduction to Research Design and Historical Failures
Central Themes of Research Methodology:
Human beings are inherently prone to cognitive errors and biases.
Scientific research methods serve as essential safeguards to protect against these cognitive errors.
Without scientific safeguards, clinical intuition and unexamined assumptions can lead to disastrous real-world consequences that harm human lives.
Case Study: Facilitated Communication:
Background: Developed in Australia and popularized in the , facilitated communication was promoted as a revolutionary breakthrough for treating infantile autism.
Theoretical Premise (Douglas Bicklin, ): Proposed that infantile autism is fundamentally a severe motor movement disorder rather than an intellectual or mental disorder. Proponents asserted that children with autism possess normal intelligence but lack the motor control to speak or type independently.
Procedure: A facilitator sits next to a child with autism, gently holding and steadying the child's hand as it moves across a computer keyboard or letter pad.
Claimed Results: Produced complete, syntactically complex sentences expressing deep emotions (e.g., "Mommy, I want you to know that I love you even though I can't speak") and complex cognitive tasks (e.g., a child requesting a medication change after reading a medical journal article; Mann, ).
Empirical Controlled Testing: Researchers placed facilitators and children in adjoining cubicles separated by a wall with an opening for hand contact over a keyboard. Researchers flashed photos on adjacent screens; on some trials, the facilitator and child saw different pictures (e.g., facilitator saw a dog, child saw a cat).
Findings (Jacobson et al., ; Romanczyk et al., ; Todd, ): In virtually of the trials, the typed word corresponded strictly to the picture flashed to the facilitator, not the picture shown to the child.
Underlying Mechanism: Facilitated communication relies on the ideomotor effect (Wegner, ), wherein unconscious thoughts direct physical movements without conscious awareness. It functions analogously to a Ouija board, where operators unknowingly guide the pointer while attributing the movement to outside forces.
Case Study: Prefrontal Lobotomy:
Background: Introduced in the early century as a mainstream treatment for schizophrenia and other severe mental disorders.
Procedure: Surgical procedure where a surgeon severs the neural fibers connecting the brain's frontal lobes to the underlying thalamus.
Scientific Recognition: Developed by Portuguese neurosurgeon Egas Moniz, who was awarded the Nobel Prize in Medicine in .
Methodological Flaw: Believers relied almost exclusively on uncontrolled, subjective clinical observations (e.g., physicians stating, "I am a sensitive observer, and my conclusion is that a vast majority of my patients get better as opposed to worse after my care"; Dawes, ).
Empirical Outcome: Controlled scientific studies demonstrated that prefrontal lobotomies were useless for treating core schizophrenia symptoms (e.g., auditory hallucinations, persecutory delusions) and created severe behavioral deficits, including extreme apathy (Deschamps et al., ; Valenstein, ).
Case Study: Surgical Organ Removal (Dr. Henry Cotton):
Historical Context: Dr. Henry Cotton, superintendent of Trenton State Hospital in New Jersey during the early century, hypothesized that severe psychological disorders were caused by localized bacterial infections.
Procedure: Surgically removed hundreds of patients' teeth, tonsils, large intestines, spleens, gallbladders, and other internal organs.
Evaluation: Relied entirely on subjective clinical judgments of patient improvement, leading to widespread mutilation and death without therapeutic benefit (Scull, ).
Modes of Thinking and Cognitive Shortcuts
Two Modes of Thinking (Kahneman, ; Stanovich & West, ):
System 1 / Intuitive Thinking (Hammond, ; Gladwell, ):
Rapid, automatic, reflexive, and emotionally driven.
Requires minimal mental effort and operates predominantly on gut hunches and first impressions.
Essential for survival and real-time navigation (e.g., swerving to avoid an oncoming vehicle or pothole).
Relies heavily on heuristics (mental shortcuts or rules of thumb).
System 2 / Analytical Thinking (Hammond, ):
Slow, deliberate, reflective, and mentally demanding.
Requires conscious mental effort and logical processing (e.g., solving math problems or understanding academic concepts).
Serves as an override mechanism to evaluate, correct, or reject flawed intuitive gut hunches (Abrami et al., ; Gilbert, ).
Skill Acquisition Continuum:
Complex cognitive and motor tasks begin in analytical thinking mode and transition to intuitive thinking mode through practice.
Example: Learning to drive a car initially requires intense analytical focus (monitoring turn signals, checking blind spots, calculating a buffer from other vehicles). Over time, driving becomes automatic and requires negligible conscious effort.
Function of Research Designs:
Formal research designs function as systematic tools engineered to harness analytical thinking.
They override the biases inherent in intuitive thinking and force researchers to systematically evaluate alternative explanations.
The Scientific Method and Foundations of Measurement
Nature of the Scientific Method:
The scientific method is not a single, rigid recipe; it represents a flexible toolbox of techniques designed to counteract cognitive biases and guard against errors in human thinking.
Modern scientific approaches incorporate emerging tools, including Artificial Intelligence (AI), to complement traditional research methods.
Sampling and Generalizability:
Random Selection: A sampling procedure in which every individual in the targeted population has an equal mathematical chance of being chosen to participate.
Generalizability: The extent to which findings from a sample apply to the broader population.
Sample Quality vs. Sample Size: A smaller, randomly selected sample (e.g., randomly selected North Americans) provides a far more accurate representation than a massive, non-random sample (e.g., residents of Nashville, Tennessee, who are unrepresentatively skewed toward country music preferences).
Polling Failures:
US Presidential Election: Non-random selection led major pollsters to incorrectly predict a victory for Hillary Clinton over Donald Trump (Jennings & Wlezien, ).
US Presidential Election: Polls indicated a tight race between Donald Trump and Kamala Harris; final results showed Trump outperformed polling by approximately in swing states (Morris, ).
Evaluating Psychological Measures:
Reliability: Consistency of measurement.
Test-Retest Reliability: Extent to which a questionnaire or test yields similar scores across multiple administrations over time.
Inter-Rater Reliability: Extent to which different independent observers or raters agree on behavioral or diagnostic observations (e.g., two thermometers yielding identical temperatures, or two clinicians agreeing on diagnostic assessments).
Validity: The extent to which a measure assesses what it claims to measure ("truth in advertising").
Relationship Between Reliability and Validity:
Reliability is a necessary condition for validity (a measure must measure consistently before it can measure accurately).
Reliability does not guarantee validity. A test can be perfectly reliable yet completely invalid.
Example: The Distance Index Middle-Width Intelligence Test (DIMWIT), calculated by subtracting index finger width from middle finger width, possesses high test-retest and inter-rater reliability, but has zero validity as a measure of intelligence.
Example: The polygraph (lie detector) demonstrates high test-retest reliability, but low validity as a lie detector because it measures physiological arousal rather than deception.
The Open Science Movement and the Replicability Crisis:
Replicability: The ability to duplicate original research findings in independent new studies using new participants and new data.
Reproducibility: The ability to duplicate statistical findings by re-analyzing the original dataset using the same analytical procedures.
Replicability Crisis: Failures to replicate high-profile findings across psychology, medicine, physics, and geology.
Key Open Science Practices:
Publicly archiving raw datasets, code, and materials in open repositories.
Directly conducting internal and independent replications before publishing.
Preregistration: Publicly registering hypotheses, methodology, and statistical analysis plans prior to data collection to prevent post-hoc modifications.
Combating the file drawer problem (publication bias; Rosenthal, ), where non-significant or null results are left unpublished.
Emphasizing meta-analyses and systematic reviews over single isolated studies.
Engaging in adversarial collaboration, where opposing theoretical teams co-design research protocols.
Observational, Case Study, and Self-Report Designs
Naturalistic Observation:
Definition: Observing and recording behavior in real-world settings without manipulating variables.
Examples:
Jane Goodall utilized naturalistic observation in Gombe, Kenya, discovering that warfare and structured aggression occur in wild chimpanzee populations.
Robert Provine (, ) eavesdropped on natural laughter incidents: found women laugh significantly more than men in social settings; less than of laughter followed humorous statements; speakers laugh significantly more than listeners.
Bowker et al. () observed youth hockey games: fans averaged total comments per game, with the vast majority being positive; fans averaged only negative comments per game (mostly directed at referees).
Modern Data Collection: Uses social media scraping, wearable devices (Apple Watch, GoPro), and mobile brain scanning.
Trade-off:
High in External Validity: Results generalize well to real-world settings.
Low in Internal Validity: Lacks experimental control over extraneous variables; prevents direct cause-and-effect inferences.
Case Study Designs:
Definition: Intensive examination of one individual or a small group over an extended period.
Applications:
Providing existence proofs (demonstrating that a psychological phenomenon can exist, such as recovered memories of trauma).
Studying rare or extreme neurological/psychological conditions (e.g., prosopagnosia [facial blindness], super-recognizers, Capgras syndrome [belief that loved ones have been replaced by identical impostors]).
Generating novel hypotheses (e.g., Aaron Beck developed Cognitive Behavioral Therapy after probing an anxious client's irrational fear of boring others).
Limitations: Cannot establish causality; vulnerable to anecdotal fallacies; findings cannot be generalized to broader populations.
Self-Report Measures and Surveys:
Wording Effects: Wording drastically alters responses.
In a survey of female homemakers, "Would you like to have a job, if this were possible?" yielded positive responses, whereas "Would you prefer to have a job, or do you prefer to do just your housework?" yielded only positive responses (Noelle-Neumann, ).
In a survey assessing public opinion on the non-existent "Agricultural Trade Act of ", of respondents offered an opinion (Bishop et al., ).
Response Sets (Systematic Response Distortions):
Socially Desirable Responding: Distorting responses to present oneself in a favorable light (e.g., overstating high school GPA; female undergraduates reporting fewer lifetime sexual partners unless connected to a fake lie detector machine).
Malingering: Faking or exaggerating psychological symptoms to achieve a specific goal (e.g., mafia boss Vincent "The Oddfather" Gigante, who faked schizophrenia for years with psychiatric hospitalizations to avoid criminal prosecution; Resnick & Noll, ).
Rating Data and Evaluative Effects:
Rating Data: Asking informed observers (peers, supervisors) to rate a target individual; circumvents personal blind spots.
Halo Effect: Tendency for an observer's rating of one positive characteristic (e.g., physical attractiveness) to bias ratings of other positive traits (e.g., intelligence, conscientiousness, competence).
Nisbet & Wilson (): Students viewed a professor displaying either a friendly or unfriendly demeanor. Students watching the friendly professor rated his accent, physical appearance, and mannerisms significantly higher.
Horns / Pitchfork Effect: Tendency for an observer's perception of one negative characteristic to spill over and negatively bias ratings of other unrelated traits (e.g., UBC students rating instructors with poor teaching skills as holding racist and sexist views; Corren, ).
Correlational Designs and Associations
Core Concepts:
Correlational Design: Research design examining the statistical association between two or more variables without direct experimental manipulation.
Correlation Coefficient (): Statistical metric ranging from to
Positive Correlation (): As variable increases, variable increases.
Zero Correlation (): Variable has no statistical relationship with variable
Negative Correlation (): As variable increases, variable decreases.
Strength of Correlation: Determined strictly by the absolute value of , independent of sign ().
Exceptions: Psychology is a science of trends and exceptions; correlation coefficients in psychology are almost strictly less than . Pointing to a single counter-example ("I know a person who smoked packs a day and lived to ") does not invalidate a statistical correlation.
Illusory Correlation:
Definition: The perception of a statistical association between two variables where no actual relationship exists.
Examples:
Lunar Lunacy Effect: The belief that full moons cause increased crime, psychiatric admissions, and births. Data show the correlation is precisely , yet of hospital medical staff endorse the belief (Shur et al., ).
Arthritis and Weather: The belief that joint pain increases during rainy weather; objective data show no correlation.
Superstitions: Baseball player Wade Boggs eating chicken before every game for years.
Cognitive Origin (The Great Fourfold Table of Life):
Matrix of 4 outcomes: Cell A (Full Moon & Crime), Cell B (Full Moon & No Crime), Cell C (No Full Moon & Crime), Cell D (No Full Moon & No Crime).
Driven by confirmation bias and availability heuristics, people focus heavily on Cell A (confirming occurrences) and ignore non-events in Cells B, C, and D.
When prophetic dreamers were instructed to keep systematic dream diaries (forcing tracking of Cell B non-events), their belief in prophetic dreams vanished (Alcock, ).
Correlation vs. Causation Fallacy:
Core Principle: Correlation does not equal causation.
Third Variable Problem ( causes both and ):
Positive correlation between birth rates in Berlin () and stork populations () is driven by overall Population Size ().
Positive correlation between daily ice cream consumption () and violent crime rates () is driven by Outdoor Temperature ().
Shark attacks correlate with tornadoes (; –).
Bruce Willis movie appearances correlate with boiler explosion deaths (; –).
Experimental Designs and Cause-and-Effect Inferences
Essential Components of an Experiment:
Random Assignment: Experimenter randomly sorts participants into groups (Experimental Group vs. Control Group), canceling out preexisting individual baseline differences.
Between-Subjects Design: Different participants are assigned to experimental vs. control conditions.
Within-Subjects Design: The same participant is evaluated before and after an experimental manipulation, acting as their own control.
Manipulation of an Independent Variable: The experimenter directly manipulates the Independent Variable (IV) and measures its precise effect on the Dependent Variable (DV).
Operational Definition: A precise, explicit, working definition of how a variable is measured or manipulated in a specific study.
Confounding Variable (Confound): Any extraneous variable that differs systematically between the experimental and control groups other than the intended independent variable, destroying internal validity.
Major Pitfalls in Experimental Design:
Placebo Effect:
Definition: Observable physical or psychological improvement resulting purely from the expectation of improvement.
History: Physician Franz Anton Mesmer claimed cures via animal magnetism. A French commission led by Benjamin Franklin tested Mesmer's claims by comparing magnetized vs. non-magnetized trees; patients fainted only when they believed a tree was magnetized.
Neurobiology: In Parkinson's disease placebo trials, expectation of recovery triggers measurable dopamine bursts in the brain. Placebos engage cannabinoid, dopamine, and opioid neurotransmitter pathways (Sheldon & Opiumoran, ).
Characteristics: Injectable placebos work faster than oral placebos; expensive placebos work better than cheap placebos; antidepressant medication outcomes display an outcome overlap with placebos (Hangartner & Platter, ).
Control: Requires single-blind designs where participants are kept blind to group assignment.
Nocebo Effect:
Definition: Physical or psychological harm resulting purely from the expectation of harm.
Examples: Allergic individuals sneezing when exposed to artificial roses; imaginary head electrical currents producing real headaches in over two-thirds of students; fake antidepressant overdoses producing severe hypotension.
Experimenter Expectancy Effect (Rosenthal Effect):
Definition: Occurs when researchers' hypotheses unintentionally bias study outcomes in subtle, non-conscious ways.
Clever Hans: Arabian stallion owned by Wilhelm von Austen that supposedly performed complex math. Psychologist Oscar Pfungst () revealed Hans was reading micro-level physical muscle tension cues emitted unintentionally by questioners when Hans reached the correct number of taps.
Rosenthal & Fode (): Students were given randomly selected rats but were falsely told they were either "maze-bright" or "maze-dull". Students with supposed "maze-bright" rats reported faster maze times.
Control: Requires Double-Blind Designs, where neither researchers nor participants know group assignments.
Demand Characteristics:
Definition: Cues in a study that allow participants to generate guesses regarding the researcher's hypotheses, leading them to alter their behavior.
Control: Uses cover stories, distractor tasks, and filler items.
Laboratory External Validity:
Mook (): High internal validity provides the prerequisite foundation for external validity.
Anderson et al. (): Meta-analysis comparing lab vs. real-world effect sizes found a strong correlation ().
Mitchell (): Generalizability varies by subfield: Industrial/Organizational Psychology (), Personality Psychology (), and Social Psychology ().
Ethical Issues in Research Design
Ethical Obligation and Value Neutrality:
Science is value-neutral (truth seeking), but research procedures are bound by human ethics.
Historical Ethical Failure: Tuskegee Syphilis Study (–):
Conducted by the United States Public Health Service on impoverished rural Black men in Alabama.
Participants were lied to about their diagnosis (told they had "bad blood") and denied informed consent.
When penicillin was established as an effective syphilis treatment in the , researchers actively withheld treatment to study the disease's natural progression.
Devastating Toll: men died directly from syphilis; died from related complications; wives contracted syphilis; children were born with congenital syphilis.
Formal apology issued by President Bill Clinton in ; last survivor Anis Hendon died in
Human Ethical Frameworks:
Research Ethics Boards (REBs): Institutional review committees in Canada operating under the Tri-Council Policy Statement (TCPS: CIHR, NSERC, SSHRC; updated , ).
Informed Consent: Subjects must be fully informed of study procedures, risks, and voluntary withdrawal rights prior to participation.
Deception and Debriefing:
Deception is permissible only when alternative non-deceptive methods are impossible, rights are unaffected, and no medical/therapeutic intervention is involved.
Stanley Milgram () obedience studies used confederates and deception regarding shock administration. Post-study surveys revealed reported negative emotional after-effects. Replicated by Berger () up to ( of original participants who passed continued to maximum ).
Debriefing: Mandatory post-experimental session explaining the true nature of the study, hypothesis, and any deceptive elements.
Indigenous Ethical Perspectives:
Christine Walsh () compared TCPS individual autonomy principles with Inuit Qaujimajatuqangit (IQ) values (Pidget Cernic, Pillamax Arnic, Abitinik, Kamatyarnik).
Highlights community-based, collective decision-making and multigenerational obligations over Western individualistic consent frameworks. TCPS now features dedicated chapters on Indigenous research collaboration.
Animal Research Ethics:
Canadian Council on Animal Care (CCAC): Oversees approximately research animals annually in Canada ( mice, fish).
Evaluated by Animal Care and Use Committees (ACUCs) including certified veterinarians and community members to enforce humane housing, care, and minimization of distress.
Statistics: Descriptive and Inferential
Descriptive Statistics:
Central Tendency:
Mean: Arithmetic average (total score divided by sample size).
Median: Exact middle score in a sequentially ordered dataset.
Mode: The most frequently occurring score in a dataset.
Impact of Skewness and Outliers:
In normal distributions, Mean, Median, and Mode align.
In skewed distributions, extreme scores (outliers) pull the Mean significantly away from the center. The Median or Mode must be reported instead.
Example: Dataset of IQs (, , , , ) vs. a dataset with an outlier (, , , , ). The outlier increases the Mean to , whereas Median and Mode remain
Variability (Dispersion):
Range: Difference between the highest and lowest scores. Intuitive but highly sensitive to extreme outliers.
Standard Deviation: Average metric of how much individual data points differ from the dataset mean. Accounts for every data point.
Inferential Statistics:
Determines whether sample findings can be generalized to the broader population.
Statistical Significance: A finding is statistically significant if the probability that it occurred by chance alone is less than in ().
Meta-Analysis: A statistical procedure combining results from multiple independent studies to determine overall pattern strength and effect size.
Practical Significance: Real-world importance. A result can be statistically significant due to a massive sample size ( yielding ) while possessing zero practical utility.
Statistical Manipulations and Misuses:
Non-Representative Central Tendency: MP Ms. Hoffs proposes a tax plan where get a cut and gets a cut. She advertises an average cut of (the Mean), masking the true Median/Mode ().
Neglecting Base Rates: Base rate is the baseline frequency of a behavior or characteristic in a population (e.g., alcohol misuse base rate is ). Prof. Glasgow compares alcohol misuse in Slosh, SK ( misusers: German descent vs. Norwegian descent). Fails to account for German base rate being times higher in that town, meaning Norwegians actually misuse alcohol at a higher proportional rate.
Peer Review, Media Evaluation, and Extrasensory Perception
Peer Review Process:
Outside expert reviewers evaluate submitted manuscripts for methodological flaws.
Flaw Identification Practice:
Dr. Scabia subliminal tape study: Lacks a control group and lacks an independent variable manipulation (not a true experiment).
Dr. Townsend anger expression therapy study: Lacks an attention-placebo control group and fails to use double-blind controls for experimenter expectancy effects.
Evaluating Popular Media Claims:
Source Evaluation: Prioritize primary sources (original peer-reviewed journal articles) over secondary news outlets.
Sharpening and Leveling: Sharpening exaggerates the central message/gist of a study; Leveling minimizes crucial context and qualifications.
Pseudosymmetry: False balance created when media outlets grant equal airtime to pseudoscientific viewpoints vs. overwhelming scientific consensus.
Extrasensory Perception (ESP):
Three Subtypes: Precognition (predicting future events paranormaly), Telepathy (reading minds), Clairvoyance (detecting hidden objects).
Rhine Zener Card Experiments (): Joseph B. Rhine claimed success ( correct out of vs. chance). Flaws: cards were physically worn (symbols visible through backs) and cards were improperly randomized. Failed to replicate.
Ganzfeld Technique: Ping-pong ball eye goggles and red light. Results display chance-level outcomes.
Daryl Bem Precognition Studies (): Claimed post-test word rehearsal retroactively boosted pre-test memory recall across of studies. Replications by Ritchie, Wiseman, and French yielded absolute zero effect.
Ad Hoc Hypotheses: Unfalsifiable explanations for failure (e.g., claiming experimenter skepticism inhibits ESP ["experimenter effect"], or below-chance performance proves ESP ["sign missing"]).
Belief Persistence: of Canadians believe in psychic powers (Angus Reid, ). Driven by illusory correlation and cognitive underestimation of coincidences.
Birthday Paradox: In a room of only people, the probability that two people share the exact same birthday exceeds (). In a room of people, probability exceeds ().
Cold Reading: Communication techniques used to convince strangers that a reader knows intimate details about them (using population stereotypes like predicting numbers or ).
The Paranormal Challenge: Center for Inquiry Investigations Group (CFIIG) offers a prize (continuing James Randi's challenge) for any verified demonstration of paranormal ability under controlled conditions. Zero applicants have passed preliminary testing.