Research Design and Statistics

Selecting and Defining Variables in Research

  • Variable: An attribute, behavior, event, or phenomenon that is capable of varying or having two states, conditions, or levels.

  • Constant: A characteristic that is restricted to one state or condition.

  • Independent Variable (IV):     * Believed to affect another variable ($X$).     * A variable that is intentionally changed or manipulated by the researcher.     * Represented as the treatment or intervention.     * Must have at least two levels ($2$).

  • Dependent Variable (DV):     * The event studied and symbols being measured ($Y$).     * Expected to change when the independent variable is changed (it depends on the IV).     * Measured by pretest and posttest tools.     * It is not manipulated by the researcher.     * Example: In a study on therapy, the "Therapy type" is the independent variable, and the "depressive symptom" is the dependent variable.

  • Operationalization: The process of strictly defining variables into measurable factors. These variables are defined in terms of the specific method by which they will be measured.

  • Measurement Methods:     * Direct measures.     * Tests.     * Observation of behavior.

  • Content Analysis: The process of organizing information into specific categories.

  • Protocol Analysis: A type of content analysis where subjects "think aloud" while solving problems.

Sampling Techniques for Observation and Research

  • Behavior Sampling: Looking at specific aspects of behavior rather than the whole.     * Requires systematic sampling and recording the frequency and duration of the behavior.     * Interval recording/sampling: Observing behavior for a period of time that is divided into intervals. The researcher records whether the behavior occurs within each specific interval.     * Event sampling: Observing and recording the behavior exactly as it occurs. This involves using a pre-coded checklist or recording the specific times when the behavior started and ended.

  • Situational Sampling:     * Goal: To observe behavior in a number of different settings.     * This helps increase the generalizability of the study findings.

  • Sequential Analysis: The process of coding behavioral sequences rather than looking at separate, isolated behaviors.

Experimental Designs and Control

  • True Experiments: An arrangement that permits maximum control over independent variables, providing the strongest basis for drawing inferences.     * Requires Random assignment of subjects.     * Requires Controlling for bias.     * Randomized control clinical trial: A true experiment conducted within the context of an intervention.

  • Quasi-experiment:     * Conditions of true experiments are only approximated.     * No random assignment!     * Variables are studied by selecting subjects who already vary in specific characteristics.     * Goal: To gain insights into the nature of a problem.

  • Random Selection Sampling:     * Definition: Each member of the population has an equal probability of being selected as a subject.     * The selection of one member does not influence the selection of any other member.     * The sample should represent the whole population based on all cases from which the sample is drawn.     * Benefits: Reduces the probability of sample bias and enhances external validity (generalizability).

  • Stratified Random Sampling:     * Used when the population varies in terms of relevant characteristics.     * Goal: Ensure all strata (subpopulations) are represented.     * Members are divided into homogeneous subgroups before sampling.     * Subjects are then randomly selected from each stratum.     * Typical strata: Gender, age, education, SES (Socioeconomic Status), cultural background, etc.

  • Cluster Sampling:     * Selecting clusters of individuals using simple or stratified sampling.     * Individuals are then selected from each cluster.     * Example: Selecting 1010 mental health hospitals first, and then selecting schizophrenia patients from those specific hospitals.

Experimental Control and Variability

  • Fundamental Questions: Is there a relationship? If yes, is the relationship causal?

  • Factors of Variability:     * Independent variable (experimental variance): Researchers want to increase this.     * Systematic error: Researchers want to control this.     * Random error: Researchers want to minimize this.

  • Methods to Increase Control:     1. Increase IV Variability: Make the levels of the independent variable as different as possible.     2. Controlling Extraneous (Confounding) Variables:         * An extraneous variable is a source of systematic error that is irrelevant to the study but has a systematic effect on the DV (it correlates with it).         * Example: Symptom severity affecting the outcome of an intervention. If not controlled, one cannot be sure if the difference is due to the intervention or the symptom severity.         * Solution 1: Random assignment to treatment groups.         * Solution 2: Holding variables constant by selecting subjects who are homogeneous regarding the variable (this reduces generalizability).         * Solution 3: Matching subjects on the variable and then randomly assigning them to groups (useful when sample size is small).         * Solution 4: Blocking (Grouping): Building the extraneous variable into the study as an independent variable.         * Solution 5: Statistical control: If a subject's status on a variable is known, the variability can be statistically removed or equalized by introducing it into the model.     3. Minimizing Random Error: Reducing random fluctuations of subjects, conditions, instruments, and procedures by standardizing procedures and ensuring subjects do not get tired.

Threats to Internal Validity

  • Internal Validity: The extent to which a causal relationship can be established between the IV and DV.

  • Maturation: Processes changing over time (growing older, stronger, wiser, tired, bored).     * Threat to one-group designs; not a threat to two-group designs (assuming both groups mature at the same rate).

  • History: Any event other than the IV occurring inside or outside the experiment that may account for results.     * Threat to one-group pre-post test designs.     * Example: A shooting in a city affecting research on attitudes toward gun violence.

  • Testing: The effects that taking a test once may have on subsequent performance.     * Threat to one-group designs, but not two-group designs if both receive the pre-test.

  • Instrumentation: Changes in the measuring instrument or measurement procedures over time.     * Not a threat in standardized tests or automated scoring.     * Examples: Rewording survey questions, experimenter remarks, or judges practicing scoring.

  • Statistical Regression: The tendency for extreme scores to revert (regress) toward the mean upon re-administration.     * The amount of regression is inversely related to test reliability.

  • Selection: Systematic differences between groups before manipulation due to the assignment of subjects.     * Threat for two-group designs; not a threat for one-group designs.     * Avoided by random sampling and random assignment.

  • Attrition: Loss of subjects during an investigation (leaving, dying). This is a threat for any design with more than one group.

  • Interaction of Selection and Other Threats: When threats apply differentially to different groups.     * Example: One group experiencing a historical event that the other did not.

External Validity

  • Definition: The extent to which results can be generalized beyond the experiment to other populations, settings, and circumstances.

  • Types:     * Population validity: Generalizing to other people.     * Ecological validity: Generalizing to other settings.     * Analogue study: Examining variables in a laboratory setting.

  • Relationship to Internal Validity: Internal validity limits external validity. If a relationship is not causal, it cannot be generalized. High internal validity does not guarantee external validity.

  • Threats to External Validity:     1. Interaction of Testing and Treatment: The pretest sensitizes subjects. Solution: Solomon four-group design or no pretest.     2. Interaction of Selection and Treatment: Subject characteristics (e.g., motivated volunteers) make them respond specifically. Solution: Representative sampling.     3. Reactivity (Reactive Arrangements): Subjects know they are being observed.         * Evaluation apprehension: Acting to avoid negative evaluation.         * Demand characteristics: Cues in the setting informing subjects how to behave.         * Experimenter expectancy: Unintentional cues from the experimenter. Solution: Single and double-blind studies.     4. Multiple Treatment Interference: Carry-over effects from being exposed to more than one condition. Solution: Counterbalanced design.

Group Research Designs

  • Between-Group Design: Each condition is administered to a different group.     * Factorial Design: More than two independent variables.     * Main Effect: The effect of one IV on the DV, disregarding other variables.     * Interaction Effect: When the effect of one IV differs at different levels of another IV. The presence of an interaction invalidates results based solely on main effects.

  • Within-Subject Design: All participants are exposed to every treatment/condition.     * Single pulse time series: Measuring the DV several times at regular intervals.     * Problem: Autocorrelation: Posttest performance correlates with pretest performance, increasing the probability of a Type I error ($false positive$).     * Solution for Carry-over: Counterbalancing.

  • Mixed Design: Combines between and within-subject designs. Contains at least one between-subject IV and one within-subject IV (e.g., 44 types of therapy measured for short and long-term effects).

Single Subject Designs

  • Characteristics: Involves a baseline phase and a treatment phase with repeated measures. Can be used on groups.

  • AB Design: Two phases: AA (Baseline/no intervention) and BB (Intervention).

  • Reversal Design (ABA, ABAB): Treatment is withdrawn after the first phase. If measurements return to baseline in the second AA phase and original treatment levels in the second BB phase, causality is more certain.     * Inappropriate when: Treatment withdrawal is unethical or the effect of the IV persists.

  • Multiple Baseline Design: Treatment is introduced in temporal sequence across:     * Behaviors: Different behaviors of the same subject.     * Settings: Same subject in different settings.     * Tasks: Same subject on different tasks.     * Subjects: Same behavior of different subjects.     * Advantage: Does not require withdrawing treatment.

Statistics: Variables and Scales of Measurement

  • Continuous Variable: Infinite number of values on a scale (e.g., time, age).

  • Discrete Variables: Countable number of values between any two values.

  • Dichotomous Variable: Has only two values (e.g., boy/girl).

  • Nominal Scale: Categorical identifiers or names. Unordered. Only mathematical operation possible is counting frequency.

  • Ordinal Scale: Categories that are rank-ordered. Limitations: Cannot determine the magnitude of difference between ranks.

  • Interval Scale: Represents quantity with equal units. Zero is just another point on the scale (no absolute zero). Addition and subtraction are possible (e.g., IQ, Fahrenheit).

  • Ratio Scale: Equal units, order, and an absolute zero (no numbers below zero). Allows multiplication and division (e.g., height, weight, Kelvin).

Descriptive Statistics and Distributions

  • Frequency Polygons: Graphs joining the midpoints of intervals.

  • Normal Curve: Symmetric, bell-shaped probability distribution.

  • Kurtosis: The sharpness of the peak.     * Platykurtic: Flatter than normal.     * Leptokurtic: More peaked than normal.

  • Skewed Distribution:     * Positively skewed: Most scores on the negative (low) side; few high scores. Mean > Median > Mode.     * Negatively skewed: Most scores on the positive (high) side; few low scores. Mean < Median < Mode.

  • Measures of Central Tendency:     * Mode: Most frequent score. Susceptible to sampling fluctuations.     * Median: Divides distribution in half. Not affected by extreme outliers.     * Mean: Arithmetic average. Least susceptible to sampling fluctuations and used in most statistical procedures, but sensitive to outliers.

  • Measures of Variability:     * Range: Difference between largest and smallest values.     * Variance: Includes all scores; the average of squared deviations from the mean.     * Sum of Squares: (XXˉ)2\sum (X - \bar{X})^2.     * Standard Deviation: Square root of the variance. Useful for comparing distributions.

Normal Distribution and Sampling Theory

  • Standard Deviation Areas under Normal Curve:     * ±1\pm 1 SD: 68.26%68.26\%     * ±2\pm 2 SD: 95.44%95.44\%     * ±3\pm 3 SD: 99.71%99.71\%     * 84%84\% of cases are below +1+1 SD.     * 16%16\% of cases are above +1+1 SD.

  • Central Limit Theorem: As sample size increases, the sampling distribution of the mean approaches a bell shape.

  • Standard Error (SESE):     * SE=σnSE = \frac{\sigma}{\sqrt{n}}     * Standard error increases when standard deviation is large or sample size (nn) is small.

Logic of Hypothesis Testing

  • Null Hypothesis (H0H_0): States there is no effect.

  • Alternative Hypothesis (H1H_1): States there is an effect.     * Nondirectional (two-tailed): Only states H0H_0 is false.     * Directional (one-tailed): Specifies the direction of change.

  • Rejection Region: Range of unlikely values representing the level of significance (Alpha, α\alpha).     * If the result falls here, H0H_0 is rejected (statistically significant).     * Typically set at α=0.05\alpha = 0.05 (95%95\% confidence it's not luck) or α=0.01\alpha = 0.01 (99%99\% confidence).

  • Decision Errors:     * Type I Error: Rejecting a true null hypothesis ($false positive$). Probability equals alpha (α\alpha).     * Type II Error: Retaining a false null hypothesis ($miss$). Probability is Beta (β\beta).     * Type I and Type II errors have an inverse relationship.

  • Statistical Power: The ability to reject a false null hypothesis.     * Increased by: Increasing α\alpha, increasing sample size, increasing IV intensity, minimizing error, using one-tailed tests, or using parametric tests.

Inferential Statistical Tests

  • Parametric Tests: Evaluate differences in population means/parameters. Used for interval/ratio data, normal distributions, and Homoscedasticity (equal variances).     * Student’s t-test: Compares means for Single samples, Independent samples (22 groups), or Correlated samples (within-subject/matched).     * Analysis of Variance (ANOVA): Compares 22 or more means. One-way (one IV) or Factorial (two or more IVs; e.g., two-way ANOVA).

  • Nonparametric Tests: Used for nominal/ordinal data or non-normal distributions (distribution-free).     * Mann-Whitney U test: Compares medians of two independent groups (ordinal data).     * Kruskal-Wallis test: Compares two or more independent groups (ordinal data).     * Wilcoxon matched-paired signed rank test: For matched/correlated ordinal data.

  • Effect Size: Magnitude of difference in SD units.     * Cohen’s d.     * Eta squared ($\eta^2$): Percent of variance accounted for by the treatment (e.g., 0.55=55%0.55 = 55\%).

Correlational and Multivariate Techniques

  • Correlation Coefficient (rr): Measures direction and strength (range 1.00-1.00 to +1.00+1.00).

  • Pearson r Assumptions: Linearity, unrestricted range, and homoscedasticity.

  • Coefficient of Determination (r2r^2): Proportions of shared variability. For r=0.6r = 0.6, r2=0.36r^2 = 0.36 (36%36\% shared variance).

  • Regression Analysis:     * Linear Regression: Minimizes error using the "least square criterion" to predict YY from XX.     * Multiple Regression: Two or more predictors, one criterion. Result is RR and R2R^2. Often used instead of ANOVA for unequal group sizes or continuous IVs.

  • Canonical Correlation: Extension of multiple regression with two or more predictors and two or more criteria.

  • Factor Analysis: Reduces many data points to a few factors to explain intercorrelation; used for subscales.

  • Cluster Analysis: Groups data based on similarities (within-group homogeneity and between-group heterogeneity). Used for identifying subgroups (e.g., ADHD subgroups).