2.4 - PSY121
Module 2.4: A Statistical Primer
Learning Objectives
Learning Objective 2.4 a: Know the key terminology of statistics.
Learning Objective 2.4 b: Understand how and why psychologists use significance tests.
Learning Objective 2.4 c: Apply your knowledge to interpret the most frequently used types of graphs.
Learning Objective 2.4 d: Analyze the choice of central tendency statistics based on the shape of the distribution.
Probability and Decision Making
Illustration of a multiple choice exam scenario:
Recorded answer: c
Doubt about the answer potentially being a.
Decision-making judgment involves probability estimation.
Key question: How likely is the original answer correct?
Evidence suggests that humans function as probability machines but do so imperfectly.
Intuitive belief among students: 75% will stick with initial choices (common intuition).
Findings from over 30 studies indicate:
Students more likely to correct their mistakes by switching answers.
Reasons behind misleading intuition:
Stronger emotional response from fear of switching to the wrong answer.
Availability heuristic increases the recall of past experiences with wrong switches, skewing estimation of probabilities.
Introduction to Statistics
Importance of statistics:
Toolset for scientists to describe events, test explanations for psychological phenomena, and predict future outcomes.
Acknowledges the limitations of intuitive heuristics.
Simplifying statistics into two general steps:
Organizing Numbers: Create tables or graphs for a bigger picture.
Hypothesis Testing: Assess significance of differences between groups or experimental conditions.
Descriptive Statistics
Purpose of descriptive statistics:.
Organize, summarize, and interpret data to provide a comprehensive view of results.
Main types of descriptive statistics:
Frequency:
Example: Analyzing GRE scores (ranges from 130 to 170).
Key questions:
Frequency distribution: How often do individual scores occur?
Score distribution shape: Normal or skewed?
Central Tendency: Measures of where data clusters include:
Mean: Arithmetic average; calculated by summing all values and dividing by the total number of values.
Median: Midpoint where 50% of observations lie above and below; valuable in skewed distributions.
Mode: Most frequently occurring score; useful for categorical data.
Variability:
Indicates degrees of dispersion within data.
High variability: Scores widely spread; Low variability: Scores closely clustered.
Importance of reporting variability alongside central tendency for comprehensive data interpretation.
Histograms and Data Representation
Histogram: Complete depiction of frequency distribution.
Vertical axis: Frequency of observations per category.
Easier interpretation through heights of bars.
Example: GRE Score distribution visualized in Figure 2.7 and described based on score ranges.
Normal Distribution:
Symmetrical curve where left and right halves are identical around the mean, often referred to as the bell curve.
Skewed Distribution:
Characterized by a large cluster of scores on one end with a long tail on the other side:
Negatively skewed: Many high scores, few low scores (e.g., quiz results where most students excel).
Positively skewed: Many low scores with few high scores (less common).
Central Tendency Statistics Based on Distribution Shape
When analyzing data:
Mean, Median, Mode offer insights, though they can differ based on data characteristics.
If data is normally distributed, the three measures will converge.
In skewed distributions:
Mean is heavily influenced by outliers.
Median remains stable, making it preferable in such cases.
Example of skewed data: Jeff Bezos’s income impacting mean but not median, justifying the use of median for clearer representation.
Variability in Data
Variability provides depth to measures of central tendency.
Standard Deviation: A key measure indicating average distance from the mean:
A larger standard deviation signifies more variability.
Example: Intelligence scores typically follow a normal distribution with a Mean of 100 and a Standard Deviation of 15 (Module 9.1).
68% of scores lie within one standard deviation from the mean (i.e., between 85 and 115).
95% of scores are found within two standard deviations.
Hypothesis Testing: Evaluating Outcomes
After summarizing data, the next step is hypothesis testing:
Hypothesis Test: Method to evaluate if differences among groups are statistically significant—meaningful or attributable to chance.
Key Terms:
Statistically Significant Difference: Implying support for the experimental hypothesis over the null hypothesis.
Noise vs. Signal: In data, variability represents 'noise' that can obscure meaningful differences ('signal').
Example: Text messaging study assessing loneliness:
Two groups: Texting vs. No texting.
Mean scores to compare, though high variability might muddle results, requiring statistical analysis for significance.
Concepts of Statistical Significance
Statistical Significance: Meaningfulness of data results measured through p-values.
Hypothesis formation:
Null Hypothesis: Assumes no difference (results due to chance).
Experimental Hypothesis: Assumes differences arise due to manipulated variables.
p-Value: Probability metric indicating the likelihood results are due to chance;
Lower p-values signify high confidence against null hypothesis.
When comparing p-values:
Historical cutoff recommended by Fisher: p < 0.05 for significance, indicating <5% likelihood results are chance-driven.
In sensitive cases, stricter thresholds (e.g., p < 0.01) can be employed.
Critique of Statistical Significance Testing
Approximation of Type I Error due to multiple comparisons
Larger sample sizes can artificially inflate significance due to small differences.
Shift towards Effect Sizes:
Enhances interpretation beyond mere significant versus non-significant; provides qualitative context.
Current standard practices promote reporting both significance and effect sizes for comprehensive data interpretation.
Summary of Learning Objectives
LO 2.4 a: Key terminology in statistics (frequency, central tendency, variability).
LO 2.4 b: Use of significance tests provides insight into the meaningfulness of differences between groups.
LO 2.4 c: Importance of data representation through graphs (e.g., histograms).
LO 2.4 d: Analyze central tendency choice based on distribution characteristics (skewness, outliers).