3020 chapter 2
Data & Distributions
CHAPTER 2 Overview
Source: McBride, The Process of Statistical Analysis in Psychology, 1st Edition, SAGE Publishing, 2017.
Learning Objectives
Difference Between Population and Sample: Understanding these fundamental concepts in statistics, especially in psychology.
Types of Data in Psychological Studies: Familiarizing oneself with the various forms of data collected in psychology.
Understanding Distribution: Learning how the shape of a distribution affects data analysis.
Using SPSS software: Exploring how computer programs assist in examining data distributions.
1. Learn the Difference Between a Population and a Sample
Defining Key Concepts
Population: A group of individuals that a researcher seeks to learn about from a research study.
Example: All University of Missouri (Mizzou) undergraduate students.
Sample: A subset of individuals chosen from the population to represent it in a research study.
Example: A specific set of Mizzou undergraduates who answered a poll question.
Importance of Population and Sample Representation
Factors Influencing Representation:
Sample size
Method of sample selection
Response rate: Of the chosen individuals, how many actually participated.
2. Understanding Sampling Error
Overview of Sampling Error
Definition: Sampling error refers to the discrepancies that arise when data is collected from a sample instead of the entire population.
Each sample represents a different subset of the population, leading to variability in scores across samples.
Estimating Population Parameters
The ultimate goal is to estimate the population mean ($ar{x}$) using the sample mean.
Margin of Error: An estimate that reflects how far the reported percentage from a sample is likely to be from the population percentage.
Estimating the Usefulness of Polls:
Sample size
Margin of error
3. Understand the Kinds of Data Collected in Psychological Studies
Operational Definitions
Definition: An operational definition specifies how a behavior is measured in a study.
It allows researchers to quantify behaviors that are not directly observable.
Multiple methods can define the same behavior, which can impact internal validity.
Examples of Operational Definitions
Measuring depression:
Example 1: Observing the frequency of smiles in one hour (fewer smiles indicate higher depression).
Example 2: Observer ratings of lethargy.
Example 3: Scores on self-reported questionnaires regarding mood.
Scales of Measurement
Nominal Scale: Categorical, non-ordered groups (e.g., classifications of mood).
Ordinal Scale: Ordered categories (e.g., rankings in competitions).
Interval Scale: Numerical responses with equal intervals but no true zero (e.g., temperature on Celsius or Fahrenheit scales).
Ratio Scale: Numerical responses with a true zero point (e.g., height, weight, age).
Survey Data Limitations
Surveys use various measurement scales but are limited to individuals' self-reports.
Social Desirability Bias: A phenomenon where respondents skew their answers to appear more favorable.
Definition: Bias stemming from a respondent's desire to be viewed favorably by others.
Example: Responses concerning sensitive topics like alcohol consumption may be understated.
Types of Social Desirability Bias
Self-Deceptive Enhancement: Respondents falsely believe in their positive behaviors despite evidence to the contrary (e.g., overstating recycling habits).
Impression Management: Conscious alteration of responses to comply with social norms (e.g., underreporting or exaggerating criminal behavior).
4. Internal and External Validity
Internal Validity
Measures the integrity of research and whether it accurately assesses the independent variable.
Minimized confounding variables lead to higher internal validity.
Examples: In a study on mice and exercise, controlling for variables like diet enhances internal validity.
External Validity
Addresses whether findings can be generalized to real-world scenarios.
Stronger control in data collection may decrease external validity.
5. Understanding Distribution
Definitions and Concepts
Distribution: A set of scores collected in the research. Analyzing distributions helps researchers summarize data.
Frequency Distribution
A method to organize data using tables or graphs that show how frequently each score occurs.
Creating Frequency Distribution Graphs:
Place scores on the X-axis.
Plot frequency counts on the Y-axis.
Creating Frequency Distribution Tables
Table Creation Steps:
List all possible scores in one column.
Indicate frequency counts in another column.
List percentages of occurrences.
Include cumulative percentages in a final column.
Shape of Distribution
The visual representation of a distribution allows researchers to assess clustering of scores.
Symmetrical Distribution: A balanced shape with mirrored sides.
Skewed Distribution: Clustering towards low or high score ends.
Positive Skew: Higher frequencies at lower scores.
Negative Skew: Higher frequencies at higher scores.
6. Introduction to SPSS
Overview of SPSS
SPSS: Widely used software for performing statistical analysis, providing a user-friendly interface compared to Excel.
Important for both descriptive and inferential statistics in psychology.
Using SPSS
Data Input Window: Similar to Excel; users define variable names and details in a separate tab.
Creating Tables:
Enter data and define variables.
Analyzing Frequencies: Access via the Descriptive Statistics menu; ensure frequency tables box is checked.
Creating Graphs: Use the Charts option to generate histograms and visual outputs.
Conclusion
Frequency distribution tables and graphs are critical in summarizing data and understanding the shape of distributions. Utilizing software like SPSS can enhance the ease and accuracy of data analysis in research.