Exhaustive Notes on 2x2 ANOVA Interpretation and SPSS Data Reporting

Descriptive Statistics vs. Inferential Statistics in Data Analysis

  • Conceptual Overview:     * Descriptive Statistics: Summarize the data as it exists. These are found on the first page of the SPSS data output and include metrics like the mean (MM) and standard deviation (SDSD).     * Inferential Statistics: Used to make predictions or test hypotheses; they determine if the observed differences between groups are statistically significant.

  • Factorial Design Groups (2×22 \times 2 Matrix):     * The study analyzes two independent variables (Job Status and Type of Attire).     * Main Effect 1 (Job Status): Assistant versus Executive.     * Main Effect 2 (Attire Comparison): Conservative versus Sexy.     * Interaction Groups: The data is broken down into four distinct crossed conditions:         * Assistant-Conservative (ACAC)         * Assistant-Sexy         * Executive-Conservative (ECEC)         * Executive-Sexy

  • Standard Deviation and Error:     * The standard deviations are relatively consistent across all groups, which indicates a low amount of unplanned error in the data.

Sample Size (nn) and Degrees of Freedom (dfdf)

  • Total Sample Size: The total number of human participants reported in the methods section is n=77n = 77.

  • Determining Degrees of Freedom (dfdf):     * In a 2×22 \times 2 ANOVA, there is one degree of freedom per effect.     * Since there are three effects total (22 main effects and 11 interaction), three degrees of freedom are subtracted from the total sample size for the calculation of significance.

  • Adjusted N for Reporting:     * Though the total n=77n = 77, the adjusted number used in reporting (the error term for the FF-statistic) is n=74n = 74 because of the three subtracted degrees of freedom (773=7477 - 3 = 74).     * Reporting Exception: A former colleague named Arvivostaxian occasionally reported the intercept, resulting in four degrees of freedom, though this is considered non-standard practice in the department.

  • Statistical Power: Removing degrees of freedom does not mean removing participants; it is a mathematical adjustment that helps the statistics account for possible error and provides more power during the ANOVA calculation in SPSS.

APA Formatting and Result Reporting Standards

  • Italics and Spacing:     * Statistical symbols such as FF, MM, SDSD, and pp must be italicized.     * There should be no space between the symbol and the parentheses surrounding the degrees of freedom (e.g., F(1,74)F(1, 74)).

  • Decimal Precision and Rounding:     * FF-scores: Report with two digits after the decimal (e.g., 57.8457.84).     * pp-values: Report with three digits after the decimal (e.g., p=0.041p = 0.041) to show as much significance as possible.     * Rounding Rule: If the third or fourth digit is a 55, round the preceding digit up.

  • Anatomy of a Significance Statement:     * Example: "We found a significant main effect of job status, F(1,74)=57.84F(1, 74) = 57.84, p < 0.001."     * The first number in the parentheses (11) indicates the degree of freedom for the main effect, and the second (7474) is the corrected total sample size.

  • Explaining Directions of Effects:     * After reporting the test statistic, descriptive statistics (MM and SDSD) must be provided to show the direction of the finding.     * Example from Data: Perception of competence for Executives (M=50.84M = 50.84, SD=4.60SD = 4.60) was higher compared to Assistants (M=43.91M = 43.91, SD=4.56SD = 4.56).

  • Hypothesis Confirmation: The final sentence of a results block should state whether the hypothesis for that specific effect was supported.

Inferencing and Measurement in Psychology

  • Direct vs. Indirect Measures:     * Indirect (Inferential): Psychology and Sociology often use indirect measures (like Likert scales) because we cannot see the mind directly; we must infer behavior from data.     * Direct (Neuroscience): Uses invasive methods like placing electrodes in specific structures (e.g., Broca's area) to measure neural firing. This requires fewer statistics because it is a direct measurement.

  • Cognitive Experiment Example: Identifying real words (e.g., Tree, House) versus non-words.     * Non-word examples: "GLIC" and "PIPL." These are used because they follow vowel/consonant patterns that look like real words, forcing the brain to process them.     * Dependent Variables (DVDV):         * DV1DV_1: Reaction times measured in milliseconds (msms).         * DV2DV_2: Accuracy measured as percent correct.

Interpreting Interactions via Graphics

  • Interaction Definition: An interaction occurs when the effect of one independent variable depends on the levels of another independent variable.

  • Visual Representation:     * Parallel Lines: Indicate that no interaction is present; the effects are independent.     * Crossing/Converging Lines: Indicate an interaction. Perfect interactions form an "X" shape, but any approach toward crossing is significant.

  • Current Study Results: The interaction in this dataset was "subtle" with a significance of p=0.04p = 0.04. While close to the alpha threshold of 0.050.05, it is still considered statistically significant.

  • Bar Graphs: Can be manipulated for clarity. Starting the Y-axis at a higher number (e.g., 3535 instead of 00) can make differences between bars look more dramatic and highlight the interaction effect.

Survey Construction: Composite Scores and Reverse Scoring

  • Composite Scores: The total score for an individual participant across all items of a measurement tool.     * Process: If a participant takes a 1010-item Likert scale and answers "55" on every item, their composite score is 5050. The average of all participants' composite scores becomes the group Mean (MM).

  • Reverse Wording: Items phrased in the opposite direction of the construct being measured.     * Purpose: To ensure validity and control for Response Set (the tendency for a participant to answer all questions the same way due to fatigue or lack of attention) and lying.

  • Reverse Scoring Math: Adjusting the numerical value of a reverse-worded item so it aligns with the rest of the scale.     * Formula: Take the Likert scale maximum plus one, then subtract the participant's answer (XX).     * For a 44-point scale: 5X5 - X. (e.g., a score of 44 becomes a 11, and a 11 becomes a 44).     * Study Specifics: Item #7 in the survey packet was the reverse-scored item.

Homogeneity of Variance and Levene's Test

  • Definition: Checking if the variance (spread) of scores is similar across all experimental conditions.

  • Heterogeneity vs. Homogeneity:     * Researchers want homogeneity within groups (consistency among participants in a single condition).     * Researchers want heterogeneity between groups (clear differences between the experimental conditions).

  • Levene's Test: A specific statistical test used to check for homogeneity of variance. It ensures there is no "weird crazy error" in one specific group compared to others.

Participant Demographics and Recruitment

  • Total Participants: n=77n = 77.

  • Recruitment: Participants were recruited from psychology classes; consequently, females outnumber males in the sample.

  • Age Profile:     * Mean Age (MM): 2323 years.     * Age Range: 1818 to 5555 years.

  • Discussion Section Use: Detailed breakdowns of student majors/careers can be used in the discussion section to suggest directions for future research or limitations of the study.

Questions & Discussion

  • Q: Does the adjusted n of 74 go in the results or methods?     * A: The total sample of 7777 goes in the methods section under "Participants." Use the adjusted number (7474) when reporting the FF-statistic degrees of freedom in the results section.

  • Q: When rounding significant digits, do we round up for a 5?     * A: Yes, rounding up is the traditional approach. It is generally preferred because higher numbers are often desired in statistical reporting.

  • Q: Why do we use reverse wording?     * A: It is primarily to control for a "response set," where a tired or bored participant might just mark "Agree" for every item without reading the content. It ensures the item is measuring their actual perception.

  • Q: What is the purpose of Levene's Test?     * A: It checks for homogeneity of variance to make sure the experimental groups are comparable and that variance is not skewing the results.