Graph Interpretation and Analysis Flashcards

Administrative Requirements and Deadlines

  • RespectUQ Module: Completion of the RespectUQ module is mandatory for all students. The deadline for completion is this Sunday at 5:00PM5:00\,\text{PM}. Failure to complete this module by the specified time will result in being locked out of the Blackboard system by next week.

Evolutionary Strategy for Graph Interpretation

  • Reading graphs is often approached through passive observation, hoping meaning will spontaneously reveal itself. This strategy is ineffective for complex or unconventional visual encodings.

  • A concrete six-step process provides a structured approach to extract maximum information from any graph, no matter how complex or "funky" its design.

Step 1: Title, Caption, and Source

  • Title: The first step is to read the graph's title to understand the primary subject matter. For example: "Median weekly earnings by educational attainment in 2022."

  • Caption/Description: Captions provide crucial context and boundary conditions for the data. In the earnings example, a caption specifies that data applies only to persons aged 2525 and over and refers specifically to full-time wage and salary workers. This excludes part-time workers and individuals under 2525.

  • Source Credibility: Assessing the origin of the data is essential for trust. A source like the US Bureau of Labor Statistics is considered highly credible, whereas data without a source should be verified before being used for policy-making or legal testimony.

Step 2: Axes, Units, and Legends

  • Axis Measurement: Identify what is being measured on each axis and the units used. In the earnings graph, the y-axis represents the level of educational attainment (ranging from "less than a high school diploma" to "doctoral degree"), and the x-axis represents median weekly earnings (ranging from 00 to $2,250\$2,250).

  • Minimum and Maximum Values: Examining the range of the axes helps to contextualize the data points.

  • Legend: Legends explain how different groups or categories are visually represented. If no legend is present, this step is skipped.

Step 3: Visual Encodings

Visual encodings are properties of a graph that change according to the data. There are eight primary encodings:

  1. Length and Height: Used in bar graphs to represent proportional values.

  2. Position: Where a point or bar is placed relative to the axes.

  3. Area: The size of a shape (e.g., circles in a scatter plot).

  4. Angle: Often used in pie charts.

  5. Color Hue: Different colors (e.g., red vs. gray) representing distinct categories.

  6. Color Shade: Intensity or lightness/darkness (e.g., light blue vs. dark blue) representing variations within a category.

  7. Shape: Using different symbols (e.g., circles, squares) for data points.

  8. Width or Thickness: The physical breadth of a line or bar.

  • Encoding vs. Decoration: Color is only a visual encoding if it changes based on the data. If all bars in a graph are the same color, color is a decorative element, not an encoding. If bars are colored differently to represent different data values or categories, it becomes a hue or shade encoding.

  • Example Analysis: In a graph ranking education by earnings, the length of the bars encodes the earnings amount, and the position is used to rank educational levels from highest to lowest achievement.

Step 4: Orienting to the Data

  • Instead of attempting to grasp the entire graph at once, select one or two specific data points to test understanding.

  • Example: In the earnings graph, a data point for "some college but no degree" shows median weekly earnings of approximately $900\$900. A second point for a "bachelor's degree" shows approximately $1,400\$1,400. Successfully identifying individual points confirms the reader understands the basic mechanics of the graph.

Step 5: Annotations

  • Annotations are specific text labels or highlights added by the creator to draw attention to significant or unusual data patterns.

  • If no annotations exist, move directly to step six.

Step 6: Zooming Out and Trend Analysis

  • The final step involves identifying relationships between variables rather than focusing on single points.

  • Variable Identification: Explicitly list the dependent and independent variables.

    • Dependent Variable: The outcome being measured (e.g., weekly earnings).

    • Independent Variable: The factor being manipulated or categorized (e.g., educational attainment).

  • Relationship Interpretation: In the earnings example, there is a positive relationship: as the level of educational attainment increases, median weekly earnings also increase.

Analyzing Complex Relationships: Homicides by Age and Martial Status

  • Data Structure:

    • Y-axis: Homicides by men, normalized as per million men per year.

    • X-axis: Age of the perpetrator in bins (253425-34, 354435-44, 455445-54, 556455-64, 65+65+).

    • Legend/Color Shade: Dark blue represents unmarried men; light blue represents married men.

  • Visual Encodings: Height (number of homicides), Position (age progression), and Color Shade (relationship status).

  • Main Effects:

    • Age: There is a negative relationship between age and homicides; as age increases, the number of homicides decreases. To see this, one must "collapse" across categories by mentally averaging the height of the married and unmarried bars for each age bin.

    • Relationship Status: Unmarried men generally commit more homicides than married men. This is observed by mentally averaging the height of all dark blue bars vs. all light blue bars.

Understanding Interactions

An interaction occurs when the effect of one independent variable on the dependent variable changes depending on the level of a second independent variable. This can be analyzed in two ways:

  1. Does the effect of Age change depending on Relationship Status?: The decline in homicides as men age is much steeper for unmarried men than for married men. For married men, the decline is less stark because they start with a lower baseline of homicides.

  2. Does the effect of Relationship Status change depending on Age?: The difference in homicide rates between married and unmarried men is largest among the youngest age group (253425-34) and becomes significantly smaller (weaker) as the men reach age 6565.

Case Study: The Veil of Darkness Test

  • Title/Source: "The veil of darkness test for police stops in Texas," published in Nature Human Behavior (Pierson et al., 2020).

  • Data Parameters: Based on 112,938112,938 stops of black and white drivers.

  • Technical Details: The vertical line at t=0t = 0 indicates dusk (darkness). A 3030-minute window between sunset and dusk is removed because it is neither fully light nor dark.

  • Visual Encodings:

    • Position (Y-axis): Percentage of stopped drivers who are black (ranging from 15%15\% to 35%35\%).

    • Position (X-axis): Time since dusk (minutes).

    • Area: The size of the circular data points representing the total number of stops in that specific time bin.

  • Patterns and Conclusions:

    • Time and Race: A higher proportion of black drivers are stopped before dusk (daylight) than after dusk (darkness).

    • Hypothesis: This provides evidence for racial motivation in police stops, as officers can see the driver's race more clearly in daylight. The counter-hypothesis that fewer black people drive at night is considered less likely.

    • Police Resources: Larger dots (more stops) occur during daylight, suggesting more police resources are active or more people are driving during those hours.

Case Study: Special Elections and Political Shifts

  • Context: This graph tracks Democratic gains in special elections (by-elections) post-January 20172017 inauguration.

  • Variables:

    • Y-axis: The difference in percentage points between the special election result and the 20162016 presidential election result.

    • X-axis: Days since inauguration (up to 400+400+ days).

    • Color Hue: Gray (Democrat win) and Red (Republican win).

    • Color Shade: Dark (flipped seat) and Light (held seat).

  • Key Findings:

    • Most special elections occur after the first 100100 days.

    • No seats were flipped in the first 100100 days.

    • Almost every flipped seat was won by Democrats. Only one instance occurred where a Republican flipped a seat.

    • Even in seats held by Republicans (light red), most data points are above the zero line, indicating a swing toward Democratic voters compared to the 20162016 presidential results.

    • The Kentucky Outlier: One data point shows a flip with an over 8080 percentage point swing toward Democrats, highlighted via annotation.

Interpreting Paradoxical Data: The "Flipped but More Republican" Points

Some data points show a seat was flipped by Democrats, yet they appear in the "more Republican" bottom half of the graph. This is possible due to the comparison of three different elections:

  1. The 2017 Special Election: This determines the current win (e.g., Democrats win with 55%55\%).

  2. The Previous Regional Election: This determines if it was a "flip" (e.g., if Republicans won the previous one with 60%60\%).

  3. The 2016 Presidential Election: This is the baseline for the y-axis. If the Republican candidate in the 20162016 Presidential race performed significantly better in that specific county than the Republican did in the 20172017 special election, the swing is measured against that specific 20162016 result.

  • While the seat was flipped from the previous local election, the Republican performed relatively better in the special election than the GOP candidate did in the 20162016 presidential race, resulting in a position in the "more Republican" half of the y-axis.

Graph Interpretation Practice: Attractiveness Data

  • Data Collection: 2020 female observers made binary (yes/no) attractiveness judgments for various photographs.

  • Y-axis: Aggregated attractiveness judgments (number of observers out of 2020 who said "yes").

  • X-axis: Stimulus identity (1010 different women), ranked by their average attractiveness scores across all their photos.

  • Findings:

    • Inter-individual Variability: Some people are rated as more attractive on average than others (Identity 1010 vs. Identity 11).

    • Intra-individual Variability: There is a significant range of attractiveness within a single person's photos. For example, some people (Identity 55) have a massive range, with some photos rated very low and others very high.

    • Overlapping Rankings: The best photo of a person ranked lowest on average (Identity 11) can be rated more attractive than the worst photo of a person ranked highest on average (Identity 1010).