Ch. 5 - Identifying Good Measurement - 1

Chapter 5: Identifying Good Measurement

  • Three common types of measures

    • Self-report measures

      • Operationalize a variable by recording people's answers to questions about themselves in a questionnaire

      • Can also include "3rd person" accounts - parent/teacher/caregiver

      • Example: Caffeine Consumption "How many caffeinated beverages have you consumed today?"

    • Observational measures

      • Operationalize a variable by recording observable behavior

      • Can be tracked automatically or by an observer

      • Example: Caffeine Consumption Sitting in a café and counting how many beverages someone drinks

    • Physiological measures

      • Operationalize a variable by recording biological data (e.g., blood solutes, brain activity, galvanic skin response, heart rate, etc.)

      • Example: Caffeine Consumption Taking a blood sample to measure caffeine or its metabolites in blood

Scales of Measurement

  • Categorical/Nominal Scales

    • Variables classified as a category (aka, nominal)

    • Labels do not relate to each other quantitatively

    • Examples: animal species, eye color, brands

  • Quantitative Scales

    • The actual values matter and have some relation to each other

    • Examples: temperature, brain activity, steps per day, weight

    • Ordinal, interval, ratio

Ordinal Scales

  • Rank-order behaviors and events - 1st, 2nd, 3rd, etc.

  • The difference between each scale unit is not necessarily uniform - only the order matters

Interval Scales

  • Distance between points is equal and meaningful - no real zero

  • Allows researcher to specify distance between observations on a given dimension

  • Examples: IQ score, shoe size, temperature

Ratio Scales

  • Same as interval but with a real zero

  • Examples: Finishing an exam in 50 min takes twice as long as finishing in 25 minutes

  • Number of correct or incorrect answers can be zero

  • Many psychological measures have no meaningful "zero" - zero memory, zero intelligence...

Reliability

  • Three types of reliability

    • Test-retest

    • Internal

    • Interrater/inter-observer

  • Evaluating the reliability of observations

    • Scatterplot

    • Correlation coefficient (r)

    • Inter-observer agreement calculations

Test-retest reliability

  • Consistent scores every time the measure is used

  • Relevant when studying constructs that are expected to be stable over time (intelligence, personality, etc.)

  • Example: IQ test at beginning of semester (Time 1) and at the end of the semester (Time 2) - scores should be relatively consistent

Internal reliability (aka internal consistency)

  • A participant provides a consistent pattern of responses, regardless of how the researcher phrased the question or runs the session

  • Example: estimate of caffeine consumption (volume of caffeinated beverages) remains consistent across different questions

  • Not to be confused with internal validity!

Page 23: Internal Reliability (aka Internal Consistency)

  • Measure of how consistent a measure is within itself

  • Example: Estimating the number of caffeinated beverages consumed daily

    • Reliable responses: 4, 5, 4, 6, 5, 6, 4, 5, 5, 4, 4, 6

    • Unreliable responses: 8, 1, 5, 10, 2, 7, 5, 1, 1, 9, 4, 3

Page 24: Interrater (Inter-Observer) Reliability

  • Consistent scores regardless of who is measuring

  • Direct observation methods used (e.g., one-way mirrors, video recording)

  • Example: Observers counting the number of people who wipe their nose before touching a door handle

Page 25: How to Evaluate Reliability

  • Methods for evaluating reliability:

    • Scatterplot

    • Correlation coefficient (r)

    • Inter-observer agreement calculations

Page 26: Using a Scatterplot to Evaluate Reliability

  • Scatterplot used