Ch. 5 - Identifying Good Measurement - 1
Chapter 5: Identifying Good Measurement
Three common types of measures
Self-report measures
Operationalize a variable by recording people's answers to questions about themselves in a questionnaire
Can also include "3rd person" accounts - parent/teacher/caregiver
Example: Caffeine Consumption "How many caffeinated beverages have you consumed today?"
Observational measures
Operationalize a variable by recording observable behavior
Can be tracked automatically or by an observer
Example: Caffeine Consumption Sitting in a café and counting how many beverages someone drinks
Physiological measures
Operationalize a variable by recording biological data (e.g., blood solutes, brain activity, galvanic skin response, heart rate, etc.)
Example: Caffeine Consumption Taking a blood sample to measure caffeine or its metabolites in blood
Scales of Measurement
Categorical/Nominal Scales
Variables classified as a category (aka, nominal)
Labels do not relate to each other quantitatively
Examples: animal species, eye color, brands
Quantitative Scales
The actual values matter and have some relation to each other
Examples: temperature, brain activity, steps per day, weight
Ordinal, interval, ratio
Ordinal Scales
Rank-order behaviors and events - 1st, 2nd, 3rd, etc.
The difference between each scale unit is not necessarily uniform - only the order matters
Interval Scales
Distance between points is equal and meaningful - no real zero
Allows researcher to specify distance between observations on a given dimension
Examples: IQ score, shoe size, temperature
Ratio Scales
Same as interval but with a real zero
Examples: Finishing an exam in 50 min takes twice as long as finishing in 25 minutes
Number of correct or incorrect answers can be zero
Many psychological measures have no meaningful "zero" - zero memory, zero intelligence...
Reliability
Three types of reliability
Test-retest
Internal
Interrater/inter-observer
Evaluating the reliability of observations
Scatterplot
Correlation coefficient (r)
Inter-observer agreement calculations
Test-retest reliability
Consistent scores every time the measure is used
Relevant when studying constructs that are expected to be stable over time (intelligence, personality, etc.)
Example: IQ test at beginning of semester (Time 1) and at the end of the semester (Time 2) - scores should be relatively consistent
Internal reliability (aka internal consistency)
A participant provides a consistent pattern of responses, regardless of how the researcher phrased the question or runs the session
Example: estimate of caffeine consumption (volume of caffeinated beverages) remains consistent across different questions
Not to be confused with internal validity!
Page 23: Internal Reliability (aka Internal Consistency)
Measure of how consistent a measure is within itself
Example: Estimating the number of caffeinated beverages consumed daily
Reliable responses: 4, 5, 4, 6, 5, 6, 4, 5, 5, 4, 4, 6
Unreliable responses: 8, 1, 5, 10, 2, 7, 5, 1, 1, 9, 4, 3
Page 24: Interrater (Inter-Observer) Reliability
Consistent scores regardless of who is measuring
Direct observation methods used (e.g., one-way mirrors, video recording)
Example: Observers counting the number of people who wipe their nose before touching a door handle
Page 25: How to Evaluate Reliability
Methods for evaluating reliability:
Scatterplot
Correlation coefficient (r)
Inter-observer agreement calculations
Page 26: Using a Scatterplot to Evaluate Reliability
Scatterplot used