Measuring Intelligence: Scales, Scales, Indices, and Psychometric Evaluation
The Binet-Simon Scale (1905)
Original Objective: Alfred Binet and Theodore Simon developed the scale with the specific goal stated as: ‐ ‐Our aim is… to take the measurement of [the child’s] intellectual powers, in order to establish whether he is normal or if he is retarded‐‐. While the term ‐‐retarded‐‐ is currently outdated and offensive, it was the standard terminology for low intelligence at the time.
Methodology: Binet and Simon took a systematic approach to measurement by writing test items designed for different age groups.
Standardization: The test was standardized across age groups. Children would take each level of the test and progressively move to the next ‐‐age‐‐ test until they encountered one they could not pass.
Mental Age: Based on the age-normed items a child could pass, they were assigned a ‐‐mental age.‐‐
Children with low intelligence were identified as having a mental age lower than their actual chronological age.
This established the first formal conceptualization of the intelligence test.
Lewis Terman and the Stanford-Binet Scale (1916)
Introduction of IQ: Lewis Terman developed the Stanford-Binet scale and introduced the concept of the Intelligence Quotient or .
The Original IQ Formula:
Calculatory Examples:
A -year-old child who passes all items designed for age : (indicated as average intelligence).
A -year-old child who passes items designed for age : This results in an above (indicated as above average).
A -year-old child who cannot pass items designed for age : This results in an below (indicated as below average intelligence).
The Age Capping Problem:
The highest items on the original scale for measuring mental age were capped at age .
Because the chronological age (the denominator) continues to rise while the mental age (the numerator) stays at a maximum of , calculated scores would naturally decrease as a person aged past .
For example, a person with a mental age of would have an of at age , but that same performance at age would result in an of approximately . Therefore, it was not possible to accurately use this specific scale for older individuals.
Modern IQ Measurement: The Deviation Score
Population Relation: Modern intelligence testing no longer uses the mental age/chronological age formula. Instead, is calculated as a deviation score in relation to the general population.
Normal Distribution: Scoring is based on a normal distribution (the bell-shaped curve).
The mean () is set at .
The standard deviation () is set at .
Classifications: This statistical bound allows for the classification of intellectual disabilities, average performance, and gifted performance.
The Stanford-Binet: Fifth Edition (SB5)
Structure: The current version yields a Full-scale based on core subtests which correlate to five distinct factors.
The Five Factors:
Fluid Reasoning (Fluid Intelligence): Relates to solving novel problems; measured via tasks like matrices.
Knowledge (Crystallised Intelligence): Measured through tasks such as vocabulary assessment.
Quantitative Reasoning: Measures general numerical ability.
Visual-Spatial Reasoning: Examines the ability to detect patterns in visual stimuli.
Working Memory: Assesses the ability to store and manipulate information in short-term memory.
Wechsler Intelligence Scales
David Wechsler: Developed alternative scales to the Stanford-Binet for different age groups.
Scale Variants:
WAIS-IV (2008): Wechsler Adult Intelligence Scale, the current version for adults.
WISC-V (2014): Wechsler Intelligence Scale for Children (used for ages to ).
WPPSI: Wechsler Preschooler & Primary Scale of Intelligence (used for children under age ).
WAIS-IV Multi-Index Structure: Full-scale (FSIQ) is broken down into four core indices:
Verbal Comprehension Index (VCI)
Working Memory Index (WMI)
Perceptual Reasoning Index (PRI)
Processing Speed Index (PSI)
WAIS-IV Subtest Examples
Verbal Comprehension Index (VCI)
Similarities: ‐‐In what ways are apples and pears alike?‐‐
Vocabulary: ‐‐What is a guitar?‐‐ (Influenced by the depth of vocabulary knowledge).
Information: ‐‐What is the capital of France?‐‐ (Tests general knowledge).
Comprehension: ‐‐Why are we tried by a jury of our peers?‐‐
Working Memory Index (WMI)
Digit Span: Participant must repeat a sequence of numbers back to the examiner (e.g., ).
Arithmetic: Mental calculations without paper and pencil. Example: ‐‐Imagine that you bought six postcards for 45\text{\cent} each. How much change would you receive back from $5?‐‐
Letter-Number Sequencing: Mentally transforming a sequence. Example: Repeat ‐‐‐‐ but place numbers in numerical order first, then letters in alphabetical order (Result: ).
Perceptual Reasoning Index (PRI)
Block Design: Participants are provided with physical blocks and must use them to copy specific patterns.
Matrix Reasoning: Finding patterns that match specific rules going across and down a grid. For example, matching a pattern where shapes change from star to pentagon vertically while colors layer in a specific sequence horizontally.
Visual Puzzles: Combining numbered shapes/parts to form a target shape (e.g., combining parts , , and to make a whole puzzle).
Picture Completion: Identifying missing elements in a image (e.g., a car missing wheels or a balloon missing a string).
Figure Weights: A complex task requiring the participant to balance scales. Example logic: if stars equal green pentagon, and a red circle equals a blue square plus a star, find the missing match for an additional star.
Processing Speed Index (PSI)
Symbol Search: A timed task where the individual identifies if specific symbols are present within a larger group.
Coding: Transposing symbols into boxes based on a numerical code key, similar to a codebreaker puzzle.
Cancellation: Moving as quickly as possible to cross out all shapes matching specific targets (e.g., red squares and yellow triangles).
Raven’s Progressive Matrices
Nature of the Test: A non-verbal test consisting of matrices presented in order of increasing difficulty.
Task: Identify the missing element that completes a visual pattern based on horizontal and vertical rules.
Construct Measured: It is considered a strong measure of reasoning ability, often referred to as Spearman’s .
Advantages:
Independent of language, reading, and writing skills.
Considered an unbiased test of intelligence.
Frequently used for job selection.
Properties of Intelligence Tests
Reliability: Measured via test-retest correlation (comparing scores from the same person on two different occasions).
IQ tests generally demonstrate high reliability, with correlations around (where represents perfect reliability).
Validity: Measured by correlating test scores with theoretically related criteria.
Example: Success in school.
IQ tests show mid-to-strong positive correlations ( to ) with academic success.
Evaluation of IQ Testing
Arguments For Testing
Diagnostic Necessity: Essential for identifying learning disabilities or assessing the impact of brain damage to ensure appropriate treatment and education.
Objectivity: Standardized tests are considered more objective and less biased than subjective methods.
Predictive Power: IQ scores are moderately to strongly correlated with several life outcomes:
IQ & School Grades: to
IQ & Years of Schooling: to
IQ & Occupational Attainment: to
IQ & Job Performance: < 0.3 to > 0.5
Arguments Against Testing
Validity Concerns: Some argue they measure learned knowledge rather than the inherent ability to learn.
Cultural Bias: Tests may be biased toward Western cultures, dominant group members, and those with rich academic backgrounds.
Labeling and Stigma: Being grouped by a score can lead to ‐‐self-fulfilling prophecies,‐‐ where individuals perform poorly because they believe the negative label attached to them.
Narrow Scope: Conventional tests fail to capture other forms of intelligence, such as emotional intelligence.