Comprehensive Study Guide: Psychological Testing, Test Construction, Psychometric Properties, and Norms
Foundations of Psychological Testing and Measurement
Psychological testing is defined as a systematic process of obtaining a sample of behavior under standardized conditions, scoring responses according to specified rules, interpreting the resulting information, and supporting an inference about a psychological characteristic. Psychological characteristics—such as anxiety, intelligence, motivation, creativity, or personality—are abstract constructs that cannot be directly observed. This creates a fundamental problem in measurement: we must infer unobservable psychological attributes from observable behaviors and test responses.
The conceptual chain of psychological measurement moves from the individual person to their test responses, to a numerical score, and finally to an inference about a psychological characteristic. A person possesses thousands of experiences, thoughts, emotions, behaviors across contexts, relationships, and historical background. A test cannot capture the entirety of a person; it captures only a limited sample of behavior related to a specific construct. For example, to assess aggressiveness, a test might sample specific behaviors using items like "I get angry easily," "I lose my temper when people disagree with me," or "I sometimes act aggressively when frustrated." Because a test captures only a sample, the core scientific question is: "How justified is the inference drawn from that sample?"
To consider a statement like "Someone has high anxiety" scientific, the assessment process must be systematic, standardized, evidence-based, scorable, interpretive, and inferential. Unstandardized procedures prevent valid comparisons. For instance, if Student A is asked "Tell me about your anxiety," Student B is asked "How often do you feel nervous in social situations?", and Student C is asked "Rate your anxiety from 1 to 5," their responses cannot be meaningfully compared because the evaluation procedure itself differed.
Standardization establishes uniform rules to ensure that differences in test scores reflect true individual differences rather than variations in test administration. Standardization mandates consistency across several dimensions: what stimulus is presented, how it is presented, what specific instructions are given, how much time is allotted, how responses are scored, and how final scores are interpreted.
Quantifying psychological responses into numerical scores allows researchers and clinicians to summarize responses, compare scores across individuals or groups, examine statistical relationships, track changes over time, and support clinical or organizational decision-making. However, a number does not interpret itself. Obtaining a score of 42 on a self-esteem questionnaire is meaningless without knowing: What does 42 represent? 42 out of what maximum? How does this score compare to a normative group? What exact construct does the questionnaire measure? How should the score be interpreted in context?
To clarify what psychological testing represents, the following contrast highlights its defining attributes:
- Psychological testing is NOT: reading personality from someone's face, asking random questions, labeling a person based on a single score, diagnosing from a single questionnaire, assuming a higher score is always better, or treating a score as absolute truth.
- Psychological testing IS: systematic, standardized, scorable, interpretive, evidence-based, and inferential.
The Comprehensive Test Construction Process
Developing a standardized psychological measurement instrument requires a rigorous, multi-stage sequence. The test construction process consists of seven distinct phases:
- Planning Phase: The test developer defines broad and specific objectives, determines the construct to be measured, evaluates the need for the test compared to existing measures, and specifies the target population. The developer resolves logistical and structural decisions including the nature of content, test length, time limits, instructions, sampling methods, testing conditions, user qualification requirements, potential for administration harm, score interpretation models, and plans for the technical manual.
- Item Writing Phase: The developer creates a large item pool based on theoretical models, literature reviews, expert opinions, qualitative interviews, focus group discussions (FGDs), or existing scales. Items are drafted across target content domains, adhering to strict guidelines regarding clarity, vocabulary, difficulty, and discrimination power.
- Preliminary Administration (Tryout) Phase: The item pool undergoes empirical tryouts across three sub-phases: pre-tryout, initial tryout, and final tryout. This phase identifies weak or ambiguous items, establishes item difficulty, determines optimal time limits, evaluates instruction clarity, examines item intercorrelations, and assesses item validity.
- Reliability Estimation Phase: The developer computes reliability coefficients using a representative sample (typically requiring a minimum sample size of ) via methods such as test-retest, split-half, or alternate forms.
- Validity Estimation Phase: The developer gathers empirical and logical evidence supporting test score interpretation across content, criterion-related, and construct validity dimensions.
- Norm Development Phase: Benchmarks are established using a large, representative standardization sample to convert raw scores into age-equivalent, grade-equivalent, percentile, or standard scores.
- Manual Preparation Phase: A comprehensive technical manual is written, documenting administration procedures, scoring keys, technical psychometric properties (reliability and validity studies), normative tables, and academic references.
Item Writing Guidelines and Format Taxonomy
When writing items, test constructors must choose appropriate formats to capture participant responses effectively. DeVellis (1991) outlined fundamental guidelines for item writing:
- Define the target construct clearly based on established theory and literature (e.g., defining Self-Compassion as treating oneself with kindness, recognizing suffering as a shared human experience, and maintaining balanced emotional awareness).
- Identify specific dimensions if the construct is multidimensional (e.g., Academic Motivation split into intrinsic and extrinsic dimensions).
- Generate an extensive item pool.
- Avoid exceptionally long items, as length impairs reading comprehension.
- Keep reading difficulty tailored to the target examinee population.
- Avoid "double-barreled" items that combine two or more distinct ideas into a single statement.
- Focus each item on a single concept using clear, concise language (e.g., prefer "I feel confident when speaking in class" over "I possess unwavering confidence while articulating my viewpoints").
- Avoid double negatives (e.g., write "I enjoy attending lectures" rather than "I do not dislike attending lectures").
- Avoid ambiguous frequency words such as usually, often, sometimes, or frequently unless strictly necessary and explicitly defined.
- Mix positively and negatively worded items sparingly to detect and counteract the acquiescence response set (the tendency to agree with statements regardless of content). For example, Positive: "I remain calm during stressful situations"; Negative: "I panic easily during stressful situations."
- Maintain sensitivity to ethnic, cultural, and demographic differences.
Item formats are broadly categorized into selected-response formats (where examinees choose from given options) and constructed-response formats (where examinees generate their own responses).
Constructed-Response Formats
Constructed-response formats include essay items, short-answer items, and completion items.
Essay items (also called free-answer items) require examinees to retrieve information from memory and write a extended response demonstrating recall, comprehension, analysis, and synthesis. Short-answer items require one or two lines, whereas long-answer items require several sentences or paragraphs. While suitable for assessing complex skills, originality, and critical thinking, essay marking is often highly unreliable and time-consuming.
To score essay items objectively, two primary scoring methods are used:

- Sorting Method: Answer sheets are sorted into distinct stacks based on overall quality or fairness of the response. The scorer assigns specific weightings to each stack and sums them to yield total scores.
- Point Score Method: The scorer creates a detailed grading key containing specific target points and assigned numerical values, comparing examinee responses directly against the key.
Completion items require test-takers to supply a missing word or phrase in a sentence. Completion items must be worded tightly so that only one specific, correct answer is accurate; ambiguous items lead to scoring inconsistency.
Selected-Response Formats
Selected-response items are systematically categorized into supply and selection types:

- Dichotomous Format: Offers two mutually exclusive alternatives (e.g., True/False, Yes/No). Selecting the target alternative earns one point. Dichotomous items are simple to administer, easy to understand, and quick to score. However, they encourage rote memorization, fail to capture complex understanding, and are susceptible to guessing, making them less reliable and precise than other formats. Items must be simple, straightforward, absolute in judgment, entirely true or entirely false, free of double negatives, and attributed to a specific authority where applicable.
- Polytomous (Polychotomous) Format: Common in multiple-choice tests, this format includes a stem, one correct (keyed) alternative, and several incorrect options termed distractors or foils. Multiple-choice items lower chance guessing probabilities compared to dichotomous items and offer rapid, objective scoring. Keyed responses must be varied in position and maintained at a consistent length relative to distractors.
- Distractor Efficiency: Psychometric theory indicates that adding plausible distractors increases item reliability. However, in practice, items rarely contain more than three or four distractors that operate efficiently. Poorly written, implausible, or overly easy distractors lower reliability and validity by allowing unprepared examinees to guess correctly.
- Matching Items: Present two columns: premises on the left and options/responses on the right. To minimize guessing, the response column should contain more items than the premise column.
- Likert Format: Measures attitude or personality dimensions by asking respondents to indicate their level of agreement on a continuum (typically 5 options: Strongly Disagree, Disagree, Neutral, Agree, Strongly Agree). Negatively worded items are reverse scored prior to summing total scores. Likert data can be subjected to factor analysis to identify item clusters measuring shared constructs.
- Category Format: Extends response choices to a wider scale (e.g., 1 to 10). To prevent respondents from altering ratings due to context shifts, the scale endpoints must be clearly defined and repeatedly emphasized.
- Checklists and Q-Sorts: Adjective checklists require respondents to select descriptive traits applying to themselves or others. The Q-sort technique (Stephenson, 1953) presents statements that respondents sort across nine piles to describe themselves or evaluate others.
- Computer Administration: Modern testing software utilizes item banks (reservoirs of pre-calibrated items) and item branching (adaptive testing where item difficulty adjusts dynamically based on ongoing accuracy).
Quantitative and Qualitative Item Analysis
Item analysis evaluates the psychometric performance of individual test items to determine which items should be retained, revised, or discarded. It incorporates analyses of item reliability, item difficulty, item discriminability, and item validity. Eliminating defective items removes chance error, making a test shorter, more efficient, more reliable, and more valid.
Item Difficulty ()
In power tests (achievement or ability assessments where examinees have sufficient time to complete all items), item difficulty is defined as the proportion of examinees who answer an item correctly. It is calculated using the formula:
where represents the item difficulty index, is the number of examinees providing the correct response, and is the total number of examinees taking the test. Subscript notation indicates specific item numbers (e.g., denotes the difficulty index for item 1).
Although termed "difficulty," higher values indicate easier items (e.g., means answered correctly), while lower values indicate harder items. In non-achievement contexts (such as personality assessment), this statistic is referred to as an item-endorsement index. Estimating difficulty via expert qualitative ranking is subjective and unreliable compared to statistical calculation.
If everyone answers an item correctly () or incorrectly (), the item provides zero variance and fails to yield individual difference information. To maximize statistical information regarding individual differences, items generally should fall within a difficulty range of to .
Item Discriminability ()
Item discriminability evaluates how adequately an item differentiates between high scorers and low scorers on the overall test. The item discrimination index is symbolized by and yields three general outcomes:
- Positive Discrimination (): A higher proportion of high total-scorers answer the item correctly compared to low total-scorers (indicates a well-functioning item).
- Negative Discrimination (): A higher proportion of low total-scorers answer the item correctly compared to high total-scorers (indicates a defective, ambiguous, or mis-keyed item that must be revised or eliminated).
- Non-discriminating (): High and low total-scorers perform identically on the item.
The Extreme Group Method calculates through the following procedural steps:
- Identify a high-performing group (e.g., the top third or top of total test scores) and a low-performing group (e.g., the bottom third or lower ).
- Calculate the proportion of correct responses in the high group ( or ) and low group ( or ).
- Subtract the low group proportion from the high group proportion:
Alternatively, this is expressed as:
where and represent correct responses in the upper and lower groups, and and represent total examinees in each group.
The following table illustrates item discriminability calculations using upper and lower class thirds:

Evaluating the items from this example:
- Item 1: , (Strong positive discrimination; retain).
- Item 2: , (Good positive discrimination; retain).
- Item 3: , (Good positive discrimination; retain).
- Item 4: , (Poor discrimination due to ceiling effect/over-easiness; consider revising).
- Item 5: , (Negative discrimination; defective item requiring elimination or major revision).
Item Analysis for Speed Tests
Standard item analysis applied to speed tests yields misleading or uninterpretable results. In speeded tests, items occurring near the end appear artificially difficult simply because examinees ran out of time before reaching them. To correct for speed effects, difficulty calculations must restrict the sample size denominator strictly to test-takers who reached the specific item ():
Even with this adjustment, analyzing speed test items presents three significant limitations:
- Later items are evaluated on progressively smaller sample sizes (), decreasing statistical reliability.
- If faster, higher-ability examinees reach later items, the sample becomes biased rather than representative.
- Because higher-ability examinees dominate the late-item sample, items near the end appear artificially easier than they truly are.
Item Characteristic Curves (ICC)
An Item Characteristic Curve (ICC) is a graphic representation of item difficulty and discrimination. The horizontal () axis represents total test performance (divided into score intervals, e.g., 51–55, 56–60, …, 96–100), and the vertical () axis represents the proportion of examinees within each interval who answered the specific item correctly:

For a "good" test item, the proportion of correct responses increases monotonically as total test score increases. Items can also be evaluated by plotting discriminability (point-biserial correlation with total score) against difficulty ():

Items falling within the shaded region—having a discriminability index above and a difficulty level between and —represent optimal candidates for final test inclusion. For example, Item 12 was passed by of respondents () and correlated with total test score, justifying its retention.
Item Response Theory (IRT) and Advanced Scaling
Item Response Theory (IRT) is a modern psychometric framework where each item possesses its own item characteristic curve defining the exact probability of answering correctly given an examinee's latent ability level. IRT models are categorized by the number of item parameters evaluated:
- 1-Parameter Model (Rasch Model): Evaluates item difficulty.
- 2-Parameter Model: Evaluates item difficulty and item discriminability.
- 3-Parameter Model: Evaluates item difficulty, item discriminability, and pseudo-chance/guessing probability (the likelihood that examinees with the lowest ability score correctly by chance).
IRT enables adaptive testing and item branching: computer algorithms sample pre-calibrated items from an item bank, pinpointing an examinee's precise ability threshold without requiring administration of every test item.
Psychometric Properties: Reliability and Classical Test Theory
Reliability refers to the degree of consistency, stability, and precision with which a test measures a psychological characteristic. It indicates the extent to which a test yields identical scores when measurement is repeated under comparable conditions.
A fundamental psychometric rule states that reliability is not an absolute property of the test instrument itself, but a property of the scores obtained from a specific administration. Reliability varies based on sample characteristics, testing conditions, administrative procedures, and test purpose. Psychologists evaluate the reliability of test scores under specific operational conditions.
In physical measurement, instruments achieve high precision; in psychology, constructs are abstract and inferred indirectly, making measurement error inevitable. Reliability reflects how effectively error variance has been minimized.
Classical Test Theory (CTT)
Classical Test Theory (CTT) posits that every observed test score consists of two components: a true score component and an error score component. This is expressed by the equation:
where represents the observed score, represents the true score, and represents measurement error.
- Observed Score (): The actual numerical score achieved by an individual on a test.
- True Score (): A hypothetical value representing the individual's actual ability level if all measurement error were eliminated. If an examinee were tested an infinite number of times under identical conditions, the average of all observed scores would equal their true score.
- Measurement Error (): Variations caused by factors unrelated to the construct being measured.
Expressed through variance components, total observed score variance is equal to true score variance plus error variance:
Mathematically, the reliability coefficient () is defined as the proportion of total observed score variance attributable to true score variance:
Reliability coefficients range from to :
- indicates perfect measurement precision (zero error variance).
- indicates that all score variation stems entirely from measurement error.
- Well-constructed standardized tests generally require reliability coefficients above . Coefficients above are required for high-stakes individual decisions such as clinical diagnosis or employment selection.
Nature of Measurement Error
Measurement errors are classified into two distinct types:
- Random Errors: Unpredictable, unsystematic fluctuations occurring from one testing occasion to another (e.g., temporary fatigue, anxiety, environmental noise, sudden illness, or momentary distraction). Random errors fluctuate in either direction, impacting score consistency and lowering reliability.
- Systematic Errors: Consistent distortions operating in a single direction (e.g., miscalibrated instruments, defective scoring keys, biased raters, or ambiguous phrasing). Systematic errors consistently distort measurement, degrading both reliability and validity.
Types of Reliability
The major empirical methods for estimating reliability are detailed below:
1. Test–Retest Reliability (Coefficient of Stability)
Evaluates score stability over time by administering the same test to the same group on two separate occasions and calculating the correlation coefficient between the score sets. It is suitable for measuring stable traits (e.g., intelligence, personality traits, aptitudes, vocational interests). The standard time gap is 15 days. Limitations include practice effects, memory retention, maturation, and temporal shifts in motivation or mood.
2. Alternate-Forms (Parallel-Forms) Reliability
Evaluates consistency across two equivalent test versions constructed from the same specification blueprint (matching in item count, content, difficulty, and statistical structure). Both forms are administered to the same examinees, and their scores are correlated. Common in educational and competitive testing, this method minimizes memory effects but is time-consuming and costly to construct.
3. Split-Half Reliability and Internal Consistency
Estimates internal consistency from a single test administration. The test is divided into two equivalent halves (typically via odd-numbered vs. even-numbered items), and scores on the two halves are correlated. Because halving test length underestimates full-length test reliability, the Spearman–Brown prophecy formula must be applied:
Internal consistency across all possible item split combinations can be evaluated using Cronbach's alpha (). This approach is economical and ideal for homogeneous tests measuring a single construct.
4. Inter-Rater (Inter-Scorer) Reliability
Measures the extent of agreement between independent scorers evaluating subjective assessments (e.g., essays, projective instruments, behavioral observations, clinical interviews). Inter-rater reliability is maximized by providing detailed scoring rubrics, standardized administration protocols, objective scoring criteria, and comprehensive examiner training.
Summary of Factors Affecting Reliability
Reliability is influenced by: test length (longer tests generally increase reliability), quality of test items, heterogeneity of the sample (wider variance increases reliability coefficients), standardization of testing conditions, and examinee characteristics.
Psychometric Properties: Validity and Scientific Evidence
Validity refers to the degree to which empirical evidence and theoretical rationale support the interpretations of test scores for their intended purposes. While reliability focuses on consistency, validity focuses on accuracy, meaningfulness, and whether conclusions drawn from scores are scientifically justified.
Types of Validity
1. Face Validity
Refers to whether a test superficially appears to measure its target construct upon informal inspection by examinees, users, or lay observers. It relies on subjective impression rather than empirical data. Face validity is the psychometrically weakest form of validity and is never sufficient on its own.
2. Content Validity
Evaluates how comprehensively and representatively test items sample the target domain of knowledge, behavior, or skills. Content validity is vital in educational achievement tests, licensing examinations, and professional competency assessments. It is established qualitatively by subject matter experts using a blueprint or table of specifications detailing target topics, cognitive levels, and content proportions.
3. Criterion-Related Validity
Evaluates how effectively test scores correlate with or predict scores on an external, objective criterion measure. It is divided into two designs:
- Concurrent Validity: The test score and external criterion are collected at approximately the same time. Used to validate new, shorter, or less expensive instruments against established gold-standard measures.
- Predictive Validity: The test score is evaluated against a future criterion measured after a specified time interval. It is crucial for tests used in selection, placement, university admissions, and career counseling.
4. Construct Validity
Evaluates whether a test accurately measures an abstract, non-observable psychological construct (e.g., intelligence, self-esteem, resilience, anxiety, emotional intelligence). Construct validity cannot be established in a single study; it requires accumulating evidence across multiple empirical sources:
- Convergent Validity: Requires test scores to correlate strongly with scores from established instruments measuring identical or theoretically related constructs.
- Discriminant (Divergent) Validity: Requires test scores to show weak or non-significant correlations with instruments measuring theoretically unrelated constructs.
- Factor Analysis: A statistical method that examines item intercorrelations to verify whether empirical factor loadings match the theoretical dimensional structure of the construct.
Factors Affecting Validity
Validity is degraded by poorly defined constructs, ambiguous test items, inadequate domain content sampling, cultural bias, language barriers, improper testing procedures, examiner bias, examinee guessing, careless responding, and unreliable scoring rules.
Norm Development, Standard Scores, and Test Scaling
Norms represent the average, typical performance of an identified group on a standardized measure. Psychological tests are interpreted by comparing individual raw scores against normative distributions.
To establish norms, test developers administer the test under standardized conditions to a large standardization sample (normative group) that accurately represents the target population. Raw scores are converted into derived scores—measures sharing standardized units that allow comparison of an individual's score against the normative sample, across different subscales, or across different tests.
Norm-Referenced vs. Criterion-Referenced Evaluation
- Norm-Referenced Tests: Compare an examinee's performance to the relative performance of a standardized comparison group. Raw scores are transformed into derived scores to establish relative percentile or standard score standing.
- Criterion-Referenced Tests: Compare an examinee's performance directly against a predetermined cut-off score, absolute standard, or explicit performance criterion (e.g., mastering of curriculum objectives), regardless of peer performance.
Steps in Developing Norms
- Define the target population based on the intended scope of the test.
- Select a representative sample from the target population.
- Standardize testing conditions and administrative controls.
- Administer the test and gather raw scores.
- Perform statistical analyses to calculate central tendency (Mean), variability (Standard Deviation), and score distributions.
Taxonomy of Norm Types and Test Scales
1. Age-Equivalent Norms
Calculated by determining the average score obtained by representative samples at specific age levels. They are suitable for cognitive abilities or physical traits that increase systematically with age during childhood and adolescence. Disadvantages: lacks uniform measurement units across growth stages (mental growth from age 3 to 4 is far greater than from age 9 to 10), growth rates decelerate near adulthood, and non-linear trait growth rates make cross-trait comparison invalid.
2. Grade-Equivalent Norms
Calculated by determining the mean raw score earned by students in each academic grade level (e.g., if fourth-graders average 23 correct problems on an arithmetic test, a raw score of 23 equals a grade equivalent of 4.0). Shortcomings: applicable only to continuous subjects taught across all grades; assumes uniform curriculum exposure; non-comparable across subjects. A fourth-grader scoring a grade equivalent of 6.9 reflects superior fourth-grade performance rather than mastery of sixth-grade academic curriculum.
3. Percentile Norms
Percentiles express an individual's relative standing by indicating the percentage of examinees in the normative sample who scored below that individual. For example, if a raw score of 26 exceeds the scores of of the standardization group, the raw score of 26 corresponds to a Percentile Rank () of 40. Percentiles are easy to compute and understand.
4. Standard Score Norms
Standard scores express an individual's distance from the normative mean in Standard Deviation () units, rendering scores across different tests directly comparable:
- Z-Score: Indicates how many standard deviations an observed raw score lies above or below the mean:
where is the raw score, is the normative mean, and is the normative standard deviation. A -score scale has a mean of and an of . Scores below the mean yield negative values. Drawback: negative numbers and decimals can be cumbersome to manage.
- T-Score (W. A. McCall, 1922): Transforms -scores to eliminate negatives and decimals using a fixed mean of and an of :
-scores range from 20 to 80 (where represents above the mean). They are widely used in clinical and personality assessment.
Sten Score: A 10-point scale commonly used in personality assessments. It has a mean of and an of approximately . Scores 1–3 indicate below-average performance, 4–7 indicate average, and 8–10 indicate above-average performance.
Stanine Score: A contraction of "Standard Nine," providing a single-digit scale ranging from 1 to 9. It has a mean of and an of (approximately ). Stanines 1–3 represent below average, 4–6 represent average, and 7–9 represent above average. Stanines facilitate comparisons of student performance across different academic content domains.
To ensure score stability and generalizability, normative samples must be sufficiently large, fully representative of the target population, and clearly specified in the technical documentation.