Threats to Psychometric Quality (9/25/2026)

0.0(0)
Studied by 0 people
call kaiCall Kai
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/26

encourage image

There's no tags or description

Looks like no tags are added yet.

Last updated 3:03 PM on 9/25/26
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

27 Terms

1
New cards

Goals of coping w/response biases

Prevent or minimize existence of bias, minimize the effects of bias, detect bias and intervene

2
New cards

Strategies for coping with response biases

Manage testing context, manage test content or scoring, use specialized tests

3
New cards

Examples of preventing/minimizing biases via managing testing context

Anonymity, minimize frustrations, warnings

4
New cards

Examples of preventing/minimizing biases via managing testing content/scoring

Simple, clear items (not for faking), neutral items/subtle items (would cause dimensionality problems), forced choice or minimal choice items, balanced scales including reverse items (would cause poor item quality), corrections for guessing (IRT 3PL)

5
New cards

Examples for detecting bias and intervening for managing testing content/scoring

Embedded validity scales (ex, MMPI, L scale (lie scale)), IRT

6
New cards

Forced-choice items

Many commercial tests employ this format. There are multiple types such as:

  • Binary preference

  • Blocks

  • Ranking

  • Graded preference

  • Proportional preference etc.

The level of social desirability should be the same or at least very similar.

7
New cards

Advantages of forced choice items

  • Reduce response biases (ex. acquiescence bias, modesty bias)

  • Increase dimensionality by preventing “halo” effects

  • Do not need verbal and non-verbal anchors in some types

  • Reduce faking


8
New cards

Disadvantages of forced choice items

  • Ipsative scores (on another flashcard)

  • To avoid/alleviate ipsative scores, need some methods including advanced IRT

  • More difficult in development (need data analysis skills for FC)

  • Need more data collection


9
New cards

(Pure) Ipsative scores

Only provide a comparison within an individual. Cannot compare your score and other people’s scores. NOT recommended to be used for recruitment and selection purposes. Helpful for coaching and development purposes. Should not be used for selection decisions.

10
New cards

Limitations of ipsative scores

May not conduct factor analysis (may conduct Principal Components analysis but the results are difficult to interpret), may not calculate internal consistency, distorts criterion-related validity estimates. FA and alpha cannot be used for even partial ipsative data. Maydeu-Olivares & Brown (2010) used multidimensional IRT to confirm dimensionality and reliability

11
New cards

To alleviate the impact of ipsative constraints

  • Increase the # of measured traits using the multidimensional FC format (30 or more)

  • Include both negative and positive items and take points away when selecting negative statements

  • Use partially order statements or do not score all ranked statements

  • Use advanced IRT models (e.g., MUPP or TIRT) to transform ipsative scores into normative scores


12
New cards

Test bias

Systematically obscures differences between groups (and people from different groups)

13
New cards

Construct bias

Group difference in score meaning. Differences in test scores (between groups) might not reflect true group differences in psychological construct. Often occurs in cross-cultural contexts but also may occur in domestic tests (i.e., “friend” and “tomodachi” example)

14
New cards

Predictive bias

Group difference in implications of score use. Scores are associated w/an important criterion to differing degrees in dif. groups (test use)

15
New cards

Test bias importance

Testing is widespread and has important consequences, so tests should differentiate among people based on real psychological differences rather than on group membership. Can compromise decisions about individuals and researchers’ study of group differences

16
New cards

Detecting construct bias

Usually detected by examining responses to items on a test. Note, the existence of group differences in test responses, on its own, does not necessarily imply bias (ex. Extraversion in the U.S. and Japan). Bias exists when there are group differences in test responses and when those differences do NOT reflect true psychological differences

17
New cards

Steps for detecting construct bias

  1. Differential reliability

  2. Differential rank order of item difficulties

  3. Differential item discrimination index

  4. Differential dimensionality

  5. Differential item functioning


18
New cards

Differential reliability

Group differences in reliability. Test scores reflect the relevant construct with better precision in one group than in another. Statistical techniques for formally testing group differences in reliability. Examine internal structures

19
New cards

Rank order of difficulties

The items that are most difficult (as compared to other items) in one group are not the most difficult in another group (may indicate construct bias)

20
New cards

Differential item discrimination

Compute an item discrimination index (for an item) separately for each group. Are those indices different across groups? If so, may indicate construct bias in that item. Particularly problematic if this occurs for many items.

21
New cards

Differential dimensionality

Dimensionality: number of dimensions/factors, connections btwn items and factors, correlations among factors.

Is factor structure dif. across groups? If so, may indicate construct bias. Best done via CFA

22
New cards

Differential item functioning

Compare psychometric properties of an item across groups (done for all items). Akin to differential item discrimination and differential dimensionality. Based on IRT. There are multiple DIF detection methods. DFIT, Likelihood ratio test (LRT) and logistic regression modeling

23
New cards

Predictive bias

If a test is more predictive (of an important outcome) for some groups than for others, then it suffers from predictive bias. Predictive bias would be caused by various factors including response bias, differences in educational systems, low reliability etc

24
New cards

Detecting predictive bias

To what degree are scores more predictive of a criterion in 1 group than in another?

  1. Administer test to people in each group

  2. Measure criterion for all those people

  3. Conduct statistical analyses to detect whether test scores are associated w/criterion differently in the groups

Usually via linear regression


25
New cards

Slope bias

Do two groups have the same slope (b)? No = “slope bias".” No = test scores are more strongly linked to the criterion in 1 group than the other group. No = test scores are more accurately predictive of the criterion in one group than in the other

26
New cards

Detecting predictive bias

In practice, researchers often use more advanced forms of regression (regression w > 1 predictor) to detect predictive bias. But basic idea is the same - do groups have dif. intercepts and dif. slopes

27
New cards

Recombination by Meade & Tonidandel (2010)

Step 1: conduct internal analyses examining differential functioning of items and the test

Step 2: before conducting regression analyses, examine both test and criterion scores for significant group mean differences via t-tests and compute descriptive statistics for each group

Step 3: Compute d effect size estimates for group mean differences for both the test and the criterion.

Step 4” conduct the regression-based differential prediction analyses by regressing the criterion onto the test, the dichotomous group variable, and the interaction between the two