1/26
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Goals of coping w/response biases
Prevent or minimize existence of bias, minimize the effects of bias, detect bias and intervene
Strategies for coping with response biases
Manage testing context, manage test content or scoring, use specialized tests
Examples of preventing/minimizing biases via managing testing context
Anonymity, minimize frustrations, warnings
Examples of preventing/minimizing biases via managing testing content/scoring
Simple, clear items (not for faking), neutral items/subtle items (would cause dimensionality problems), forced choice or minimal choice items, balanced scales including reverse items (would cause poor item quality), corrections for guessing (IRT 3PL)
Examples for detecting bias and intervening for managing testing content/scoring
Embedded validity scales (ex, MMPI, L scale (lie scale)), IRT
Forced-choice items
Many commercial tests employ this format. There are multiple types such as:
Binary preference
Blocks
Ranking
Graded preference
Proportional preference etc.
The level of social desirability should be the same or at least very similar.
Advantages of forced choice items
Reduce response biases (ex. acquiescence bias, modesty bias)
Increase dimensionality by preventing “halo” effects
Do not need verbal and non-verbal anchors in some types
Reduce faking
Disadvantages of forced choice items
Ipsative scores (on another flashcard)
To avoid/alleviate ipsative scores, need some methods including advanced IRT
More difficult in development (need data analysis skills for FC)
Need more data collection
(Pure) Ipsative scores
Only provide a comparison within an individual. Cannot compare your score and other people’s scores. NOT recommended to be used for recruitment and selection purposes. Helpful for coaching and development purposes. Should not be used for selection decisions.
Limitations of ipsative scores
May not conduct factor analysis (may conduct Principal Components analysis but the results are difficult to interpret), may not calculate internal consistency, distorts criterion-related validity estimates. FA and alpha cannot be used for even partial ipsative data. Maydeu-Olivares & Brown (2010) used multidimensional IRT to confirm dimensionality and reliability
To alleviate the impact of ipsative constraints
Increase the # of measured traits using the multidimensional FC format (30 or more)
Include both negative and positive items and take points away when selecting negative statements
Use partially order statements or do not score all ranked statements
Use advanced IRT models (e.g., MUPP or TIRT) to transform ipsative scores into normative scores
Test bias
Systematically obscures differences between groups (and people from different groups)
Construct bias
Group difference in score meaning. Differences in test scores (between groups) might not reflect true group differences in psychological construct. Often occurs in cross-cultural contexts but also may occur in domestic tests (i.e., “friend” and “tomodachi” example)
Predictive bias
Group difference in implications of score use. Scores are associated w/an important criterion to differing degrees in dif. groups (test use)
Test bias importance
Testing is widespread and has important consequences, so tests should differentiate among people based on real psychological differences rather than on group membership. Can compromise decisions about individuals and researchers’ study of group differences
Detecting construct bias
Usually detected by examining responses to items on a test. Note, the existence of group differences in test responses, on its own, does not necessarily imply bias (ex. Extraversion in the U.S. and Japan). Bias exists when there are group differences in test responses and when those differences do NOT reflect true psychological differences
Steps for detecting construct bias
Differential reliability
Differential rank order of item difficulties
Differential item discrimination index
Differential dimensionality
Differential item functioning
Differential reliability
Group differences in reliability. Test scores reflect the relevant construct with better precision in one group than in another. Statistical techniques for formally testing group differences in reliability. Examine internal structures
Rank order of difficulties
The items that are most difficult (as compared to other items) in one group are not the most difficult in another group (may indicate construct bias)
Differential item discrimination
Compute an item discrimination index (for an item) separately for each group. Are those indices different across groups? If so, may indicate construct bias in that item. Particularly problematic if this occurs for many items.
Differential dimensionality
Dimensionality: number of dimensions/factors, connections btwn items and factors, correlations among factors.
Is factor structure dif. across groups? If so, may indicate construct bias. Best done via CFA
Differential item functioning
Compare psychometric properties of an item across groups (done for all items). Akin to differential item discrimination and differential dimensionality. Based on IRT. There are multiple DIF detection methods. DFIT, Likelihood ratio test (LRT) and logistic regression modeling
Predictive bias
If a test is more predictive (of an important outcome) for some groups than for others, then it suffers from predictive bias. Predictive bias would be caused by various factors including response bias, differences in educational systems, low reliability etc
Detecting predictive bias
To what degree are scores more predictive of a criterion in 1 group than in another?
Administer test to people in each group
Measure criterion for all those people
Conduct statistical analyses to detect whether test scores are associated w/criterion differently in the groups
Usually via linear regression
Slope bias
Do two groups have the same slope (b)? No = “slope bias".” No = test scores are more strongly linked to the criterion in 1 group than the other group. No = test scores are more accurately predictive of the criterion in one group than in the other
Detecting predictive bias
In practice, researchers often use more advanced forms of regression (regression w > 1 predictor) to detect predictive bias. But basic idea is the same - do groups have dif. intercepts and dif. slopes
Recombination by Meade & Tonidandel (2010)
Step 1: conduct internal analyses examining differential functioning of items and the test
Step 2: before conducting regression analyses, examine both test and criterion scores for significant group mean differences via t-tests and compute descriptive statistics for each group
Step 3: Compute d effect size estimates for group mean differences for both the test and the criterion.
Step 4” conduct the regression-based differential prediction analyses by regressing the criterion onto the test, the dichotomous group variable, and the interaction between the two