Psychological Testing and Assessment: Reliability, Validity, Norms, and Test Development

0.0(0)
Studied by 0 people
call kaiCall Kai
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/38

flashcard set

Earn XP

Description and Tags

Flashcards covering psychological testing and assessment topics including types of reliability, face, content, criterion-related, and construct validity, rating errors, test standardization, norms, and item-writing guidelines.

Last updated 6:40 PM on 10/7/26
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

39 Terms

1
New cards

When is parallel forms reliability appropriate to use?

It is appropriate for tests that have two forms.

2
New cards

When is test-retest reliability appropriate to use?

It is appropriate for tests that are designed to be administered to an individual more than once.

3
New cards

Which reliability estimate is appropriate for tests demonstrating factorial purity?

Cronbach's coefficient alpha.

4
New cards

When is split-half reliability appropriate to use?

It is appropriate for tests with items carefully ordered according to difficulty.

5
New cards

When should inter-rater reliability be utilized?

It is used for tests that involve some degree of subjective scoring.

6
New cards

What is the Kuder-Richardson 20 (KR20\text{KR}_{20}) formula used for?

It is the statistic used for calculating the reliability of a test in which items are dichotomous or scored as 00 or 11 (forced-choice items).

7
New cards

What is face validity?

The simplest and least stringent form of validity, referring to whether a test looks valid at face value or superficial appearance to test users, examiners, and examinees.

8
New cards

Why do inkblot tests typically demonstrate low face validity?

Because test takers question whether the test really measures personality.

9
New cards

How is content validity defined and evaluated?

It is the extent to which a test covers the behavioral/conceptual domain to be measured; it is evaluated through item inspection by a panel of experts rather than statistical analysis.

10
New cards

What is a table of specification?

A blueprint of the test in terms of the number of items per difficulty, topic importance, or taxonomy.

11
New cards

What is construct underrepresentation?

The failure to capture important components of a construct (e.g., an English test containing only vocabulary items but no grammar items).

12
New cards

What is construct-irrelevant variance?

A situation that occurs when scores are influenced by factors irrelevant to the construct, such as test anxiety, reading speed, reading comprehension, or illness.

13
New cards

What are the three essential characteristics of a criterion?

  1. Relevant
  2. Valid and Reliable
  3. Uncontaminated
14
New cards

When does criterion contamination occur?

It occurs if the criterion is based on predictor measures, meaning the criterion used is a criterion of what is supposed to be the criterion.

15
New cards

What is concurrent validity?

A form of criterion-related validity established when test scores (predictor) are correlated with scores of a different measure (criterion) obtained at the same time to estimate present standing.

16
New cards

What is predictive validity?

A form of criterion-related validity demonstrated when scores on a test can predict future behavior or scores on another test taken in the future.

17
New cards

What is incremental validity?

The degree to which an additional predictor explains something about the criterion measure that is not explained by predictors already in use.

18
New cards

What is the correlation coefficient between a predictor and criterion called?

The validity coefficient (often calculated using Pearson rr).

19
New cards

What is a construct in psychological assessment?

An informed scientific idea developed or hypothesized to describe or explain a behavior; an unobservable, presupposed trait built by mental synthesis.

20
New cards

Which statistical tools can provide evidence that a test is homogeneous (measuring a single construct)?

Coefficient alpha, Spearman Rho (correlating item to item), and Pearson or point biserial (item-total correlation).

21
New cards

What is the method of contrasted groups in construct validation?

A method demonstrating construct validity by showing that test scores differ significantly between distinct groups, often tested using a tt-test.

22
New cards

What is convergent validity?

Construct validation demonstrated when a test correlates highly with other variables or tests with which it theoretically should correlate (e.g., extraversion correlating with sociability, or a life satisfaction test correlating with the "Satisfaction with Life Scale" by Ed Deiner, Ph.D.).

23
New cards

What is divergent (discriminant) validity?

Construct validation demonstrating that a test has low or negative correlation with measures of unrelated constructs from which it should differ (e.g., optimism negatively correlating with pessimism, or optimism having a weak correlation with gender identity).

24
New cards

What is the difference between Exploratory Factor Analysis (EFA) and Confirmatory Factor Analysis (CFA)?

Exploratory Factor Analysis is used for summarizing data, whereas Confirmatory Factor Analysis is used for generalization of factors.

25
New cards

What is cross-validation, and what is validity shrinkage?

Cross-validation is the revalidation of a test to a criterion based on another group different from the original group; validity shrinkage is the resulting decrease in validity after cross-validation.

26
New cards

What is the difference between co-validation and co-norming?

Co-validation is the validation of more than one test from the same group, whereas co-norming is the norming of more than one test from the same group.

27
New cards

What is the difference between severity error and leniency error?

Severity (strictness) error is an evaluation error due to a rater being overly critical, whereas leniency (generosity) error occurs when a rater is too forgiving and insufficiently critical.

28
New cards

What is central tendency error?

A rating error wherein the rater exhibits reluctance to rate at either extreme, causing all or most ratings to cluster in the middle of the rating continuum.

29
New cards

What is proximity error in rating scales?

A rating error committed due to the proximity or similarity of the traits being rated.

30
New cards

What is the halo effect?

A rating error wherein the rater views the object of the rating with extreme favor and bestows ratings that are inflated in a positive direction.

31
New cards

In impression management, what is the difference between acquiescence and non-acquiescence?

Acquiescence is the tendency to always agree with items, whereas non-acquiescence is the tendency to always disagree.

32
New cards

In impression management, what is the purpose behind faking-good versus faking-bad?

Faking-good is done to gain something, whereas faking-bad is done to avoid something.

33
New cards

In the context of psychological testing norms, who is a 'Norman'?

The test developer who will use the norms.

34
New cards

In developmental norms, what is the difference between basal age and ceiling age?

Basal age is the chronological age level at which maximum credits are achieved, while ceiling age represents the level where credits cease.

35
New cards

What is the formula for Intelligence Quotient (IQIQ) in developmental norms?

IQ=MACA×100IQ = \frac{\text{MA}}{\text{CA}} \times 100, where MA\text{MA} is Mental Age and CA\text{CA} is Chronological Age.

36
New cards

What are the basic premises regarding variables in test standardization?

The independent variable is the individual being tested, the dependent variable is their behavior (Behavior=person×situation\text{Behavior} = \text{person} \times \text{situation}), and standardization controls the situational factor so the person factor stands out.

37
New cards

What is the relationship between reliability and validity?

Reliability is a prerequisite for validity (a measurement cannot be valid unless it is reliable), but a measurement does not need to be valid to be considered reliable.

38
New cards

According to item-writing guidelines, why should 'double-barreled' items be avoided?

Because they convey two or more ideas at the same time, confusing the respondent.

39
New cards

Why should test developers consider mixing positively and negatively worded items?

To counteract the acquiescence response set, where respondents tend to agree with most items without fully understanding them.