A. PA Psychometric Properties and Principles

0.0(0)
Studied by 0 people
call kaiCall Kai
Locked
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/118

encourage image

There's no tags or description

Looks like no tags are added yet.

Last updated 10:04 AM on 7/30/26
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

119 Terms

1
New cards

Reliability

extent to which a method yields the same results under similar conditions

2
New cards

Reliability Coefficient

statistic that quantifies reliability, ranging from 0 to 1

3
New cards

True Score

measurement of a quantity if there were no measurement error at all

4
New cards

Carryover Effects

measurement processes that alter what is measured

5
New cards

Practice Effects

test itself provides an opportunity to learn and practice the ability being measured (increase of score due to test taker)

6
New cards

Test Sophistication

increase of score due to the test

7
New cards

Fatigue Effects

repeated testing reduces overall mental energy or motivation to perform on a test

8
New cards

Construct Score

person's standing on a theoretical variable independent of any particular measurement

9
New cards

Variance

useful in describing sources of test score variability; the standard deviation squared

10
New cards

True Variance

variance from true differences

11
New cards

Error Variance

variance from irrelevant, random sources; may increase or decrease a test score by varying amounts

12
New cards

Bias

degree to which a measure predictably overestimates or underestimates a quantity

13
New cards

Measurement Error

inherent uncertainty associated with any measurement, even after care has been taken to minimize preventable mistakes

14
New cards

Error

refers to the component of the observed test score that does not have to do with the test taker's ability

15
New cards

Random Error

source of error in measuring a targeted variable caused by unpredictable fluctuations and inconsistencies of other variables in the measurement process

16
New cards

Systematic Error

source of error in measuring a variable that is typically constant or proportionate to what is presumed to be the true value of the variable being measured

17
New cards

Item Sampling

refer to variation among items within a test as well as to variation among items between tests

18
New cards

Test

Retest Reliability

19
New cards

Coefficient of Stability

estimate of test

20
New cards

Parallel/Alternate Forms Reliability

evaluates the correlation between 2 different forms of a test

21
New cards

Coefficient of Equivalence

estimate of alternate

22
New cards

Parallel Forms Reliability

for each form of the test, the means and the variances of observed test scores are equal

23
New cards

Alternate Forms Reliability

different versions of a test that have been constructed so as to be parallel

24
New cards

Split

Half Reliability

25
New cards

Odd

Even Reliability

26
New cards

Spearman

Brown Formula

27
New cards

Average Proportional Distance

measure used to evaluate internal consistency of a test that focuses on the degree of differences that exists between item scores

28
New cards

Interrater Reliability

degree of agreement or consistency between two or more scorers with regard to a particular measure

29
New cards

Coefficient of Inter

Scorer Reliability

30
New cards

Dynamic

a trait, state, or ability presumed to be ever

31
New cards

Static

a trait, state, or ability presumed to be relatively unchanging (ex. intelligence)

32
New cards

Speed Tests

contains items of uniform level of difficulty and within a time limit

33
New cards

Power Tests

difficult items, time limit is long enough to allow test takers to attempt all items

34
New cards

Criterion

Referenced Tests

35
New cards

Classical Test Theory (CTT)

true score model of measurement

36
New cards

Domain Sampling Theory

estimate the extent to which specific sources of variation under defined conditions are contributing to the test scores

37
New cards

Generalizability Theory

based on the idea that a person's test scores vary from testing to testing because of the variables in the testing situations

38
New cards

Universe

test situation

39
New cards

Facets

number of items in the test, amount of review, and the purpose of test administration

40
New cards

Decision Study

developers examine the usefulness of test scores in helping the test user make decisions

41
New cards

Item Response Theory (IRT)

the probability that a person with X ability will be able to perform at a level of Y in a test

42
New cards

Difficulty

attribute of not being easily accomplished, solved, or comprehended

43
New cards

Discrimination

degree to which an item differentiates among people with higher or lower levels of the trait, ability or etc.

44
New cards

Dichotomous

can be answered with only one of two alternative responses

45
New cards

Polytomous

3 or more alternative responses

46
New cards

Standard Error of Measurement

provides a measure of the precision of an observed test score

47
New cards

Confidence Interval

a range or band of test scores that is likely to contain true scores

48
New cards

Standard Error of the Difference

can aid a test user in determining how large a difference should be before it is considered statistically significant

49
New cards

Standard Error of Estimate

refers to the standard error of the difference between the predicted and observed values

50
New cards

Validity

a judgment or estimate of how well a test measures what it supposed to measure

51
New cards

Inferences

logical result or deduction

52
New cards

Validation

the process of gathering and evaluating evidence about validity

53
New cards

Validation Studies

yield insights regarding a particular population of test takers as compared to the norming sample described in a test manual

54
New cards

Face Validity

test appears to measure what is meant to measure

55
New cards

Content Validity

concerned with the extent of how test items represent the behavior domain to be measured

56
New cards

Test Blueprint

a plan regarding the types of information to be covered by the items, the no. of items tapping each area of coverage, the organization of the items, and so forth

57
New cards

Underrepresentation

failure to capture components

58
New cards

Irrelevant Variance

other factors influenced the construct

59
New cards

Construct Validity

ability of the test to measure what it is meant to measure

60
New cards

Method of Contrasted Groups

demonstrate that scores on the test vary in a predictable way as a function of membership in a group

61
New cards

Factor Analysis

statistical tool used to analyze interrelationships among constructs

62
New cards

Factor Loading

conveys info about the extent to which the factor determine the test score or scores

63
New cards

Criterion

Related Validity

64
New cards

Criterion

external factor used as basis; has to be valid, reliable, and uncontaminated

65
New cards

Concurrent Validity

extent to which test scores may be used to estimate an individual's present standing on a criterion; criterion is readily available and administered at the same time

66
New cards

Predictive Validity

predict future behavior / scores on another test

67
New cards

Validity Coefficient

correlation coefficient that provides a measure of the relationship between test scores and scores on the criterion measure

68
New cards

Incremental Validity

degree to which an additional predictor explains something about the criterion measure that is not explained by predictors already in use

69
New cards

Criterion Contamination

occurs when the criterion measure includes aspects of performance that are not part of the job or when the measure is affected by "construct

70
New cards

Bias

factor inherent in a test that systematically prevents accurate, impartial measurement

71
New cards

Rating

numerical or verbal judgment that places a person or an attribute along a continuum identified by a scale of numerical or word descriptors known as Rating Scale

72
New cards

Rating Error

intentional or unintentional misuse of the scale

73
New cards

Leniency Error

rater is lenient in scoring (Generosity Error)

74
New cards

Severity Error

rater is strict in scoring

75
New cards

Central Tendency Error

rater's rating would tend to cluster in the middle of the rating scale

76
New cards

Halo Effect

tendency to give high score due to failure to discriminate among conceptually distinct and potentially independent aspects of a ratee's behavior

77
New cards

Fairness

the extent to which a test is used in an impartial, just, and equitable way

78
New cards

Norm

test performance data of a particular group of test takers that are designed for use as a reference when evaluating and interpreting individual test scores

79
New cards

Normative Sample

group of people whose performance on a particular test is analyzed for reference

80
New cards

Norming

process of deriving norms

81
New cards

Norman

person responsible for deriving norms

82
New cards

Percentile Norms

raw data from a test's standardization sample converted to percentile form

83
New cards

Percentile

expression of the percentage of people whose score on a test or measure falls below a particular raw score

84
New cards

Percentage Correct

refers to the number of items that were answered correctly multiplied by 100 and divided by the total number of items

85
New cards

Developmental Norms

norms developed on the basis of any trait, ability, skills, or other characteristic that is presumed to develop, deteriorate, or affected by stage of life

86
New cards

Age Norms

age

87
New cards

Grade Norms

indicate the average performance of different test takers in a given school grade; developed by administering the test to representative samples of children over a range of consecutive grade levels

88
New cards

National Norms

derived from a normative sample that was nationally representative of the population

89
New cards

National Anchor Norms

equivalency table for scores on the 2 tests which provides the tool for such comparison

90
New cards

Subgroup Norms

a normative sample can be segmented by any of the criteria initially used in selecting subjects for the sample

91
New cards

Local Norms

typically developed by test users; provide normative information with respect to the local population's performance on some test

92
New cards

Fixed Reference Group Scoring System

distribution of scores obtained on the test from one group of test takers (future reference group) is used as basis for the calculation of test scores for future administrations of the test (ex. SAT)

93
New cards

Standardization

process of establishing uniform procedures for administering, scoring, and interpreting a psychological test

94
New cards

Sample

portion of people deemed to be representative of the whole population

95
New cards

Purposive Sampling

Researcher intentionally selects people with specific characteristics

96
New cards

Snowball Sampling

Existing participants refer others

97
New cards

Convenience Sampling

Choosing participants who are easy to reach

98
New cards

Test Utility

usefulness or practical value of testing to improve efficiency

99
New cards

Expectancy Data

provide an indication that a test taker will score within some interval of scores on a criterion measure

100
New cards

Taylor

Russell Tables