1/38
Flashcards covering psychological testing and assessment topics including types of reliability, face, content, criterion-related, and construct validity, rating errors, test standardization, norms, and item-writing guidelines.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
When is parallel forms reliability appropriate to use?
It is appropriate for tests that have two forms.
When is test-retest reliability appropriate to use?
It is appropriate for tests that are designed to be administered to an individual more than once.
Which reliability estimate is appropriate for tests demonstrating factorial purity?
Cronbach's coefficient alpha.
When is split-half reliability appropriate to use?
It is appropriate for tests with items carefully ordered according to difficulty.
When should inter-rater reliability be utilized?
It is used for tests that involve some degree of subjective scoring.
What is the Kuder-Richardson 20 (KR20) formula used for?
It is the statistic used for calculating the reliability of a test in which items are dichotomous or scored as 0 or 1 (forced-choice items).
What is face validity?
The simplest and least stringent form of validity, referring to whether a test looks valid at face value or superficial appearance to test users, examiners, and examinees.
Why do inkblot tests typically demonstrate low face validity?
Because test takers question whether the test really measures personality.
How is content validity defined and evaluated?
It is the extent to which a test covers the behavioral/conceptual domain to be measured; it is evaluated through item inspection by a panel of experts rather than statistical analysis.
What is a table of specification?
A blueprint of the test in terms of the number of items per difficulty, topic importance, or taxonomy.
What is construct underrepresentation?
The failure to capture important components of a construct (e.g., an English test containing only vocabulary items but no grammar items).
What is construct-irrelevant variance?
A situation that occurs when scores are influenced by factors irrelevant to the construct, such as test anxiety, reading speed, reading comprehension, or illness.
What are the three essential characteristics of a criterion?
When does criterion contamination occur?
It occurs if the criterion is based on predictor measures, meaning the criterion used is a criterion of what is supposed to be the criterion.
What is concurrent validity?
A form of criterion-related validity established when test scores (predictor) are correlated with scores of a different measure (criterion) obtained at the same time to estimate present standing.
What is predictive validity?
A form of criterion-related validity demonstrated when scores on a test can predict future behavior or scores on another test taken in the future.
What is incremental validity?
The degree to which an additional predictor explains something about the criterion measure that is not explained by predictors already in use.
What is the correlation coefficient between a predictor and criterion called?
The validity coefficient (often calculated using Pearson r).
What is a construct in psychological assessment?
An informed scientific idea developed or hypothesized to describe or explain a behavior; an unobservable, presupposed trait built by mental synthesis.
Which statistical tools can provide evidence that a test is homogeneous (measuring a single construct)?
Coefficient alpha, Spearman Rho (correlating item to item), and Pearson or point biserial (item-total correlation).
What is the method of contrasted groups in construct validation?
A method demonstrating construct validity by showing that test scores differ significantly between distinct groups, often tested using a t-test.
What is convergent validity?
Construct validation demonstrated when a test correlates highly with other variables or tests with which it theoretically should correlate (e.g., extraversion correlating with sociability, or a life satisfaction test correlating with the "Satisfaction with Life Scale" by Ed Deiner, Ph.D.).
What is divergent (discriminant) validity?
Construct validation demonstrating that a test has low or negative correlation with measures of unrelated constructs from which it should differ (e.g., optimism negatively correlating with pessimism, or optimism having a weak correlation with gender identity).
What is the difference between Exploratory Factor Analysis (EFA) and Confirmatory Factor Analysis (CFA)?
Exploratory Factor Analysis is used for summarizing data, whereas Confirmatory Factor Analysis is used for generalization of factors.
What is cross-validation, and what is validity shrinkage?
Cross-validation is the revalidation of a test to a criterion based on another group different from the original group; validity shrinkage is the resulting decrease in validity after cross-validation.
What is the difference between co-validation and co-norming?
Co-validation is the validation of more than one test from the same group, whereas co-norming is the norming of more than one test from the same group.
What is the difference between severity error and leniency error?
Severity (strictness) error is an evaluation error due to a rater being overly critical, whereas leniency (generosity) error occurs when a rater is too forgiving and insufficiently critical.
What is central tendency error?
A rating error wherein the rater exhibits reluctance to rate at either extreme, causing all or most ratings to cluster in the middle of the rating continuum.
What is proximity error in rating scales?
A rating error committed due to the proximity or similarity of the traits being rated.
What is the halo effect?
A rating error wherein the rater views the object of the rating with extreme favor and bestows ratings that are inflated in a positive direction.
In impression management, what is the difference between acquiescence and non-acquiescence?
Acquiescence is the tendency to always agree with items, whereas non-acquiescence is the tendency to always disagree.
In impression management, what is the purpose behind faking-good versus faking-bad?
Faking-good is done to gain something, whereas faking-bad is done to avoid something.
In the context of psychological testing norms, who is a 'Norman'?
The test developer who will use the norms.
In developmental norms, what is the difference between basal age and ceiling age?
Basal age is the chronological age level at which maximum credits are achieved, while ceiling age represents the level where credits cease.
What is the formula for Intelligence Quotient (IQ) in developmental norms?
IQ=CAMA×100, where MA is Mental Age and CA is Chronological Age.
What are the basic premises regarding variables in test standardization?
The independent variable is the individual being tested, the dependent variable is their behavior (Behavior=person×situation), and standardization controls the situational factor so the person factor stands out.
What is the relationship between reliability and validity?
Reliability is a prerequisite for validity (a measurement cannot be valid unless it is reliable), but a measurement does not need to be valid to be considered reliable.
According to item-writing guidelines, why should 'double-barreled' items be avoided?
Because they convey two or more ideas at the same time, confusing the respondent.
Why should test developers consider mixing positively and negatively worded items?
To counteract the acquiescence response set, where respondents tend to agree with most items without fully understanding them.