1/118
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Reliability
extent to which a method yields the same results under similar conditions
Reliability Coefficient
statistic that quantifies reliability, ranging from 0 to 1
True Score
measurement of a quantity if there were no measurement error at all
Carryover Effects
measurement processes that alter what is measured
Practice Effects
test itself provides an opportunity to learn and practice the ability being measured (increase of score due to test taker)
Test Sophistication
increase of score due to the test
Fatigue Effects
repeated testing reduces overall mental energy or motivation to perform on a test
Construct Score
person's standing on a theoretical variable independent of any particular measurement
Variance
useful in describing sources of test score variability; the standard deviation squared
True Variance
variance from true differences
Error Variance
variance from irrelevant, random sources; may increase or decrease a test score by varying amounts
Bias
degree to which a measure predictably overestimates or underestimates a quantity
Measurement Error
inherent uncertainty associated with any measurement, even after care has been taken to minimize preventable mistakes
Error
refers to the component of the observed test score that does not have to do with the test taker's ability
Random Error
source of error in measuring a targeted variable caused by unpredictable fluctuations and inconsistencies of other variables in the measurement process
Systematic Error
source of error in measuring a variable that is typically constant or proportionate to what is presumed to be the true value of the variable being measured
Item Sampling
refer to variation among items within a test as well as to variation among items between tests
Test
Retest Reliability
Coefficient of Stability
estimate of test
Parallel/Alternate Forms Reliability
evaluates the correlation between 2 different forms of a test
Coefficient of Equivalence
estimate of alternate
Parallel Forms Reliability
for each form of the test, the means and the variances of observed test scores are equal
Alternate Forms Reliability
different versions of a test that have been constructed so as to be parallel
Split
Half Reliability
Odd
Even Reliability
Spearman
Brown Formula
Average Proportional Distance
measure used to evaluate internal consistency of a test that focuses on the degree of differences that exists between item scores
Interrater Reliability
degree of agreement or consistency between two or more scorers with regard to a particular measure
Coefficient of Inter
Scorer Reliability
Dynamic
a trait, state, or ability presumed to be ever
Static
a trait, state, or ability presumed to be relatively unchanging (ex. intelligence)
Speed Tests
contains items of uniform level of difficulty and within a time limit
Power Tests
difficult items, time limit is long enough to allow test takers to attempt all items
Criterion
Referenced Tests
Classical Test Theory (CTT)
true score model of measurement
Domain Sampling Theory
estimate the extent to which specific sources of variation under defined conditions are contributing to the test scores
Generalizability Theory
based on the idea that a person's test scores vary from testing to testing because of the variables in the testing situations
Universe
test situation
Facets
number of items in the test, amount of review, and the purpose of test administration
Decision Study
developers examine the usefulness of test scores in helping the test user make decisions
Item Response Theory (IRT)
the probability that a person with X ability will be able to perform at a level of Y in a test
Difficulty
attribute of not being easily accomplished, solved, or comprehended
Discrimination
degree to which an item differentiates among people with higher or lower levels of the trait, ability or etc.
Dichotomous
can be answered with only one of two alternative responses
Polytomous
3 or more alternative responses
Standard Error of Measurement
provides a measure of the precision of an observed test score
Confidence Interval
a range or band of test scores that is likely to contain true scores
Standard Error of the Difference
can aid a test user in determining how large a difference should be before it is considered statistically significant
Standard Error of Estimate
refers to the standard error of the difference between the predicted and observed values
Validity
a judgment or estimate of how well a test measures what it supposed to measure
Inferences
logical result or deduction
Validation
the process of gathering and evaluating evidence about validity
Validation Studies
yield insights regarding a particular population of test takers as compared to the norming sample described in a test manual
Face Validity
test appears to measure what is meant to measure
Content Validity
concerned with the extent of how test items represent the behavior domain to be measured
Test Blueprint
a plan regarding the types of information to be covered by the items, the no. of items tapping each area of coverage, the organization of the items, and so forth
Underrepresentation
failure to capture components
Irrelevant Variance
other factors influenced the construct
Construct Validity
ability of the test to measure what it is meant to measure
Method of Contrasted Groups
demonstrate that scores on the test vary in a predictable way as a function of membership in a group
Factor Analysis
statistical tool used to analyze interrelationships among constructs
Factor Loading
conveys info about the extent to which the factor determine the test score or scores
Criterion
Related Validity
Criterion
external factor used as basis; has to be valid, reliable, and uncontaminated
Concurrent Validity
extent to which test scores may be used to estimate an individual's present standing on a criterion; criterion is readily available and administered at the same time
Predictive Validity
predict future behavior / scores on another test
Validity Coefficient
correlation coefficient that provides a measure of the relationship between test scores and scores on the criterion measure
Incremental Validity
degree to which an additional predictor explains something about the criterion measure that is not explained by predictors already in use
Criterion Contamination
occurs when the criterion measure includes aspects of performance that are not part of the job or when the measure is affected by "construct
Bias
factor inherent in a test that systematically prevents accurate, impartial measurement
Rating
numerical or verbal judgment that places a person or an attribute along a continuum identified by a scale of numerical or word descriptors known as Rating Scale
Rating Error
intentional or unintentional misuse of the scale
Leniency Error
rater is lenient in scoring (Generosity Error)
Severity Error
rater is strict in scoring
Central Tendency Error
rater's rating would tend to cluster in the middle of the rating scale
Halo Effect
tendency to give high score due to failure to discriminate among conceptually distinct and potentially independent aspects of a ratee's behavior
Fairness
the extent to which a test is used in an impartial, just, and equitable way
Norm
test performance data of a particular group of test takers that are designed for use as a reference when evaluating and interpreting individual test scores
Normative Sample
group of people whose performance on a particular test is analyzed for reference
Norming
process of deriving norms
Norman
person responsible for deriving norms
Percentile Norms
raw data from a test's standardization sample converted to percentile form
Percentile
expression of the percentage of people whose score on a test or measure falls below a particular raw score
Percentage Correct
refers to the number of items that were answered correctly multiplied by 100 and divided by the total number of items
Developmental Norms
norms developed on the basis of any trait, ability, skills, or other characteristic that is presumed to develop, deteriorate, or affected by stage of life
Age Norms
age
Grade Norms
indicate the average performance of different test takers in a given school grade; developed by administering the test to representative samples of children over a range of consecutive grade levels
National Norms
derived from a normative sample that was nationally representative of the population
National Anchor Norms
equivalency table for scores on the 2 tests which provides the tool for such comparison
Subgroup Norms
a normative sample can be segmented by any of the criteria initially used in selecting subjects for the sample
Local Norms
typically developed by test users; provide normative information with respect to the local population's performance on some test
Fixed Reference Group Scoring System
distribution of scores obtained on the test from one group of test takers (future reference group) is used as basis for the calculation of test scores for future administrations of the test (ex. SAT)
Standardization
process of establishing uniform procedures for administering, scoring, and interpreting a psychological test
Sample
portion of people deemed to be representative of the whole population
Purposive Sampling
Researcher intentionally selects people with specific characteristics
Snowball Sampling
Existing participants refer others
Convenience Sampling
Choosing participants who are easy to reach
Test Utility
usefulness or practical value of testing to improve efficiency
Expectancy Data
provide an indication that a test taker will score within some interval of scores on a criterion measure
Taylor
Russell Tables