1/42
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Characteristis of assessment measures
Specific domains of functioning are sampled and can be administered to individuals, groups or
organisations.
Measures are administered under standardised conditions.
Systematic methods are applied to score or evaluate assessment protocols.
Guidelines are available to understand and interpret the results of an assessment measures. Such
guidelines make provision for the comparison of an individual’s performance to that of an appropriate
norm group or criterion.
Assessment measures should be supported by evidence that they are valid and reliable for the intended
purpose.
This evidence is provided in a test manual.
Categories of measurement levels
Categorical data (discrete or distinct categories)
Nominal
Ordinal
Continuous data (measured on a continuum)
Interval
ratio
Nominal
Not ranked, categories are the same, represented by name
group membership
i.e. gender, country,
statistics: frequencey, mode, correlation coefficient
Ordinal
More refined, numbers assigned to objects, rank ordered
determination of rank
social class, job level, psy. construct
above, median, percentile rank
Interval
Equal numerical differences
interval between 96 - 98 is same as 98 - 100
IQ scores, temperature
above and median, SD, Pearson correlation coefficient
ratio
most refined, equal differences and absolute zero
determination of equality of rations
score of 20 is 20x greater than 1
distance ruler
stats: Above and coefficient of variation
Measurement errors - Random sampling errors
Differ due to chance factors
Larger samples yield more accurate
results
Reliability is affected
Variability of errors of measurement
that function in a random manner
Guessing
Complex language
Test anxiety
Disruptions
Measurement errors - systematic errors
The execution of the research design/ application of
the measure
Measurement error/ bias that may negatively affect
scores obtained from norm groups
Minimised if principles of sound scale construction
are adhered to
Can disadvantage one group of test-takers
Incorrect calibration of a scale to obtain weight of
participants
Timing errors, scoring, administration
Response bias
Constant errors of measurement that occur when all
test scores are extremely high or extremely low
Test conditions, cultural groups
ACCEPTABLE ERRORS OF MEASUREMENT
Standard error of measurement (lower = more precise)
confidence levels (express uncertainty)
reliability coefficient
categories of statistical meausures
Measures of central tendency
measures of variability
measures of association
Measures of variability
range
variance (how scores deviate from mean)
SD (Square root of variance, low scores = low variability) / together = similar scores; scattered = scores range
correlation coefficient
r - indicates direction and strength of a relationship between 2 variables
Norms
types
criterion referenced
norm referenced
Norm subgroups
Applicant (similarities between norm groups)
Incumbent (comparative aspects, e.g. gender, age)
co-norming of measures
Assessment practitioner wants to compare scores obtained on different but related measures to test for possible learning or memory effects.
Process where two or more are related, but different measures are administered and standardised as a unit on the same norm group.
Standardisation
Systematic and scientific development of tests
Construction of adequate norms to compare individual scores
Test items, administration, scoring and interpretation remain consistent during each administration
Norm tables (how was it selected, demograhics, date)
construct referenced norms
criterion-referenced norms
reliability
Consistency implies certain amount of measurement error
Classical test theory: Observed score = true score + error
Measure in same way over time
Dependability
Factors affecting reliability
intrinsic (in test): length of test, Homogeneity of items, difficulty values of items, discrimination vales, test instructions, item selection, reliability of scorer
extrinsic (participant): group variability, guessing and chance factors, environmental conditions, momentary fluctuations, response bias
Random / Systematic error
Respondent error
non-response error
self-selection bias
response bias (Extremity bias, Stringency or leniency bias [rater], Acquiescience buas [Agreement with everything], Halo effect [?], Social desirability, purpose falsification [purposefully misrepresent facts, conceal personal information, respond only randomly], unconscious misrepresentation [may not be able to recall instructions, unintentionally incorrect])
Systematic error
or ADMINISTRATIVE ERROR
instructions
assessment conditions
interpretation of instructions
scoring and timing
validity
What the test measures and how well it does so
accuracy of the measure
the stability and predictions made for future results
unitary validity: all validity facets
Validity coefficient: correlation between predictor and criterion
Factors affecting validity
Reliability
Differential impact of subgroups
validity coefficinet must be constant for subgroups that differ - biographic factors may have differential impact on validity coefficient
Sample homogeneity
Similar scores = restriction of range
wider range = higher validity coefficient
Linear relationship between predictor and criterion
Relationship must be linear
Criterion contamination
Affect magnitude, may lose validity
Moderator variables
gender, age, SES - may lead to discrimination or bias
Developing a psy. measure
Planning phase
item developmen(“Item writing phase”)
Assembling and pre-testing the experimental version of the measure
Item analysis
revising and standardising the final version of the measure (Standardisation phase)
technical evaluation and establishing norms
publishing and ongoing refinement
Developing a psy measure - Phase 1
Planning
Blueprint
— Establish test development team
Aim? Purpose?
target population & nature of the measure
Operationalise construct (rational and analytic)
define content
Criterion keying approach to ensure theoretically grounded measure
develop test plan = specifications to guide development
determine test format (objective - MCQ, subjective - open ended, sentence completion)
determine test mode (pencil-paper, computer-based, performance-based)
other aspects to consider
unintentional biases (multicultural environment)
response set (i.e. agreeing to all items)
language
length (impact time & quality of responses)
developing a measure - Phase 2
Item development
develop or source items
take content, domain, cultural and linguistic aspects into account
clear wording, appropriate vocabulary
one theme per item
mix up MCQ questions
ensure colour when developing measure for children
develop more items than needed
Review items
expert review to judge appropriateness
pilot testing
revise and re-write items as needed
developing a measure - Phase 3
Assesmbling and pre-testing experimental version
arrange items
finalise length
specify answer mechanisms & protocols
develop administration instructions
pre-test experimental version - large population
developing a measure - phase 4
item analysis
examine each item to see if it serves the purpose of the measure
item difficulty values (P value = no of correct items / total population —> high p value = easy test across different dimensions of the measure)
discrimination values
after this investigate bias
IRT - most accurate, does not depend on ability level of test-taker
investigate DIF
3 criteria !
content validity
difficulty and discrimination
bias and fairness
—> Identify items for final pool
Developing a measure - phase 5
Revising and standardising the final version of the measure
revise items ans test
select items for final version
refine instructions and scoring procedures
administer final version to a large target population to establish reliability, validity and standardisation
developing a measure - phase 6
technical evaluation and establishing norms
evaluate psychometric properties
devise norm tables, set cut off scores, performance standards
compare test scores to one another to determine the meaning of the scores
developing a measure - phase 7
Publishing and ongoing refinement
compile test manual
submit the measure for classification
publish and market
revise, define and update continuously
Test manual must include
Purpose of test
Who the test is for
Length of the test
Reading grade level required
Administration and scoring instructions
Outline of development process
Types of reliability + validity
Norms and tables
evaluating a measure
How long ago was it developed
Quality of manual contents
Clarity on instructions
Cultural appropriateness
Investigations for multicultural use
Adequacy of psychometric properties
Nature of norms
Cross cultural comparison = equivalence
determining the number of items in a test
content coverage
reliability - do add. items contribute?
validity - How well is the construct measured?
test-taker fatigue
administrative constraints
Item difficulty and discrimination (items should vary in difficulty)
item quality (each item must be clear, relevant, and bias free)
factor structure (FA - identify optimal no. of items per factor)
cultural sensitivity (different cultural styles?)
item purpose (Purpose can influence optimal test length)
item randomisation (minimise order effects & increase reliability)
Bias
Systematic error in measurement process that
leads to inaccurate or unfair assessment of
individuals
Fairness
Equitable treatment of individuals
A fair test should provide accurate + unbiased
representation of the individual
construct bias
Psychological construct being measured is not equivalent across different groups
Conceptualisation
Cultural differences
method bias
Procedures used in testing process systematically affects the response/scores of a certain group
Test administration
Scoring procedures
item bias
dif
qualitative review
stats
response bias
systematic way of responding to items that leads to inaccuracies
response styles
cultural bias
characteristics of a good test
validity
reliability
limitations
management
review
standardisation
the assessment process
Intake interview
multidimensional (various sources of info.)
informed conset
limtations to confidentiality
Assessment battery
tailored to individual needs
gather info from sources and cluster info
culturally appropriate test battery
assessment day
build rapport, ensure consent, setting
after
score, interpret, report
cluster info and results
provide appropriate recommendations tailored to client abilities
Provide feedback
understanding
paymnt
Assessment process - prep
Prep
Intake interview (gain background information and informed consent)
Select measures according to client needs
check materials
be familiar with instructions
satisfactory assessment condtions
personal circumstance test taker
planning test sequence
address linguistic issues
address test sophistication
Assessment process - administration
rapport
test anxiety
provide instructions
time limits
manage irregularities
record behaviour and clinical observations
mental disabilities?
assessment process - aftermath
collect & secure materials
record process notes, score and interpret
contact, provide feedback, recommendations
reasons to adapt tests
enhance fairness
reduce cost and save time
comparative studies
globalisation
consider when adapting tests…
construct equivalence
language use
psychometric properties
nature of tasks