1/26
Practice flashcards covering core concepts, assessment principles, taxonomies, test design, validity, reliability, item analysis, and statistical distributions from the Assessment and Evaluation of Learning lecture.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
What is the difference between measurement and evaluation in education?
Measurement is the process of quantifying the degree to which someone or something possesses a given trait, while evaluation is the systematic collection and analysis of data to make a judgment about the desirability of changes in students.
What is the distinction between immediate outcomes and deferred outcomes in Outcomes-Based Education (OBE)?
Immediate outcomes are competencies and skills acquired upon completion of a lesson, subject, grade, or program (such as communication or problem-solving skills), whereas deferred outcomes refer to the long-term application of skills in professional and workplace practice (such as success in professional career planning).
How does William Spady define outcomes in transformational Outcomes-Based Education (OBE)?
Spady defines outcomes as deferred, long-term, cross-curricular outcomes that relate directly to a student's future life roles, such as being a productive worker, responsible citizen, or parent.
What are the four principles of Outcomes-Based Education (OBE)?
What concept describes aligning assessment tasks and specific evaluation criteria directly to intended learning outcomes?
Constructive Alignment
What are the six levels of the Cognitive Domain in Bloom's Taxonomy?
What are the five levels of the Affective Domain in Bloom's Taxonomy?
What are the five levels of the Psychomotor Domain in Bloom's Taxonomy?
What three systems comprise Marzano's New Taxonomy model of thinking skills?
What are the differences among Assessment FOR Learning, Assessment OF Learning, and Assessment AS Learning?
Assessment FOR learning is used before or during instruction to guide teaching (e.g., placement, formative, diagnostic); Assessment OF learning evaluates achievement at the end of instruction (e.g., summative); Assessment AS learning focuses on training teachers and learners on how to assess.
How do Traditional Assessment and Authentic Assessment compare regarding Action, Setting, Focus, and Outcome?
In Traditional Assessment, Action is selecting a response, Setting is contrived, Focus is teacher-structured, and Outcome is indirect evidence. In Authentic Assessment, Action is performing a task, Setting is simulation, Focus is student-structured, and Outcome is direct evidence.
What is the difference between Demonstration-type and Creation-type performance-based tasks?
Demonstration-type tasks require no physical product (such as cooking demonstrations or entertaining tourists), whereas Creation-type tasks require tangible products (such as project plans or research papers).
What do the components of the GRASPS acronym stand for in setting criteria for performance tasks?
G - Goal; R - Role; A - Audience; S - Situation; P - Product; S - Standards and Criteria
What sequence of steps is involved in the portfolio development process depicted in this diagram?
According to this Venn diagram, how do Checklists, Rating Scales, and Rubrics differ in their purpose?
Checklists show observed traits of a work or performance, Rating Scales show the degree of quality of a work or performance, and Rubrics combine both traits as modified checklists and rating scales.
What is the key structural difference between a Holistic Rubric and an Analytic Rubric?
A Holistic Rubric gives a single rating describing the overall quality of the entire performance or product, whereas an Analytic Rubric rates identified dimensions or criteria independently to provide a detailed assessment.
How do Norm-Referenced Tests and Criterion-Referenced Tests differ in score interpretation?
Norm-Referenced Tests interpret results by comparing one student with other students under competition for limited high scores, whereas Criterion-Referenced Tests interpret results by comparing a student against a fixed set of criteria with no competition.
What distinguishes a Power Test from a Speed Test?
A Power Test consists of items with increasing levels of difficulty taken with ample time to measure difficulty capacity, whereas a Speed Test consists of items with uniform difficulty taken within a time limit to measure speed and accuracy.
What are the four phases in the Development and Validation of an Assessment Instrument?
Phase I: Planning Stage; Phase II: Item Writing Stage; Phase III: Try Out Stage; Phase IV: Evaluation Stage
How are Difficulty Index ranges interpreted in item analysis?
0.00−0.20: Very difficult item; 0.21−0.40: Difficult item; 0.41−0.60: Moderately difficult item; 0.61−0.80: Easy item; 0.81 and above: Very easy item
How are Discrimination Index ranges interpreted in item analysis?
−1.00−−0.60: Questionable item; −0.59−−0.20: Not discriminating item; −0.21−0.20: Moderately discriminating item; 0.21−0.60: Discriminating item; 0.61−1.00: Very discriminating item
What is the difference between Concurrent Validity and Predictive Validity?
Concurrent Validity describes present status by correlating test scores with external measures administered concurrently, while Predictive Validity describes future performance by correlating test scores with measures administered after a longer time interval.
What is the difference between Convergent Validity and Divergent Validity?
Convergent Validity is established when an instrument correlates with tests measuring similar traits, while Divergent Validity is established when an instrument describes only the intended trait and does not correlate with tests measuring distinct traits.
Which method for measuring internal consistency reliability involves scoring odd- and even-numbered items separately?
Split Half method
What two types of skewed distributions are shown in this diagram, and how do their scores distribute?
What three types of kurtosis are shown in this diagram, and what are their respective K values?
A. Mesokurtic (Normal) where K=0; B. Leptokurtic (peaked/steeper) where K>0; C. Platykurtic (flatter) where K<0
What are the definitions of z-score, stanine, and t-score?
z-score represents standard deviations above or below the mean; Stanine divides a normal distribution into 9 segments numbered 1 through 9; t-score places a score in a normal distribution with a mean of 50 and a standard deviation of 10.