CH05 Performance Measurement Notes
MODULE 5.1 Basic Concepts in Performance Measurement
• Performance measurement is ubiquitous
– Examples span classrooms, households, sports, politics, parenting, and work.
– Consultant Gerald Tannenbaum likens client enthusiasm for performance-measurement projects to a dental visit: necessary but not eagerly anticipated.
• Primary organizational uses for performance information
– Criterion data: validate selection tests by correlating test scores with objective/ratings correlations (Heneman, Bommer et al.).
– Employee development: identify strengths/weaknesses, craft improvement & training plans.
– Motivation & satisfaction: establish standards, evaluate success, give feedback → stronger motivation.
– Rewards: link pay/bonuses to performance (Rynes et al.).
– Transfer / Promotion / Layoff: choose employees for moves, advancement, or downsizing.
• Three data classes (Ch. 4 review)
– Objective: quantitative counts (sales, output, defects).
– Personnel: absences, tardiness, accidents.
– Judgmental: supervisory ratings.
– Correlations between classes are modest ( uncorrected; corrected) → measures are not interchangeable.
• Hands-on performance measurement
– Standardized work samples/simulations performed under controlled conditions.
– U.S. Army tank-crew simulation: radio use, internal comms, cannon positioning, weapon disassembly.
– Walk-through testing: employee verbally describes task execution while touring workplace.
• Electronic Performance Monitoring (EPM)
– ~78 % of firms (2001 AMA) monitor e-mail / web; 40 M workers (2000).
– Pros: objective, job-related, detailed logs; Cons: privacy invasion, stress, morale loss.
– Acceptance ↑ when tasks monitored are job-relevant, employees have voice, advance warning, ability to delay monitoring.
– Research: sparse field data; lab studies show skilled workers improve under monitoring; frequent EPM → higher task performance & OCBs in call centers (Bhave, 2014).
– Example: Ques Tec system evaluating MLB umpires – changed strike-zone calls and angered pitchers & umpires.
• Performance Management vs. Performance Appraisal (Banks & May)
– PM integrates definition, measurement, and communication linking behavior to strategic goals.
– Differences:
* Frequency: continuous vs. annual.
* Development: jointly by managers & employees vs. HR-imposed.
* Feedback: whenever needed vs. post-appraisal only.
* Roles: shared understanding vs. supervisor-dictated.
– Critical success factors: ongoing expectations dialogue, senior-leader modeling, manager feedback training.
• Key terms: objective performance measure, judgmental performance measure, hands-on performance measurement, walk-through testing, electronic performance monitoring, performance management.
MODULE 5.2 Performance Rating — Substance
• Process theories: Landy & Farr (1980) process model → observation → storage → retrieval → judgment.
– Later models emphasize raters’ cognition (memory, information processing).
• Levels of focus
– Overall performance ratings: administratively simple (akin to GPA) but psychologically complex.
* Negative info weighs more heavily (Ganzach).
* Overall ratings reflect Task Perf., OCB, CWB (Rotundo & Sackett) fairly consistently across jobs.
– Trait ratings: deprecated – traits (e.g., persistence) are predictors, not performance itself; legally weak.
– Task-based ratings: derived from job analysis; defensible.
– Critical-incident methods: lists of effective/ineffective behavioral examples (Flanagan) → basis for BARS.
– OCB & Adaptive Performance ratings: research supports adding these dimensions; should be validated via job analysis for legal defensibility.
• Structural characteristics of rating scales
Behavioral definition of dimension.
Defined meaning of response categories (anchored scale points).
Unambiguous interpretability for users.
– Only scale (f) in Figure 5.2 possessed all three.
• Rating formats
– Graphic Rating Scales: visual high→low continuum; effectiveness depends on good anchors; more points (e.g., 9) preferred by ratees (Bartol et al.).
– Checklists
* Weighted: hidden item weights summed (e.g., instructor behaviors w/ values 1-5).
* Forced-choice: rater selects best descriptors from sets balanced on social desirability → reduces leniency; meta-analysis shows ≥50 % validity gain (Bartram, 2007).
– Behavioral Anchored Rating Scales (BARS): behavioral descriptions at each scale point; time-consuming but high face validity.
– Behavioral Observation Scales (BOS): rate frequency of specific behaviors (Almost Never 1 → Almost Always 5); easier to develop; preferred by users.
– Employee Comparison Methods
* Simple ranking
* Paired comparison ; impractical as grows.
* CARS (Computer Adaptive Rating Scales): adaptive forced-choice using CAT logic → fewer comparisons.
• Take-aways
– Choose format based on feedback utility, legal defensibility, and resource constraints.
– Well-defined dimensions & behavioral anchors + trained raters = effective regardless of format.
• Key terms: task performance, OCB, CWB, duties, critical incidents, graphic rating scale, checklist, weighted checklist, forced-choice, BARS, BOS, employee comparison, simple ranking, paired comparison.
MODULE 5.3 Performance Rating — Process
• Rating sources
– Supervisors: most common; avoidance due to time, delivering negatives, fear of litigation.
– Peers: better for typical performance & OCB; issues when used for raises/promotion.
– Self: boosts justice perceptions; tendency toward inflation unless ratings are to be discussed.
– Subordinates: good for leadership behaviors; must be anonymous; useful for dev. not admin.
– Customers/Suppliers: capture service & interpersonal facets.
– 360-Degree feedback: integrates multiple sources for richer picture.
• Common rating distortions (errors/biases)
– Central tendency: clustering at midpoint.
– Leniency / Severity: systematically high or low.
– Halo: same score across dimensions.
• Rater-training approaches
– Administrative: how to use form.
– Psychometric: explain errors; can hurt accuracy (Bernardin & Pence).
– Frame-of-Reference (FOR): teach multidimensional performance, anchor meaning, practice/feedback → improves accuracy.
• Reliability & validity issues
– Low inter-rater reliability () due to each source seeing different behaviors – not inherently bad.
– Validity supported by job-analysis-based dimensions, quality anchors, and trained raters.
• Key terms: 360-degree feedback, rating errors, central tendency error, leniency error, severity error, halo error, psychometric training, frame-of-reference training.
MODULE 5.4 Social & Legal Context of Performance Evaluation
• Motivation & Politics in rating
– Raters may purposely distort to serve self, subordinate, or organizational goals (Banks & Murphy; Longnecker et al.).
– Stakeholder goals (Cleveland & Murphy):
* Rater: task, interpersonal, strategic, self-image.
* Ratee: info-gathering; info-dissemination.
* Organization: between-person (pay, promotion), within-person (development), systems-maintenance.
– Goal conflict arises when one system tries to satisfy divergent goals → possible solution: multiple systems or separate admin vs. developmental processes.
• Performance feedback principles
– Workers seek feedback to reduce uncertainty; prefer positives.
– Separate feedback sessions from salary discussions (up to 6 mo). Limit negative points per meeting – too many → defensiveness (Kay et al.).
– Avoid praise-criticism-praise sandwich – employees focus on the negative.
– Acceptance of negative feedback ↑ when:
* Supervisor observed enough behavior.
* Agreement on duties & standards.
* Focus on improvement plans.
– Destructive criticism (Baron): sarcastic, personal; elicits anger; best repaired via apology + explanation.
• 360-Degree feedback implementation guidelines (Harris)
Ensure rater anonymity (aggregate).
Jointly select raters.
Use for development only.
Train raters & feedback providers.
Follow-up coaching.
– Positive change most likely when feedback signals need for change, recipient is open, sets goals, and believes change is feasible (Smither et al.).
• Culture & appraisal
– High power-distance / collectivist cultures (e.g., Colombia, Venezuela) show:
* Supervisors most discrepant raters; peers least.
* Interpersonal behaviors valued > instrumental.
* Self-ratings modest; upward ratings lenient.
– Hofstede predictions: collectivists favor group appraisal; high PD resists upward feedback; masculinity & uncertainty avoidance shape style of feedback.
• Legal considerations
– Forced-distribution systems (e.g., Ford’s A/B/C, GE “rank & yank”): led to Ford settlement; risk of age, gender, race claims; may boost short-term task performance but hurt OCB & morale.
– Court focus is fairness, not psychometrics. Key favorable factors (Werner & Bolino):
* Job analysis basis.
* Written instructions & rater training.
* Appeal mechanisms.
* Multiple raters.
– Malos’ substantive & procedural safeguards (see Tables 5.6–5.7): objective, behavior-based, controllable criteria; standardized, documented, reviewable, appealable processes.
– Research shows negligible systemic bias against women, minorities, older workers when scales well-developed; stereotypes diminish with individuating info.
• Key terms: destructive criticism, forced-distribution rating system, policy capturing.