CH05 Performance Measurement Notes

MODULE 5.1 Basic Concepts in Performance Measurement

• Performance measurement is ubiquitous
– Examples span classrooms, households, sports, politics, parenting, and work.
– Consultant Gerald Tannenbaum likens client enthusiasm for performance-measurement projects to a dental visit: necessary but not eagerly anticipated.

• Primary organizational uses for performance information
– Criterion data: validate selection tests by correlating test scores with r≈+.20 to +.39r \approx +.20 \text{ to } +.39 objective/ratings correlations (Heneman, Bommer et al.).
– Employee development: identify strengths/weaknesses, craft improvement & training plans.
– Motivation & satisfaction: establish standards, evaluate success, give feedback → stronger motivation.
– Rewards: link pay/bonuses to performance (Rynes et al.).
– Transfer / Promotion / Layoff: choose employees for moves, advancement, or downsizing.

• Three data classes (Ch. 4 review)
– Objective: quantitative counts (sales, output, defects).
– Personnel: absences, tardiness, accidents.
– Judgmental: supervisory ratings.
– Correlations between classes are modest (r≈+.20r \approx +.20 uncorrected; r≈+.39r \approx +.39 corrected) → measures are not interchangeable.

• Hands-on performance measurement
– Standardized work samples/simulations performed under controlled conditions.
– U.S. Army tank-crew simulation: radio use, internal comms, cannon positioning, weapon disassembly.
– Walk-through testing: employee verbally describes task execution while touring workplace.

• Electronic Performance Monitoring (EPM)
– ~78 % of firms (2001 AMA) monitor e-mail / web; 40 M workers (2000).
– Pros: objective, job-related, detailed logs; Cons: privacy invasion, stress, morale loss.
– Acceptance ↑ when tasks monitored are job-relevant, employees have voice, advance warning, ability to delay monitoring.
– Research: sparse field data; lab studies show skilled workers improve under monitoring; frequent EPM → higher task performance & OCBs in call centers (Bhave, 2014).
– Example: Ques Tec system evaluating MLB umpires – changed strike-zone calls and angered pitchers & umpires.

• Performance Management vs. Performance Appraisal (Banks & May)
– PM integrates definition, measurement, and communication linking behavior to strategic goals.
– Differences:
* Frequency: continuous vs. annual.
* Development: jointly by managers & employees vs. HR-imposed.
* Feedback: whenever needed vs. post-appraisal only.
* Roles: shared understanding vs. supervisor-dictated.
– Critical success factors: ongoing expectations dialogue, senior-leader modeling, manager feedback training.

• Key terms: objective performance measure, judgmental performance measure, hands-on performance measurement, walk-through testing, electronic performance monitoring, performance management.


MODULE 5.2 Performance Rating — Substance

• Process theories: Landy & Farr (1980) process model → observation → storage → retrieval → judgment.
– Later models emphasize raters’ cognition (memory, information processing).

• Levels of focus
– Overall performance ratings: administratively simple (akin to GPA) but psychologically complex.
* Negative info weighs more heavily (Ganzach).
* Overall ratings reflect Task Perf., OCB, CWB (Rotundo & Sackett) fairly consistently across jobs.
– Trait ratings: deprecated – traits (e.g., persistence) are predictors, not performance itself; legally weak.
– Task-based ratings: derived from job analysis; defensible.
– Critical-incident methods: lists of effective/ineffective behavioral examples (Flanagan) → basis for BARS.
– OCB & Adaptive Performance ratings: research supports adding these dimensions; should be validated via job analysis for legal defensibility.

• Structural characteristics of rating scales

  1. Behavioral definition of dimension.

  2. Defined meaning of response categories (anchored scale points).

  3. Unambiguous interpretability for users.
    – Only scale (f) in Figure 5.2 possessed all three.

• Rating formats
– Graphic Rating Scales: visual high→low continuum; effectiveness depends on good anchors; more points (e.g., 9) preferred by ratees (Bartol et al.).
– Checklists
* Weighted: hidden item weights summed (e.g., instructor behaviors w/ values 1-5).
* Forced-choice: rater selects best descriptors from sets balanced on social desirability → reduces leniency; meta-analysis shows ≥50 % validity gain (Bartram, 2007).
– Behavioral Anchored Rating Scales (BARS): behavioral descriptions at each scale point; time-consuming but high face validity.
– Behavioral Observation Scales (BOS): rate frequency of specific behaviors (Almost Never 1 → Almost Always 5); easier to develop; preferred by users.
– Employee Comparison Methods
* Simple ranking
* Paired comparison Comparisons=n(n−1)2\text{Comparisons}=\frac{n(n-1)}{2}; impractical as nn grows.
* CARS (Computer Adaptive Rating Scales): adaptive forced-choice using CAT logic → fewer comparisons.

• Take-aways
– Choose format based on feedback utility, legal defensibility, and resource constraints.
– Well-defined dimensions & behavioral anchors + trained raters = effective regardless of format.

• Key terms: task performance, OCB, CWB, duties, critical incidents, graphic rating scale, checklist, weighted checklist, forced-choice, BARS, BOS, employee comparison, simple ranking, paired comparison.


MODULE 5.3 Performance Rating — Process

• Rating sources
– Supervisors: most common; avoidance due to time, delivering negatives, fear of litigation.
– Peers: better for typical performance & OCB; issues when used for raises/promotion.
– Self: boosts justice perceptions; tendency toward inflation unless ratings are to be discussed.
– Subordinates: good for leadership behaviors; must be anonymous; useful for dev. not admin.
– Customers/Suppliers: capture service & interpersonal facets.
– 360-Degree feedback: integrates multiple sources for richer picture.

• Common rating distortions (errors/biases)
– Central tendency: clustering at midpoint.
– Leniency / Severity: systematically high or low.
– Halo: same score across dimensions.

• Rater-training approaches
– Administrative: how to use form.
– Psychometric: explain errors; can hurt accuracy (Bernardin & Pence).
– Frame-of-Reference (FOR): teach multidimensional performance, anchor meaning, practice/feedback → improves accuracy.

• Reliability & validity issues
– Low inter-rater reliability (rxx′≈.50−.60r_{xx'}\approx .50-.60) due to each source seeing different behaviors – not inherently bad.
– Validity supported by job-analysis-based dimensions, quality anchors, and trained raters.

• Key terms: 360-degree feedback, rating errors, central tendency error, leniency error, severity error, halo error, psychometric training, frame-of-reference training.


MODULE 5.4 Social & Legal Context of Performance Evaluation

• Motivation & Politics in rating
– Raters may purposely distort to serve self, subordinate, or organizational goals (Banks & Murphy; Longnecker et al.).
– Stakeholder goals (Cleveland & Murphy):
* Rater: task, interpersonal, strategic, self-image.
* Ratee: info-gathering; info-dissemination.
* Organization: between-person (pay, promotion), within-person (development), systems-maintenance.
– Goal conflict arises when one system tries to satisfy divergent goals → possible solution: multiple systems or separate admin vs. developmental processes.

• Performance feedback principles
– Workers seek feedback to reduce uncertainty; prefer positives.
– Separate feedback sessions from salary discussions (up to 6 mo). Limit negative points per meeting – too many → defensiveness (Kay et al.).
– Avoid praise-criticism-praise sandwich – employees focus on the negative.
– Acceptance of negative feedback ↑ when:
* Supervisor observed enough behavior.
* Agreement on duties & standards.
* Focus on improvement plans.
– Destructive criticism (Baron): sarcastic, personal; elicits anger; best repaired via apology + explanation.

• 360-Degree feedback implementation guidelines (Harris)

  1. Ensure rater anonymity (aggregate).

  2. Jointly select raters.

  3. Use for development only.

  4. Train raters & feedback providers.

  5. Follow-up coaching.
    – Positive change most likely when feedback signals need for change, recipient is open, sets goals, and believes change is feasible (Smither et al.).

• Culture & appraisal
– High power-distance / collectivist cultures (e.g., Colombia, Venezuela) show:
* Supervisors most discrepant raters; peers least.
* Interpersonal behaviors valued > instrumental.
* Self-ratings modest; upward ratings lenient.
– Hofstede predictions: collectivists favor group appraisal; high PD resists upward feedback; masculinity & uncertainty avoidance shape style of feedback.

• Legal considerations
– Forced-distribution systems (e.g., Ford’s A/B/C, GE “rank & yank”): led to $10.5 M\$10.5\text{ M} Ford settlement; risk of age, gender, race claims; may boost short-term task performance but hurt OCB & morale.
– Court focus is fairness, not psychometrics. Key favorable factors (Werner & Bolino):
* Job analysis basis.
* Written instructions & rater training.
* Appeal mechanisms.
* Multiple raters.
– Malos’ substantive & procedural safeguards (see Tables 5.6–5.7): objective, behavior-based, controllable criteria; standardized, documented, reviewable, appealable processes.
– Research shows negligible systemic bias against women, minorities, older workers when scales well-developed; stereotypes diminish with individuating info.

• Key terms: destructive criticism, forced-distribution rating system, policy capturing.