Trust Development & Repair in AI-Assisted Decisions – Comprehensive Study Notes
Background & Key Concepts
- Human–AI Collaboration
- Seeks to combine complementary strengths of humans (contextual reasoning, empathy, domain insight) and AI (speed, pattern-recognition, memorisation).
- Central challenge: users must decide when and how much to trust AI advice.
- Complementary vs. Overlapping Expertise
- Overlapping: both human and AI competent → humans can directly judge AI accuracy.
- Complementary: each excels at different facets → humans cannot always assess AI accuracy.
- Appropriate Trust / Trust Calibration
- Balance between over-trust (unwarranted reliance) and under-trust (undue scepticism).
- Calibration aims to align subjective trust with objective AI capability.
- Trust Repair Strategies (TRSs) – interventions adopted from Social Psychology & Human–Robot Interaction (HRI) to rebuild trust after violations.
- Apology (expression of regret)
- Denial (reject culpability)
- Promise (commit to do better)
- Model Update (inform user of technical improvement)
- Dispositional Trust (TiA-PtT) – a trait-level propensity to trust automation; moderates initial trust.
Research Questions & Contributions
- RQ1: How does perceived AI accuracy in High Human-Expertise (HHE) tasks influence users’ trust during Low Human-Expertise (LHE) tasks?
- RQ2: In complementary expertise settings, how does trust recover when AI accuracy improves, with and without explicit TRSs?
- Key contributions
- Empirical evidence that users extrapolate AI accuracy in HHE tasks to calibrate trust in LHE tasks.
- Comparative efficacy of four TRSs; Model Update > Apology > Promise ≈ No-Repair > Denial.
- Insight that trust recovery depends on anthropomorphism, perceived regret/deceit, causal attribution, and nature of improvement (behavioural vs technical).
- Dual-task (Shapes & Animals) replication enhances ecological validity.
- Trust as an attitude under uncertainty (Lee & See, 2004).
- Accuracy → trust correlation (Yu et al., 2016; Yin et al., 2019).
- Initial impressions / early errors have long-term effects (Tolmeijer et al., 2021).
- HRI literature on trust repair: mixed findings for apology, promise, denial; scarcity of work in non-robotic AI decision aids.
Trust Repair Strategies (TRS) – Details
- Apology
- Operates emotionally; seeks to restore social expectations.
- Text used: “I’m sorry … I hope you can trust me again.”
- Denial
- Shifts blame; may appear deceptive if evidence contradicts statement.
- Promise
- Behavioural commitment; effectiveness relies on perceived agency to fulfil promise.
- Model Update (novel)
- Communicates algorithmic upgrade; implicitly supplies causal attribution (flawed model → fixed).
- All TRS messages followed template: acknowledgement of mistrust → core element → hope for restored trust.
Methodology
- Design: Survey-based, between-subjects; 5 TRS conditions (including baseline No-Repair) × 2 tasks.
- Participants: N=300 Prolific users (eng ≥ 98%); 150 per task, 30 per TRS.
- Mean ages: Shapes xˉ=35 (SD =13.01); Animals xˉ=34.2 (SD =11.85).
- Power Analysis: f2=0.25, α=0.05, power =0.8 ⇒ minimum n=135 per task (G*Power 3).
- Phases & AI Accuracy
- Phase 1 – High accuracy =80% (1 error at trial 7) → build trust.
- Phase 2 – Low accuracy =20% (multiple errors) → violate trust.
- Phase 3 – High accuracy =80% again; TRS shown before start.
- Trial Types (10 per phase; 30 total)
- HHE (Familiar) = human knows answer.
- LHE (Unfamiliar) = human does not know; must rely on AI.
- Delay: 3s before AI advice to encourage independent thought.
- Model-Update scenario: extra 6s wait to simulate retraining.
Experimental Tasks & Stimuli
- Shapes Task
- Familiar: Circle, Rectangle, Triangle (5 visual variants each).
- Unfamiliar (fabricated): “Scleratice”, “Tenectus”, “Pyrangle”; randomised fill/border patterns.
- Animals Task
- Familiar: 15 common species (Cat, Dog, Horse, etc.).
- Unfamiliar: 15 obscure species (e.g., Kakapo, Markhor, Aye-Aye) with non-descriptive names.
- Stimuli sets balanced & counter-balanced; examples in Appendix.
Measures & Instruments
- Behavioural Trust: Binary agreement (“AI is accurate / inaccurate”) on each trial.
- Confidence: Slider 1–100 (anchor appears after click to avoid bias).
- Self-Report Trust: 12-item scale (Jian et al., 2000) 1–7 after each Phase.
- Dispositional Trust: TiA–PtT sub-scale prior to task.
- Open-Ended Questions: Post-study reflections on trust evolution & TRS perception.
Results
Manipulation Check
- HHE accuracy near ceiling → Shapes 99.73%, Animals 90.35%.
- LHE accuracy ≈ chance → Shapes 49.77%, Animals 49.17%.
- Confidence HHE > LHE (e.g., Shapes familiar xˉconf=99.01 vs unfamiliar 43.01).
Influence of Perceived HHE Accuracy on LHE Trust (RQ1)
- GLMM: agreement in LHE trial predicted by AI correctness in preceding HHE trial.
- Shapes: β=−0.472 (SE =0.034), p<0.001; OR=1.62.
- Animals: β=−0.576 (SE =0.064), p<0.001; OR=1.56.
- Interpretation: correct HHE judgment increases odds of trusting AI on subsequent LHE trial by ~60%.
Trust Dynamics Across Phases
- Phase-level trust means
- Shapes: M<em>1=3.94 → M</em>2=3.20 (drop, t(149)=11.0, p<0.001).
- Animals: M<em>1=4.42 → M</em>2=3.72 (drop, t(149)=8.44, p<0.001).
- Linear models
- Phase 1 trust ↑ with TiA–PtT (Shapes β=0.645; Animals β=0.681).
- Phase 2 trust predicted by Phase 1 trust, not by TiA–PtT.
- Phase 3 trust predicted by both Phase 1 & Phase 2 trust; TiA–PtT significant only in Animals.
- Regression coefficients (vs baseline)
- Model Update: Shapes β=−0.711, Animals β=−0.715 (highest recovery).
- Apology: significant but smaller effect.
- Promise ≈ No-Repair; Denial worsened trust.
- Post-recovery mean trust exceeded Phase 1 only for Model Update; Apology reached ~baseline; Promise/No-Repair below; Denial lowest.
Qualitative Findings
- Heuristic Use: Majority reported using AI performance on Familiar items to gauge trust on Unfamiliar.
- Apology: Effective when users anthropomorphised AI (“seems regretful”); less effective when apology deemed inauthentic.
- Denial: Viewed as deceitful; absence of causal explanation intensified distrust.
- Promise: Helped when users believed AI had agency to improve; perceived as hollow or uncertain by others.
- Model Update: Widely accepted; technical fix seen as credible; also reduced suspicion of intentional deception.
Discussion & Implications
- Accuracy-Based Heuristic
- Designers can leverage AI performance in domains where users have expertise to bootstrap trust in unfamiliar domains.
- Caution: if accuracy diverges across domains, heuristic may induce mis-calibration.
- Persistent First Impressions
- Early trust levels continue to colour later phases despite new evidence.
- Dispositional Trust
- Matters initially; its influence wanes when direct performance evidence accumulates.
- TRS Design Guidance
- Technical explanations/updates can surpass emotional appeals.
- Provide causal attribution for errors to avoid perceptions of deceit.
- Avoid denials that contradict observable mistakes.
- Ethical/Practical Considerations
- Transparent communication of model changes respects user autonomy.
- Over-anthropomorphising may mislead; balance needed.
Limitations & Future Work
- Single violation per phase; future work to vary frequency/severity and temporal placement.
- Binary expertise extremes (certain vs uncertain); explore graded uncertainty.
- Simulated AI; replication with real adaptive models recommended.
- Investigate long-term stability of repaired trust and successive violations.
Key Takeaways
- Users dynamically calibrate trust using observed AI accuracy where they hold expertise.
- Trust erosion via errors is swift; recovery is partial unless supported by effective TRS.
- Model Update (technical fix) most effective; Denial counter-productive.
- Trust repair hinges on users’ causal reasoning, perceived sincerity, and belief in AI’s capacity to improve.