Trust Development & Repair in AI-Assisted Decisions – Comprehensive Study Notes

Background & Key Concepts

  • Human–AI Collaboration
    • Seeks to combine complementary strengths of humans (contextual reasoning, empathy, domain insight) and AI (speed, pattern-recognition, memorisation).
    • Central challenge: users must decide when and how much to trust AI advice.
  • Complementary vs. Overlapping Expertise
    • Overlapping: both human and AI competent → humans can directly judge AI accuracy.
    • Complementary: each excels at different facets → humans cannot always assess AI accuracy.
  • Appropriate Trust / Trust Calibration
    • Balance between over-trust (unwarranted reliance) and under-trust (undue scepticism).
    • Calibration aims to align subjective trust with objective AI capability.
  • Trust Repair Strategies (TRSs) – interventions adopted from Social Psychology & Human–Robot Interaction (HRI) to rebuild trust after violations.
    • Apology (expression of regret)
    • Denial (reject culpability)
    • Promise (commit to do better)
    • Model Update (inform user of technical improvement)
  • Dispositional Trust (TiA-PtT) – a trait-level propensity to trust automation; moderates initial trust.

Research Questions & Contributions

  • RQ1: How does perceived AI accuracy in High Human-Expertise (HHE) tasks influence users’ trust during Low Human-Expertise (LHE) tasks?
  • RQ2: In complementary expertise settings, how does trust recover when AI accuracy improves, with and without explicit TRSs?
  • Key contributions
    • Empirical evidence that users extrapolate AI accuracy in HHE tasks to calibrate trust in LHE tasks.
    • Comparative efficacy of four TRSs; Model Update > Apology > Promise ≈ No-Repair > Denial.
    • Insight that trust recovery depends on anthropomorphism, perceived regret/deceit, causal attribution, and nature of improvement (behavioural vs technical).
    • Dual-task (Shapes & Animals) replication enhances ecological validity.

Related Work

  • Trust as an attitude under uncertainty (Lee & See, 2004).
  • Accuracy → trust correlation (Yu et al., 2016; Yin et al., 2019).
  • Initial impressions / early errors have long-term effects (Tolmeijer et al., 2021).
  • HRI literature on trust repair: mixed findings for apology, promise, denial; scarcity of work in non-robotic AI decision aids.

Trust Repair Strategies (TRS) – Details

  • Apology
    • Operates emotionally; seeks to restore social expectations.
    • Text used: “I’m sorry … I hope you can trust me again.”
  • Denial
    • Shifts blame; may appear deceptive if evidence contradicts statement.
  • Promise
    • Behavioural commitment; effectiveness relies on perceived agency to fulfil promise.
  • Model Update (novel)
    • Communicates algorithmic upgrade; implicitly supplies causal attribution (flawed model → fixed).
  • All TRS messages followed template: acknowledgement of mistrust → core element → hope for restored trust.

Methodology

  • Design: Survey-based, between-subjects; 55 TRS conditions (including baseline No-Repair) × 22 tasks.
  • Participants: N=300N = 300 Prolific users (eng ≥ 98%98\%); 150150 per task, 3030 per TRS.
    • Mean ages: Shapes xˉ=35\bar{x}=35 (SD =13.01=13.01); Animals xˉ=34.2\bar{x}=34.2 (SD =11.85=11.85).
  • Power Analysis: f2=0.25f^{2}=0.25, α=0.05\alpha = 0.05, power =0.8=0.8 ⇒ minimum n=135n=135 per task (G*Power 3).
  • Phases & AI Accuracy
    1. Phase 1 – High accuracy =80%=80\% (1 error at trial 7) → build trust.
    2. Phase 2 – Low accuracy =20%=20\% (multiple errors) → violate trust.
    3. Phase 3 – High accuracy =80%=80\% again; TRS shown before start.
  • Trial Types (10 per phase; 30 total)
    • HHE (Familiar) = human knows answer.
    • LHE (Unfamiliar) = human does not know; must rely on AI.
  • Delay: 3  s3\;\text{s} before AI advice to encourage independent thought.
  • Model-Update scenario: extra 6  s6\;\text{s} wait to simulate retraining.

Experimental Tasks & Stimuli

  • Shapes Task
    • Familiar: Circle, Rectangle, Triangle (5 visual variants each).
    • Unfamiliar (fabricated): “Scleratice”, “Tenectus”, “Pyrangle”; randomised fill/border patterns.
  • Animals Task
    • Familiar: 15 common species (Cat, Dog, Horse, etc.).
    • Unfamiliar: 15 obscure species (e.g., Kakapo, Markhor, Aye-Aye) with non-descriptive names.
  • Stimuli sets balanced & counter-balanced; examples in Appendix.

Measures & Instruments

  • Behavioural Trust: Binary agreement (“AI is accurate / inaccurate”) on each trial.
  • Confidence: Slider 11001\text{–}100 (anchor appears after click to avoid bias).
  • Self-Report Trust: 1212-item scale (Jian et al., 2000) 171\text{–}7 after each Phase.
  • Dispositional Trust: TiA–PtT\text{TiA–PtT} sub-scale prior to task.
  • Open-Ended Questions: Post-study reflections on trust evolution & TRS perception.

Results

Manipulation Check

  • HHE accuracy near ceiling → Shapes 99.73%99.73\%, Animals 90.35%90.35\%.
  • LHE accuracy ≈ chance → Shapes 49.77%49.77\%, Animals 49.17%49.17\%.
  • Confidence HHE > LHE (e.g., Shapes familiar xˉconf=99.01\bar{x}_{conf}=99.01 vs unfamiliar 43.0143.01).

Influence of Perceived HHE Accuracy on LHE Trust (RQ1)

  • GLMM: agreement in LHE trial predicted by AI correctness in preceding HHE trial.
    • Shapes: β=0.472\beta = -0.472 (SE =0.034=0.034), p<0.001; OR=1.62\text{OR}=1.62.
    • Animals: β=0.576\beta = -0.576 (SE =0.064=0.064), p<0.001; OR=1.56\text{OR}=1.56.
  • Interpretation: correct HHE judgment increases odds of trusting AI on subsequent LHE trial by ~60%60\%.

Trust Dynamics Across Phases

  • Phase-level trust means
    • Shapes: M<em>1=3.94M<em>1=3.94M</em>2=3.20M</em>2=3.20 (drop, t(149)=11.0, p<0.001).
    • Animals: M<em>1=4.42M<em>1=4.42M</em>2=3.72M</em>2=3.72 (drop, t(149)=8.44, p<0.001).
  • Linear models
    • Phase 1 trust ↑ with TiA–PtT\text{TiA–PtT} (Shapes β=0.645\beta=0.645; Animals β=0.681\beta=0.681).
    • Phase 2 trust predicted by Phase 1 trust, not by TiA–PtT\text{TiA–PtT}.
    • Phase 3 trust predicted by both Phase 1 & Phase 2 trust; TiA–PtT\text{TiA–PtT} significant only in Animals.

Effectiveness of TRSs (RQ2)

  • Regression coefficients (vs baseline)
    • Model Update: Shapes β=0.711\beta=-0.711, Animals β=0.715\beta=-0.715 (highest recovery).
    • Apology: significant but smaller effect.
    • Promise ≈ No-Repair; Denial worsened trust.
  • Post-recovery mean trust exceeded Phase 1 only for Model Update; Apology reached ~baseline; Promise/No-Repair below; Denial lowest.

Qualitative Findings

  • Heuristic Use: Majority reported using AI performance on Familiar items to gauge trust on Unfamiliar.
  • Apology: Effective when users anthropomorphised AI (“seems regretful”); less effective when apology deemed inauthentic.
  • Denial: Viewed as deceitful; absence of causal explanation intensified distrust.
  • Promise: Helped when users believed AI had agency to improve; perceived as hollow or uncertain by others.
  • Model Update: Widely accepted; technical fix seen as credible; also reduced suspicion of intentional deception.

Discussion & Implications

  • Accuracy-Based Heuristic
    • Designers can leverage AI performance in domains where users have expertise to bootstrap trust in unfamiliar domains.
    • Caution: if accuracy diverges across domains, heuristic may induce mis-calibration.
  • Persistent First Impressions
    • Early trust levels continue to colour later phases despite new evidence.
  • Dispositional Trust
    • Matters initially; its influence wanes when direct performance evidence accumulates.
  • TRS Design Guidance
    • Technical explanations/updates can surpass emotional appeals.
    • Provide causal attribution for errors to avoid perceptions of deceit.
    • Avoid denials that contradict observable mistakes.
  • Ethical/Practical Considerations
    • Transparent communication of model changes respects user autonomy.
    • Over-anthropomorphising may mislead; balance needed.

Limitations & Future Work

  • Single violation per phase; future work to vary frequency/severity and temporal placement.
  • Binary expertise extremes (certain vs uncertain); explore graded uncertainty.
  • Simulated AI; replication with real adaptive models recommended.
  • Investigate long-term stability of repaired trust and successive violations.

Key Takeaways

  • Users dynamically calibrate trust using observed AI accuracy where they hold expertise.
  • Trust erosion via errors is swift; recovery is partial unless supported by effective TRS.
  • Model Update (technical fix) most effective; Denial counter-productive.
  • Trust repair hinges on users’ causal reasoning, perceived sincerity, and belief in AI’s capacity to improve.