Electronic Thesis/Dissertation
 

Trust Calibration in Human Autonomous System Interaction

Open Access Deposited

Although large language models have shown encouraging diagnostic accuracy on static medical question-answering benchmarks, they routinely fall short of offering the self-correction mechanisms, calibrated uncertainty quantification, and structural safety assurances needed for practical deployment. The discrepancy between clinical reliability and benchmark performance is one of the most significant unresolved issues in AI systems in medicine. Regardless of its benchmark score, a system that scores 87% on the US Medical Licensing Examination but produces confident incorrect diagnoses in atypical presentations, anchors on a first false impression, and hallucinates pharmacological contraindications is not therapeutically effective. This gap arises in part because clinical diagnosis is inherently a sequential reasoning process. A model must gather evidence across multiple interactions, revise hypotheses as new test results arrive, and commit to a diagnosis only when sufficient evidence supports it. This thesis investigates whether these failures can be reduced through reasoning structure and decision monitoring built around the model rather than into it. The work starts from a systematic review of the trust calibration literature, which shaped three core design principles. Each of the three distinct elements of trustworthiness which are ability, integrity, and benevolence, needs to be examined separately at every stage of decision-making. Preventing catastrophic errors is more important than minor accuracy gains because trust declines more quickly after failure than it recovers after success. The decision to build a system that corrects its own unsafe behaviors rather than announcing them was influenced by findings showing that, after repeated failures, quiet behavioral correction is more effective than verbal explanation. This thesis presents a multi-layer reasoning architecture that transforms a single-agent large language model clinical consultation into a trust-calibrated diagnostic system. The framework extends the AgentClinic interactive diagnostic benchmark with a multi-expert debate pipeline and six safety components designed to improve reasoning reliability in sequential clinical decision making. An evidence classifier distinguishes pathognomonic findings from non-specific observations and prevents later consensus from overriding diagnostic evidence. A Partial Observability based belief tracker maintains an explicit differential diagnosis across the encounter and includes computational detectors that identify anchoring before final decisions are made. A specialist panel engages in an adversarial debate with a critic that challenges majority reasoning. A trust monitor evaluates each proposed action across ten dimensions and vetoes decisions that fall below a phase-calibrated threshold. When a veto occurs, a correction module generates a replacement action grounded in the evidence record rather than the failed reasoning trace. Together these mechanisms transform the original single-agent baseline into a substantially expanded safety-oriented architecture incorporating multiple layers of reasoning verification and decision oversight. The system is evaluated through a five-arm ablation study comprising 1,060 simulation runs across 106 clinical diagnostic scenarios and two model classes, a specialist fine-tuned model Llama-3.1-8B-UltraMedical and a general-purpose model Llama-3.1-8B-Instruct. The results reveal a model-adaptive impact. Safety layers improve the general model by 11.3% points, from 38.7% to 50.0%, but occasionally reduce accuracy on specialist models for cases already well-represented in the training distribution. Passive trust monitoring outperforms active veto correction, achieving 50.0% compared with 47.2%, because the correction step sometimes replaces correct diagnoses with confident but incorrect alternatives due to linguistic noise in the 8B model's correction reasoning. Of 38 scenarios universally failed by baseline configurations, 39.5% were correctly resolved by the layer-enhanced architecture, demonstrating that structured reasoning rather than model scale is the primary determinant of clinical AI trustworthiness in resource-constrained settings. These findings position the framework for Phase 2 deployment on the Pepper social robot.

Author Language Date created Type of Work License
  • All rights reserved
Rights statement GW Unit Degree Advisor Committee Member(s) Persistent URL

Notice to Authors

If you are the author of this work and you have any questions about the information on this page, please use the Contact form to get in touch with us.

Thumbnail Title Date Uploaded Visibility Actions
Preview of Nayak_gwu_0075M_17939.pdf Nayak_gwu_0075M_17939.pdf 2026-06-24 Open Access