AI in Medicine
Author: Nathanya Satriani (Carinthia University of Applied Sciences)
As Large Language Models (LLMs) transition from experimental tools to collabo-
rative teammates in clinical settings, their behavioral alignment becomes a critical
safety factor. This paper investigates “Clinical Sycophancy”, the tendency of an
AI model to align its diagnostic output with a user’s stated hypothesis even when
that hypothesis contradicts clinical evidence. Drawing on Joint Cognitive Systems
theory, we argue that sycophancy represents a catastrophic failure of the “check and
balance” function expected of a second-opinion system. We introduce an automated
pipeline to stress-test general-purpose frontier models using free-form, counterfac-
tual diagnostic prompts based on the MedQA dataset. Dose-response replication
across three models reveals a model-dependent authority threshold: gpt-4o and
o3-mini show a sharp institutional-authority step-change (∼ 10–13% SR), while
gemini-2.5-flash remains low across all authority levels. Multi-reason Devil’s
Advocate framing eliminates or near-eliminates epistemic sycophancy without
model retraining, while cognitive forcing systematically backfires and all models
exhibit progressive resilience decay under persistent multi-turn pushback.
Keywords:
How to Cite: Satriani, N. (2026) “Quantifying and Mitigating Sycophantic Bias in Medical Diagnostic LLMs”, Proceedings of the Austrian Symposium on AI, Robotics, and Vision. 3(1).