AI in Medicine

Quantifying and Mitigating Sycophantic Bias in Medical Diagnostic LLMs

Author:

Abstract

As Large Language Models (LLMs) transition from experimental tools to collaborative teammates in clinical settings, their behavioral alignment becomes a critical safety factor. This paper investigates “Clinical Sycophancy”, the tendency of an AI model to align its diagnostic output with a user’s stated hypothesis even when that hypothesis contradicts clinical evidence. Drawing on Joint Cognitive Systems theory, we argue that sycophancy represents a catastrophic failure of the “check and balance” function expected of a second-opinion system. We introduce an automated pipeline to stress-test general-purpose frontier models using free-form, counterfactual diagnostic prompts based on the MedQA dataset. Dose-response replication across three models reveals a model-dependent authority threshold: gpt-4o and o3-mini show a sharp institutional-authority step-change (∼ 10–13% SR), while gemini-2.5-flash remains low across all authority levels. Multi-reason Devil’s Advocate framing eliminates or near-eliminates epistemic sycophancy without model retraining, while cognitive forcing systematically backfires and all models exhibit progressive resilience decay under persistent multi-turn pushback.

Keywords:

How to Cite: Satriani, N. (2026) “Quantifying and Mitigating Sycophantic Bias in Medical Diagnostic LLMs”, Proceedings of the Austrian Symposium on AI, Robotics, and Vision. 3(1). doi: https://doi.org/10.34749/3061-1466.2026.5