When the Chatbot Becomes the Doctor: On the Right Level of Trust in Medical AI Advisory
by Vanessa Yarza Navarro-Schär
A pastor from Florida is suing OpenAI. He accuses the company of having its chatbot, ChatGPT, advise him incorrectly in a serious medical situation. According to the lawsuit, the system repeatedly classified his symptoms as harmless. It discouraged him from seeing a doctor and reinforced his decision to stay at home, even when the people around him urged him to go to the hospital. Over several weeks, his condition deteriorated. On 13 July 2025, he was admitted to intensive care with a severe pulmonary embolism (New York Times, 2026).
The case is extreme. But it does not stand alone. Relying on chatbots for health questions has become an everyday practice and affects a very large number of users each week. This makes visible what research has been describing for some time: the success of medical AI advisory depends not only on the capability of the system but on users' trust in that system (Chen et al., 2025). This question arises across all age groups, but it arises with particular acuity when older people turn to such tools, a group that increasingly uses medical AI advisory and who, as a population, tend to have greater needs for health information.
A Well-Documented Phenomenon
That people extend trust to AI systems in the medical context is empirically well established. The findings of Shekar et al. (2025) illustrate the extent of this trust. In an experimental setting, many respondents rated AI-generated medical answers as at least as trustworthy as physician answers. AI systems can also generate responses that appear coherent while being factually incorrect. In high-risk medical settings, such hallucinations carry serious consequences (Bélisle-Pipon, 2024). Because these responses mimic human reasoning, users tend to ascribe greater competence to the system than is warranted (Bélisle-Pipon, 2024). Trust in AI refers to a user's readiness to depend on the system and to bear the risk of its errors, grounded in the belief that it is capable and acts with good intent (Bach et al., 2024; Li et al., 2025). What matters is not whether trust is present, but whether it is present at the right level. Two misalignments are central here.
Overtrust and Undertrust
Overtrust occurs when trust is higher than the performance of the AI system justifies (Aroyo et al., 2021; Ullrich et al., 2021). The consequence: users examine the answers less critically and adopt them almost unfiltered (Ullrich et al., 2021). In the medical context, this carries heavy weight. Whoever follows an incorrect recommendation or postpones necessary care may put off care until the condition has become dangerous (Shekar et al., 2025). The case from Florida stands for overtrust.
Undertrust is the reverse error. Trust falls short of the performance of the system, often as a reaction to a preceding AI error (Klingbeil et al., 2024). This too is problematic: the system is avoided, and its benefit goes unused (Gillath et al., 2021).
Both states are therefore misalignments. The goal is thus neither as much trust as possible nor as little trust as possible.
Calibrated Trust as the Goal
The aim is calibrated trust. This means that users' trust and the actual performance of the AI system match one another. Calibrated trust takes shape when the interaction supplies enough information for users to correct their expectations, lowering trust where the system falls short and raising it where the system performs (Schlicker et al., 2025).
In domains as consequential as medicine, calibrated trust becomes a requirement rather than a nicety. Both misalignments are dangerous here. Overtrust can lead users to follow an incorrect recommendation. Undertrust can lead them to reject helpful support. In both cases, harmful decisions loom (Lee & See, 2004). The question of how trust can be calibrated is therefore not only a question of acceptance. It is a matter of managing risk in clinical decisions (Goddard et al., 2012).
Open Research Questions
How calibrated trust can be established in everyday medical practice is only partly understood. From our perspective, several directions suggest themselves.
First, and central for us: the role of age. Older people are a growing user group of medical AI advice. They frequently have health questions, and access barriers in the healthcare system reinforce the incentive to turn to a chatbot to answer them. It remains open how overtrust and undertrust in AI systems are distributed across the lifespan. It also remains open whether older users could require different calibration aids than younger ones.
Second: the role of diagnostic cues. Recalibration hinges on the signals a system provides about its own reliability (Schlicker et al., 2025). Which signals prove effective in everyday medical practice, and how they should be designed so that they foster neither overtrust nor undertrust, remains largely unresolved (Kox et al., 2025).
Third: the social background. In the case from Florida, the chatbot stepped between the user and the real world. What role family, professionals, and community play in the calibration of trust deserves closer investigation.
Do You Observe This as Well?
At the MIT AgeLab, we investigate how people build, lose, and restore trust in technological systems across the lifespan. If, in your organization, your practice, or your research, you observe how trust in medical AI advisory develops, and you would like to investigate this phenomenon together, we welcome the exchange.
Photo credit: Kiichiro Sato/Associated Press via The New York Times
References
Aroyo, A. M., De Bruyne, J., Dheu, O., Fosch-Villaronga, E., Gudkov, A., Hoch, H., Jones, S., Lutz, C., Sætra, H., Solberg, M., & Tamò-Larrieux, A. (2021). Overtrusting robots: Setting a research agenda to mitigate overtrust in automation. Paladyn, Journal of Behavioral Robotics, 12(1), 423–436. https://doi.org/10.1515/pjbr-2021-0029\
Bach, T. A., Khan, A., Hallock, H., Beltrão, G., & Sousa, S. (2024). A Systematic Literature Review of User Trust in AI-Enabled Systems: An HCI Perspective. International Journal of Human–Computer Interaction, 40(5), 1251–1266. https://doi.org/10.1080/10447318.2022.2138826\
Bélisle-Pipon, J.-C. (2024). Why we need to be careful with LLMs in medicine. Frontiers in Medicine, 11, 1495582. https://doi.org/10.3389/fmed.2024.1495582\
Chen, Y., Luo, S., & Yin, Y. (2025). Are you willing to forgive generative AI doctors? Trust repair after failures in online health consultation services. Frontiers in Psychology, 16, 1668633. https://doi.org/10.3389/fpsyg.2025.1668633\
Elkins, A. C., & Derrick, D. C. (2013). The Sound of Trust: Voice as a Measurement of Trust During Interactions with Embodied Conversational Agents. Group Decision and Negotiation, 22(5), 897–913. https://doi.org/10.1007/s10726-012-9339-x\
Gillath, O., Ai, T., Branicky, M. S., Keshmiri, S., Davison, R. B., & Spaulding, R. (2021). Attachment and trust in artificial intelligence. Computers in Human Behavior, 115, 106607. https://doi.org/10.1016/j.chb.2020.106607 Rosenbluth, T. (2026, 22. Juli). ChatGPT led to a man’s near-fatal health crisis, lawsuit claims. The New York Times.
Goddard, K., Roudsari, A., & Wyatt, J. C. (2012). Automation bias: A systematic review of frequency, effect mediators, and mitigators. Journal of the American Medical Informatics Association, 19(1), 121–127.
Klingbeil, A., Grützner, C., & Schreck, P. (2024). Trust and reliance on AI — An experimental study on the extent and costs of overreliance on AI. Computers in Human Behavior, 160, 108352. https://doi.org/10.1016/j.chb.2024.108352
Kox, E., Hennekens, M., Metcalfe, J., & Kerstholt, J. (2025). Trust Violations due to Error or Choice: The Differential Effects on Trust Repair in Human–Human and Human–Robot Interaction. ACM Transactions on Human-Robot Interaction, 14(4), 1–27. https://doi.org/10.1145/3743694
Lee, J. D., & See, K. A. (2004). Trust in Automation: Designing for Appropriate Reliance. Human Factors.
Li, H., Liu, Y., Du, J., & Pan, Y. (2025). Understanding User Trust in Interactive AI Learning Environments: The Role of Social Perception and User Experience. IEEE Access, 13, 194995–195015. https://doi.org/10.1109/ACCESS.2025.3624627
Schlicker, N., Baum, K., Uhde, A., Sterz, S., Hirsch, M. C., & Langer, M. (2025). How do we assess the trustworthiness of AI? Introducing the trustworthiness assessment model (TrAM). Computers in Human Behavior, 170, 108671.
Shekar, S., Pataranutaporn, P., Sarabu, C., Cecchi, G. A., & Maes, P. (2025). People Overtrust AI-Generated Medical Advice despite Low Accuracy. NEJM AI, 2(6). https://doi.org/10.1056/AIoa2300015
Ullrich, D., Butz, A., & Diefenbach, S. (2021). The Development of Overtrust: An Empirical Simulation and Psychological Analysis in the Context of Human–Robot Interaction. Frontiers in Robotics and AI, 8, 554578. https://doi.org/10.3389/frobt.2021.554578

