A University of Arizona study published in Scientific Reports on September 5, 2026, found that all seven leading large language models tested become more vulnerable to misinformation during extended, multi-turn conversations. The research exposes failure patterns that remain hidden in single-question interactions, raising safety questions as AI moves deeper into high-stakes professional environments.
"When generative AI came out in November 2022, there was a lot of potential for regulation, but that has since fallen by the wayside. People are recognizing the onus is now left to the users," said senior study author Dr. Marvin Slepian, Regents Professor of medicine and biomedical engineering at the University of Arizona.
How the models performed under pressure
The research team tested ChatGPT (GPT-3.5, GPT-4o, and GPT-4o-mini), Claude 3.5 Sonnet, Gemini 1.5 Pro, Llama-3-70B, and DeepSeek-R1 across three dimensions: fallibility, persuadability, and correctability. ChatGPT 3.5 proved most vulnerable to reaffirming misinformation when confronted with repeated false statements during a conversation. Claude 3.5 Sonnet resisted best.
All seven models struggled more with obscure topics. The pattern suggests that more training data on a subject builds stronger resistance to false claims. DeepSeek was the most persuadable, largely because its sarcastic responses could not be reliably interpreted as agreement or rejection. Four models-ChatGPT 4o, ChatGPT 4o-mini, Gemini 1.5 Pro, and DeepSeek-corrected errors 100% of the time when researchers gave them a second chance.
Four failure pathologies
The team identified four distinct ways models failed to affirm factual information. One behavior, which Slepian labeled "reverberation," involved models oscillating between accepting and rejecting the same false statement within a single conversation. "If one were relying on the model for critical decision-making, one might-depending upon the phase of the oscillation-'fire the missile' or 'cut off the leg,' or not, based simply on chance," Slepian said.
As a cardiologist and member of the Sarver Heart Center, Slepian characterizes these failures as pathologies. "How can we use fickle systems that are not reproducible? These need to be fixed, but this study has spanned three years, and there are still the same unfixed characteristics," he said. Closed models like ChatGPT and Claude make internal diagnosis impossible. Through the Arizona Center for Accelerated Biomedical Innovation, Slepian's team is now building diagnostic tools for open models as part of their AI Pathology Lab.
For researchers and scientists integrating LLMs into their workflows, understanding these failure modes is increasingly critical. Courses covering Generative AI and LLM Courses can help technical professionals recognize when models are likely to falter under sustained conversational pressure.
Why this matters for science and research professionals
The study confirms that single-prompt evaluations of AI reliability paint an incomplete picture. Multi-turn conversations-the kind researchers use when iterating on hypotheses, analyzing data, or drafting manuscripts-introduce vulnerabilities that one-off queries miss. Oscillation between correct and incorrect answers means a model's output on a given turn may be essentially random, a risk factor in any workflow where decisions compound across exchanges. Professionals using AI for Science & Research should build verification checkpoints into extended LLM interactions rather than treating the conversation as a single continuous thread of trusted output.
Your membership also unlocks: