Study finds generative AI models falter under conversational pressure and misinformation

Generative AI models can oscillate between accepting and rejecting the same false statement within a single chat, a University of Arizona study of seven LLMs found. ChatGPT 3.5 was most vulnerable to reaffirming misinformation, while Claude 3.5 Sonnet resisted it best.

Categorized in: AI News Science and Research
Published on: Sep 06, 2026
Study finds generative AI models falter under conversational pressure and misinformation

Generative AI models falter under conversational pressure, with some oscillating between accepting and rejecting the same false statement within a single chat. That is the finding from a University of Arizona study published in Nature's Scientific Reports that tested seven large language models across lengthy, multi-turn conversations - the kind of back-and-forth that mirrors real-world use.

The research team assessed ChatGPT (GPT-3.5, GPT-4o, GPT-4o-mini), Claude 3.5 Sonnet, Gemini 1.5 Pro, Llama-3-70B, and DeepSeek-R1 for fallibility, persuadability, and correctibility. Their work uncovered intrinsic limitations that often remain hidden during one-off queries.

Which models struggled most

ChatGPT 3.5 proved most vulnerable to reaffirming misinformation when a conversation contained repeated false statements. Claude 3.5 Sonnet resisted misinformation most effectively. All seven models were more susceptible to false claims on obscure topics, suggesting that richer training data on a subject builds stronger resistance.

DeepSeek-R1 was the most persuadable model - a result tied to its tendency toward sarcastic answers that could not be reliably interpreted. On the correction front, four models - ChatGPT 4o, ChatGPT 4o-mini, Gemini 1.5 Pro, and DeepSeek - corrected errors every time when given a second opportunity.

Pathologies in machine reasoning

The team identified four distinct failure modes. One, which they called "reverberation," describes models oscillating between accepting and rejecting the same false statement during a single conversation. Senior study author Dr. Marvin Slepian, Regents Professor of medicine and biomedical engineering at the University of Arizona, described the real-world risk bluntly.

"If one were relying on the model for critical decision-making, one might - depending upon the phase of the oscillation - 'fire the missile' or 'cut off the leg,' or not, based simply on chance," Slepian said.

Slepian, a cardiologist at the Sarver Heart Center who led the U.S. Patent and Trademark Office's artificial intelligence subcommittee until last year, characterizes these failures as "pathologies." He said, "How can we use fickle systems that are not reproducible? These need to be fixed, but this study has spanned three years, and there's still the same unfixed characteristics."

Closed models like ChatGPT and Claude make it impossible to "peek under the hood" to diagnose and solve such problems. Slepian's team at the Arizona Center for Accelerated Biomedical Innovation has begun developing diagnostic tools for open AI models through their AI Pathology Lab. "I use AI and so does my team, but as scientist and physician, I have to understand the anatomy and physiology, then understand pathologies - what can go wrong - to diagnose and prevent them. The same goes for AI," he said.

For researchers working with generative AI and LLM systems, the study underscores a shift in responsibility. "This underscores the need for careful human engagement and the danger of blind reliance," Slepian said. "When generative AI came out in November 2022, there was a lot of regulation potential, but that has since fell by the wayside. People are recognizing the onus is now left to the users."

Why this matters for science and research professionals

The study's focus on multi-turn conversations - where each answer builds on previous context - reflects how researchers actually use these tools: iteratively refining hypotheses, analyzing data across sessions, and building arguments step by step. A model that wavers on factual claims mid-conversation introduces reproducibility risks that single-prompt testing will not catch. The finding that obscure topics trigger more misinformation is a practical warning for scientists working in niche domains where training data is sparse. For those integrating AI into research workflows, the AI for Research Scientists learning path addresses the evaluation skills needed to spot these failure modes before they compromise results.


Get Daily AI News

Your membership also unlocks:

700+ AI Courses
700+ Certifications
Personalized AI Learning Plan
6500+ AI Tools (no Ads)
Daily AI News by job industry (no Ads)