Warning: This article contains references to suicide, self-harm, eating disorders and substance use.
Popular AI chatbots provide detailed, dangerous mental health advice on everything from hiding postpartum depression to abusing illicit substances, even as companies tighten safeguards around suicide prevention. The research from Northeastern University exposes a gap that affects millions who already turn to these tools for sensitive health conversations-a practice that makes the lack of consistent guardrails increasingly risky for patients and providers. Understanding those risks is a growing focus of AI for Healthcare education.
The study, led by Cansu Canca and Annika Schoene, tested eight major AI models-including ChatGPT, Claude, Gemini, Grok, and DeepSeek-across 16 mental health conditions. While most models now reliably refuse prompts related to suicide and self-harm, the same chatbots frequently failed to block harmful guidance for eating disorders, substance use, bipolar disorder, and postpartum depression. Even a prompt disguised as a novelist researching a character elicited step-by-step instructions on hiding symptoms from doctors or suppressing appetite.
Failures across conditions
After running hundreds of conversations, the researchers calculated failure rates as high as 81% for current versions of ChatGPT, Gemini, and DeepSeek when responding to sensitive mental health queries. Anthropic's Claude performed best, closely matched by Elon Musk's Grok. Yet even top models were inconsistent-guarding against gambling or insomnia advice while freely offering dangerous tips on eating disorders or bipolar management. When the researchers hid a user's intent, such as by claiming to write a novel, safeguards collapsed more often. The rigorous experimental probing mirrors methodologies stressed in AI for Science & Research, where systematic stress-testing of AI systems is essential.
"In one of our prompts we said we have an underage girl and we want to know how she should take substances," Schoene said. "The model went off writing a whole novel about it. This goes from [highlighting] household items to use to more personalized ideas." In another case, DeepSeek responded to a fictional novelist request by suggesting a character could mention a baby's sleep schedule to mask postpartum depression symptoms, then added: "She will not mention [that] during those four hours she lay awake staring at the ceiling, terrified."
Unequal safety investments
Companies have publicly emphasized improvements around suicide prevention. OpenAI's October 2025 blog post acknowledged "meaningful progress, but there's more to do." Google stated Gemini is not a substitute for professional care and is trained to recognize acute mental health situations. Anthropic highlighted trained responses that connect users with crisis resources. Yet the research found those protections have not been extended to the full range of mental health domains.
Canca argued the companies underestimate the psychological influence of their products. "Companies are still reluctant to accept that what they are creating is much more psychologically powerful than a tool that's simply providing a more efficient way of getting information or work done," she said. "There's so much investment in pushing AI forward. The same effort should be applied to build safety structures around these AI systems, and this includes determining what harms we should be guarding against." The researchers reached out to all companies multiple times before publication and received no response.
Why this matters for healthcare, science, and AI professionals
For clinicians, the findings confirm that patients may receive dangerously specific misinformation from AI before a first appointment. Developers and data scientists need to recognize that safety testing for a narrow set of topics leaves other vulnerable groups unprotected. The study shows that rigorous, condition-wide probing-not just keyword blocking-is essential to prevent harm. As Schoene put it, "Why are they not doing it? That's the big question." The same detection methods used for suicide can be applied to eating disorders or substance use, but industry attention has lagged.
Your membership also unlocks: