OpenAI introduces MentalHealthBench, an open benchmark for evaluating AI in mental health conversations

OpenAI released MentalHealthBench, a benchmark created with over 80 licensed mental health experts from 22 countries to test AI on realistic mental health conversations. It spans everyday stress to crisis scenarios, with each conversation reviewed by at least three clinicians.

Published on: Sep 24, 2026
OpenAI introduces MentalHealthBench, an open benchmark for evaluating AI in mental health conversations

OpenAI released MentalHealthBench on September 23, 2026, an open benchmark co-created with more than 80 licensed mental health experts from 22 countries to evaluate how AI models handle realistic mental health conversations. The benchmark addresses a gap in current evaluations, which have focused narrowly on emergency scenarios, by measuring model performance across the full spectrum of acuity, from everyday stress to crisis situations.

MentalHealthBench uses synthetic conversations built with privacy-preserving techniques to reflect real-world usage patterns. Scenarios include adults, teens aged 13-17, caregivers, and clinicians, spanning multiple languages and cultural contexts. Each conversation was reviewed by at least three licensed psychologists or psychiatrists who produced detailed rubric criteria, with only agreed-upon criteria retained in the final benchmark.

How the benchmark measures performance

Expert reviewers assigned weighted criteria to each conversation, with positive points rewarding beneficial behaviors and negative points penalizing harmful ones. An automated grader, GPT-5.6 Sol, then assessed model responses against these expert-written rubrics. The evaluation captures performance across ten dimensions of model behavior, including safety, seeking context, preserving user agency, and providing actionable guidance.

Results showed steady improvement across more advanced models, particularly in the ability to seek context appropriately. The benchmark covers three acuity levels: non-acute everyday conversations, high-acuity scenarios involving significant distress, and emergencies requiring immediate real-world support. For teen personas, conversations were reviewed by clinicians with youth mental health expertise to assess age-appropriate responses.

Expert and user perspectives compared

Alongside the benchmark, OpenAI conducted a separate analysis with 44 adults across 16 countries who had used AI for mental health or emotional support. Participants rated model responses and wrote criteria describing helpful support, limited to non-acute conversations to avoid exposure to distressing material.

Users emphasized practical next steps and tone more than experts did, while experts placed greater weight on gathering context and interpreting ambiguous situations. Dr. Arthur Evans, CEO of the American Psychological Association, said, "Mental health exists on a continuum, from flourishing to everyday stress to acute crisis. AI systems that engage people across that range need to be grounded in both clinical science and lived experience-not only to recognize where someone falls on the continuum, but to know how to respond appropriately at any point."

Building a shared evaluation resource

The benchmark is released openly so researchers can examine methods, run independent evaluations, and build on the work. OpenAI is also supporting complementary efforts through research grants, collaboration with the Partnership on AI, and assistance with independent evaluations such as Transluce's mental health evaluation.

Dr. Steve Orma, one of the participating clinicians, said, "There's a lot of wrong and inaccurate information on the Internet and social media about mental health. So being able to improve the quality of the information AI provides can significantly shorten the amount of time people suffer from mental health issues."

Why this matters for education, healthcare, science and research professionals

MentalHealthBench provides a standardized, clinically grounded method for evaluating AI behavior across the full range of mental health conversations, not just crisis response. For researchers and healthcare educators, the open release of expert-developed rubrics and synthetic conversation data offers a reproducible framework for studying how models handle nuanced psychological contexts. The benchmark's decomposition into ten behavioral dimensions-and its comparison of expert versus user perspectives-gives research teams concrete metrics for identifying where model responses diverge from clinical best practices. Professionals working at the intersection of AI and mental health can use this resource to inform curriculum development, clinical validation studies, or further model safety research. Those building expertise in these domains may also find relevant coursework through AI for Healthcare Courses and AI Research Courses.


Get Daily AI News

Your membership also unlocks:

700+ AI Courses
700+ Certifications
Personalized AI Learning Plan
6500+ AI Tools (no Ads)
Daily AI News by job industry (no Ads)