OpenAI’s Mental Health AI Test Reveals Gaps in Context and Urgency

OpenAI’s Mental Health AI Test Reveals Gaps in Context and Urgency

OpenAI releases open benchmark for evaluating AI mental health support. Image: Zulfugar Karimov/Unsplash

OpenAI’s MentalHealthBench tests responses to 1,215 synthetic mental health conversations, revealing gaps in how AI models seek context and gauge urgency.

Verfasst von
Aminu Abdullahi
Aminu Abdullahi
Sep 25, 2026

AI chatbots are getting better at sounding caring, but OpenAI wants to know whether they are actually giving people the right help when conversations get difficult.

On Wednesday, OpenAI introduced MentalHealthBench, an open evaluation dataset developed with over 80 licensed psychologists and psychiatrists across 22 countries. The public benchmark tests how well AI systems handle emotionally sensitive dialogues, moving beyond simple emergency safety filters to measure real-world conversational competence across 1,215 synthetic scenarios.

Inside the clinical test

OpenAI designed the benchmark to test conversations ranging from everyday concerns to mental health emergencies. According to the company’s research release, non-acute everyday stress comprises 53.5% of the synthetic chats, high-acuity distress makes up 18.2%, and urgent emergencies account for 28.3%. Scenarios reflect four distinct user profiles: adults, teenagers aged 13 to 17, caregivers, and clinicians.

To score the models, participating experts across nearly 20 mental health subspecialties crafted 5,262 rubric criteria, weighting behaviors from -10 to +10 based on clinical importance. An automated grader, GPT-5.6 Sol, assessed model replies across 10 behavioral dimensions, including context-seeking, agency, and empathy.

OpenAI evaluated major systems on the dataset, finding that frontier models show notable gaps. GPT-6 Astra led evaluations with a 57.3% task-clipped score, followed by GPT-6 Sol at 53.9% and Anthropic’s Claude Opus 5.5 at 52.4%. Older models trailed, with GPT-4o scoring 32.1%. These figures measure responses to synthetic conversations against expert-written criteria; they do not show how often a model helps someone in a real mental health conversation.

“ChatGPT is not a therapist, and is not here to replace a clinician,” Dr. Declan Grabb, mental health safety research lead at OpenAI, told The Deep View. “That being said, when I talk to mental health clinicians across the globe, the most responsible and safe thing to do is if people are coming to AI to ask these questions, we absolutely need to have an expert opinion on how you should navigate them.”

Advertisement

The gap between clinical care and digital convenience

OpenAI also surveyed 44 adults who have used AI for emotional support, revealing an institutional rift: while clinicians prioritized cautious context-gathering and slow assessment, everyday users wanted rapid, practical next steps and an empathetic conversational tone.

This tension creates an engineering tightrope. An algorithm optimized to sound immediately comforting can easily validate distorted thinking or rush into giving flawed lifestyle advice before understanding a user’s background.

Conversely, an AI that acts like a cautious, interrogating clinician risks alienating a vulnerable person who turned to a chatbot precisely to avoid clinical friction. By attempting to bridge this divide with an open benchmark, the industry must reckon with whether an automated system can offer genuine emotional utility without leaning into superficial sycophancy.

What this means for everyday chatbot users

For the estimated one billion people using ChatGPT weekly, this benchmark highlights clear guardrails and limitations. Chatbots remain software, not clinical professionals, and top models still struggle to ask necessary clarifying questions or calibrate urgency across long conversations.

OpenAI has added crisis resources, Trusted Contact, and protections for younger users. For IT and HR teams evaluating chatbots that people may use for sensitive conversations, MentalHealthBench offers a way to inspect model behavior, but a benchmark score alone is not proof that a product can provide clinical care. Teams should check how their chosen tool handles urgent risk, directs users to human support, and protects sensitive disclosures.

Read more: What should HR and IT leaders check before adopting AI mental health tools? The guide examines clinical safeguards, privacy, and the need for human oversight in workplace deployments.

Aminu Abdullahi

Aminu Abdullahi is a B2C and B2B technology and finance writer with more than six years of experience covering enterprise IT, cybersecurity, cloud computing, artificial intelligence, fintech, business software, and emerging technologies. He has written for a wide range of technical and business audiences, from IT professionals and cybersecurity leaders to small business owners, executives, and technology buyers. His work has appeared in publications including: TechRepublic eWEEK Channel Insider Geekflare Enterprise Networking Planet eSecurity Planet CIO Insight Webopedia With a background in computer science, Aminu specializes in translating complex technical subjects into clear, practical, and accessible content. His writing helps readers understand emerging technologies, evaluate business software, strengthen cybersecurity strategies, and make more informed decisions about technology investments. Across his work, Aminu focuses on the real-world impact of technology, connecting technical innovation with business value, operational efficiency, security, and long-term digital transformation.