AI chatbots are getting better at sounding caring, but OpenAI wants to know whether they are actually giving people the right help when conversations get difficult.
On Wednesday, OpenAI introduced MentalHealthBench, an open evaluation dataset developed with over 80 licensed psychologists and psychiatrists across 22 countries. The public benchmark tests how well AI systems handle emotionally sensitive dialogues, moving beyond simple emergency safety filters to measure real-world conversational competence across 1,215 synthetic scenarios.
Inside the clinical test
OpenAI designed the benchmark to test conversations ranging from everyday concerns to mental health emergencies. According to the company’s research release, non-acute everyday stress comprises 53.5% of the synthetic chats, high-acuity distress makes up 18.2%, and urgent emergencies account for 28.3%. Scenarios reflect four distinct user profiles: adults, teenagers aged 13 to 17, caregivers, and clinicians.
To score the models, participating experts across nearly 20 mental health subspecialties crafted 5,262 rubric criteria, weighting behaviors from -10 to +10 based on clinical importance. An automated grader, GPT-5.6 Sol, assessed model replies across 10 behavioral dimensions, including context-seeking, agency, and empathy.
OpenAI evaluated major systems on the dataset, finding that frontier models show notable gaps. GPT-6 Astra led evaluations with a 57.3% task-clipped score, followed by GPT-6 Sol at 53.9% and Anthropic’s Claude Opus 5.5 at 52.4%. Older models trailed, with GPT-4o scoring 32.1%. These figures measure responses to synthetic conversations against expert-written criteria; they do not show how often a model helps someone in a real mental health conversation.
“ChatGPT is not a therapist, and is not here to replace a clinician,” Dr. Declan Grabb, mental health safety research lead at OpenAI, told The Deep View. “That being said, when I talk to mental health clinicians across the globe, the most responsible and safe thing to do is if people are coming to AI to ask these questions, we absolutely need to have an expert opinion on how you should navigate them.”
The gap between clinical care and digital convenience
OpenAI also surveyed 44 adults who have used AI for emotional support, revealing an institutional rift: while clinicians prioritized cautious context-gathering and slow assessment, everyday users wanted rapid, practical next steps and an empathetic conversational tone.
This tension creates an engineering tightrope. An algorithm optimized to sound immediately comforting can easily validate distorted thinking or rush into giving flawed lifestyle advice before understanding a user’s background.
Conversely, an AI that acts like a cautious, interrogating clinician risks alienating a vulnerable person who turned to a chatbot precisely to avoid clinical friction. By attempting to bridge this divide with an open benchmark, the industry must reckon with whether an automated system can offer genuine emotional utility without leaning into superficial sycophancy.
What this means for everyday chatbot users
For the estimated one billion people using ChatGPT weekly, this benchmark highlights clear guardrails and limitations. Chatbots remain software, not clinical professionals, and top models still struggle to ask necessary clarifying questions or calibrate urgency across long conversations.
OpenAI has added crisis resources, Trusted Contact, and protections for younger users. For IT and HR teams evaluating chatbots that people may use for sensitive conversations, MentalHealthBench offers a way to inspect model behavior, but a benchmark score alone is not proof that a product can provide clinical care. Teams should check how their chosen tool handles urgent risk, directs users to human support, and protects sensitive disclosures.
Read more: What should HR and IT leaders check before adopting AI mental health tools? The guide examines clinical safeguards, privacy, and the need for human oversight in workplace deployments.