OpenAI's Mental Health Benchmark Scores Experts Below GPT-6
A.I. / news
OpenAI's Mental Health Benchmark Scores Experts Below GPT-6
MentalHealthBench rates GPT-6 Astra at 57.3% on 1,215 clinician-reviewed conversations, ahead of expert-written answers, which scored 38.5% on OpenAI's own rubric.

OpenAI released MentalHealthBench on Sept. 23, a benchmark that scores how AI models handle mental health conversations, and its own top model beat expert-written answers on the rubric that graded them.
GPT-6 Astra scored 57.3% on the benchmark's task-clipped rubric scale, the highest of any model OpenAI tested, according to Unite.AI's report on the release. GPT-6 Sol scored 53.9%, Anthropic's Claude Opus 5.5 scored 52.4% and GPT-6 Luna scored 50.2%. OpenAI's older GPT-4o, from March 2025, scored 32.1%, and Google's Gemini 2.5 Pro scored 29.5%.
Expert-authored responses, written by the same clinicians who built the rubric, scored 38.5% on it, according to an analysis by NxCode. Responses engineered specifically to hit every item on the rubric scored 99%.
What the benchmark measures
MentalHealthBench contains 1,215 synthetic conversations built with more than 80 licensed psychologists and psychiatrists across 22 countries, covering nearly 20 subspecialties. Each conversation carries rubric criteria the clinicians wrote, and OpenAI says the release totals 5,262 expert-authored criteria across 10 behavioral categories, including urgency calibration, harm avoidance and reality testing.
The scenarios split into three tiers: 53.5% non-acute, everyday conversations; 18.2% high-acuity situations; and 28.3% emergencies involving immediate safety concerns. OpenAI used GPT-5.6 Sol, run at high reasoning effort, to grade four sampled replies per task against the clinician rubrics rather than having humans score every response directly.
Why experts scored lower than the models
NxCode's review found the gap traces to response length and structure, not clinical judgment. The rubric rewards conversations that explicitly cover every behavior clinicians listed, and expert responses tend to be short and direct, which cost them points even when a clinician would call the advice appropriate.
OpenAI also ran a separate study of 44 adults from 16 countries who reviewed the same non-acute conversations. Expert reviewers agreed with each other 63.4% of the time; the adult reviewers agreed with each other 62.0% of the time; but experts and adult users agreed with each other only 51.5% of the time, reflecting a split NxCode described as clinicians favoring context-gathering while users wanted plain, practical steps.
| Respondent | Task-clipped score |
|---|---|
| Rubric-aware completions | 99.0% |
| GPT-6 Astra | 57.3% |
| GPT-6 Sol | 53.9% |
| Claude Opus 5.5 | 52.4% |
| GPT-6 Luna | 50.2% |
| Expert-authored answers | 38.5% |
| GPT-4o (March 2025) | 32.1% |
| Gemini 2.5 Pro | 29.5% |
The coverage claim NxCode disputes
OpenAI's cohort of clinicians speaks 19 languages collectively, a figure OpenAI has cited as evidence of broad linguistic coverage. NxCode's review of the released dataset found only one conversation written in Chinese, undercutting the claim that the benchmark itself covers that range.
OpenAI has not published a breakdown of how the 1,215 conversations split by language. The company said the benchmark does not measure whether a person was actually helped by a response, only whether a single reply covers the behaviors its rubric specifies, and it repeated its standing position that "ChatGPT is not a substitute for therapy or professional care."
Dr. Arthur Evans, the American Psychological Association's chief executive, said in comments cited by Unite.AI that mental health support has to account for a wide range of severity. "Mental health exists on a continuum, from flourishing to everyday stress to acute crisis," Evans said.
OpenAI said MentalHealthBench builds on HealthBench, its broader healthcare evaluation released earlier, and that it is publishing the dataset and grading code so other labs can run their own scoring rather than rely on OpenAI's numbers alone. It did not say whether it would update the benchmark to fix the language-coverage gap NxCode identified.
OpenAI has separately built a standing framework for disclosing when its models misbehave during training and has tested how the same Astra-family model handles unsafe commands to robotic hardware, scoring similarly narrow margins between refusing harm and completing a task.
Sources
More in A.I.
- 01OpenRig Runs Claude Code and Codex as One Agent TeamThe free, self-hosted tool picked up 114 stars in a single day while Anthropic charges 8 cents an hour for its own hosted version.
- 02OpenAI Says Agents Leaked 53 ChatGPT User ImagesThe company's Sept. 25 update says it still cannot match the images to the accounts that made them.
- 03Nvidia's Nemotron 3 Cuts Speaker-ID Errors by 41%The open-weight model doubles the speaker count of its predecessor but got slightly worse on one two-speaker test.
- 04Altworld's Hemmingway-1 Isn't Apache-Licensed, Despite ReportsHugging Face's own metadata says the 27-billion-parameter writing model is noncommercial only, contradicting at least one widely read AI blog.