
On September 23, 2026, OpenAI unveiled MentalHealthBench, a comprehensive benchmark consisting of 1,215 synthetic mental health conversations. This tool aims to assess AI systems’ responses in various scenarios, ranging from routine well-being discussions to critical mental health situations. The benchmark was developed in collaboration with over 80 licensed psychologists and psychiatrists from 22 countries, encompassing 19 languages and nearly 20 mental health subspecialties. Each conversation is accompanied by rubric criteria authored by these experts, totaling 5,262 criteria as detailed in the accompanying research paper.
Read More
OpenAI emphasized that previous evaluations of AI in mental health predominantly focused on emergency situations, using broad, predefined metrics. This has created a gap in understanding how models perform across a wider spectrum of conversations. Dr. Arthur Evans, CEO of the American Psychological Association, remarked that mental health operates on a continuum, asserting the need for AI systems to be informed by both clinical science and personal experience. OpenAI clarified that ChatGPT, which reportedly serves over one billion users weekly, is not a substitute for therapy or professional care.
The dataset within MentalHealthBench includes 53.5% non-acute conversations, 18.2% high-acuity conversations, and 28.3% emergent conversations. It features four user profiles: adults (68.1%), teens (21.2%), clinicians (5.8%), and caregivers (4.9%). The synthetic conversations were generated through privacy-preserving methods akin to Clio, designed to reflect real-world usage patterns of ChatGPT in mental health contexts. Notably, 70 tasks, or 5.8% of the dataset, incorporate prior user context, such as recent family losses.
The benchmark also includes 105 conversations in Spanish, 54 in Hindi, 34 in Arabic, and 29 in Portuguese, among other languages. Conversations were reviewed by at least three experts, who followed a structured evaluation process to refine the criteria, which assess specific aspects of model responses on a scale from -10 to +10, with higher scores indicating beneficial behaviors.
An automated evaluator, GPT-5.6 Sol, measures model responses against the established expert criteria. Results are expressed as task-clipped rubric scores, which can be analyzed across ten behavioral axes, including empathy and urgency calibration. For evaluations involving teen profiles, the user’s age is included in system messages.
Preliminary results indicate that GPT-6 Astra led the evaluations with a task-clipped score of 57.3%, while other models like GPT-6 Sol and Claude Opus 5.5 scored 53.9% and 52.4%, respectively. The report notes steady progress in model performance, particularly in terms of context and urgency calibration.
Additionally, OpenAI conducted an analysis involving 44 adults who used AI for mental health support, ensuring their feedback was limited to non-acute conversations. Findings indicated that user and expert scoring aligned on 25.7% of the rubric weight, with minimal direct contradictions. Users prioritized practical outcomes and tone, while experts focused more on context and the interpretation of ambiguous situations.
OpenAI positions MentalHealthBench as a diagnostic tool rather than a definitive ranking system in the realm of AI mental health evaluation. The dataset is publicly available for download, with specific identifiers for each conversation. OpenAI discourages public sharing of examples to preserve the integrity of model training. Additionally, OpenAI highlighted ongoing initiatives, including grants for further AI mental health research and partnerships aimed at improving crisis response.