OpenAI has launched MentalHealthBench, an open benchmark aimed at assessing how AI systems engage in realistic mental health and emotional support dialogues. This benchmark was developed in collaboration with over 80 licensed psychologists and psychiatrists from 22 countries, who speak 19 languages and cover nearly 20 mental health subspecialties.

Read More

MentalHealthBench goes beyond simply measuring if a model avoids unsafe responses. It evaluates behaviors such as clinical accuracy, context awareness, user agency maintenance, actionable guidance, empathy, urgency recognition, and harm avoidance. Declan Grabb, a safety expert at OpenAI, has described the initiative on LinkedIn as a means to explore AI responses in genuine mental health conversations, ensuring that each synthetic dialogue has been reviewed by three clinicians.

The benchmark, now publicly available, allows researchers and developers to scrutinize its methodology, conduct their evaluations, and build upon the work. The synthetic conversations utilized in MentalHealthBench depict various scenarios involving adults, teenagers aged 13 to 17, caregivers, and clinicians, in multiple languages and regions. Some scenarios also include background information about the synthetic user, enabling researchers to test models' contextual responses.

Of the scenarios, 53.5% are classified as non-acute, addressing everyday stress and emotional issues. High-acuity interactions account for 18.2%, while 28.3% are emergency situations requiring immediate support. More than half of the conversations consist of over five messages, allowing for a comprehensive evaluation of models across dialogues rather than isolated queries. OpenAI emphasizes that this scenario mix is constructed for evaluation purposes and does not reflect the frequency of different types of conversations in ChatGPT, which remains a tool, not a substitute for professional care.

Dr. Arthur Evans, CEO of the American Psychological Association, indicated the importance of assessing AI systems within the broader spectrum of mental health discussions, emphasizing the continuum from everyday stress to acute crisis. He remarked that AI interactions need to be rooted in both clinical science and lived experience.

Mental health experts determined the scoring criteria for model responses in each synthetic conversation. Individual criteria are scored from -10 to +10, rewarding behaviors deemed beneficial while penalizing harmful ones, with larger penalties for significantly detrimental responses. Each conversation underwent review by at least three experts, with criteria retained only when at least two agreed, and the third did not oppose.

OpenAI employs an automated system, GPT-5.6 Sol, to evaluate responses against these expert-defined rubrics. This enables MentalHealthBench to analyze performance across ten areas, including context assessment, actionable guidance, clinical accuracy, and empathy.

Preliminary evaluations indicated that frontier models achieved higher scores, with GPT-6 Astra leading at 57.3%, followed by GPT-6 Sol at 53.9%, and Claude Opus 5.5 at 52.4%. Although older models scored lower, with GPT-4o from March 2025 at 32.1% and Gemini 2.5 Pro at 29.5%, these scores reflect benchmark performance rather than clinical efficacy.

Results can also be analyzed based on conversation urgency, user profile, and specific behaviors, providing insights into differing model performances. OpenAI conducted another study involving 44 participants from 16 countries who had previously engaged with AI for mental health support. This comparison revealed that users focused more on practical steps and tone, while clinicians prioritized context and interpretation of ambiguous situations. However, this user study did not alter the benchmark's scoring criteria, which are firmly based on expert consensus.

MentalHealthBench is now openly accessible for researchers and developers. OpenAI plans to integrate the benchmark into further mental health research and safety evaluations as AI systems advance.