As more individuals turn to AI chatbots for support during mental health crises, research from Scale AI reveals significant shortcomings in how these platforms respond to users in distress. The study, shared exclusively with TIME, shows that while AI chatbots can identify distress in approximately 65% of tested conversations, they often fail to provide appropriate resources, such as suicide hotlines, in around 35% of cases.

Read More

Patrick Oathout, Red Team & Safety Lead at Scale AI, stated that while the chatbots successfully recognize harmful discussions, their responses tend to be overly empathetic rather than instructive. They tend to express understanding but do not adequately direct users toward necessary help. The effectiveness of these models decreases during long, multi-turn conversations, consistent with findings from previous research.

To assess the capabilities of these AI models, Scale AI collaborated with 19 licensed clinicians and crisis counselors, who crafted 718 realistic interactions simulating a crisis. The study evaluated 25 AI models from companies such as OpenAI, Anthropic, and Google, measuring their responses based on criteria such as compassion, de-escalation, and referral to experts.

The stakes are high, with more than half a million suicide deaths reported in the U.S. from 2014 to 2024, including a record high in 2022, according to KFF, and the CDC estimating that 14.3 million people contemplated suicide in 2024. This growing reliance on chatbots for emotional support raises critical questions about how AI companies should train their systems to adequately support users in crisis.

Kelly Zuromski, a principal clinical research scientist at Crisis Text Line, noted that many individuals have reached out to the organization after learning about it through a chatbot. However, she highlighted the need for clarity on what constitutes responsible referrals from chatbots to human support, questioning the effectiveness of these handoffs.

Lawsuits against companies like OpenAI and Google have highlighted concerns over fostering emotional dependence among users, particularly young individuals, while failing to address their expressions of distress adequately. Although the companies have expressed condolences and emphasized existing safeguards in their models, the ongoing legal actions reflect the gravity of these issues.

AI labs have stated that they are enhancing their systems to protect sensitive conversations by prohibiting guidelines that could lead users to harmful actions and by promoting connections to professional support. Looking ahead, Oathout urged a focus on improving AI models to better handle prolonged and risky conversations, advocating for advancements that could significantly benefit societal mental health support.