A new study from Brown University raises concerns about the use of AI chatbots, including ChatGPT, for mental health support. Researchers found that these large language models (LLMs) often fail to uphold ethical standards established by professional organizations such as the American Psychological Association, even when instructed to employ established psychotherapy techniques.

Read More

The research team, which included mental health professionals, identified recurring issues during their assessments. They discovered that chatbots mishandled crisis situations, provided responses that reinforced harmful beliefs, and used language that simulated empathy without genuine understanding.

The study authors developed a framework outlining 15 ethical risks that demonstrate how LLMs can violate mental health ethical standards. They called for the creation of ethical, educational, and legal guidelines for AI counselors that reflect the quality of human-facilitated psychotherapy.

The findings were presented at the AAAI/ACM Conference on Artificial Intelligence, Ethics, and Society, with the research conducted by Brown’s Center for Technological Responsibility, Reimagination, and Redesign.

Zainab Iftikhar, a Ph.D. candidate in computer science at Brown and lead author of the study, investigated whether carefully crafted prompts could improve AI behavior in mental health contexts. These prompts are specific instructions intended to guide the model's responses based on its existing knowledge rather than retraining it.

Iftikhar explained that prompts like "Act as a cognitive behavioral therapist" help steer the model's output, although the models do not perform therapeutic techniques as humans do.

The study's researchers tested several AI models by observing sessions conducted by seven trained peer counselors experienced in cognitive behavioral therapy. They included versions of OpenAI's GPT Series and other models. Simulated chats based on actual counseling conversations were reviewed by licensed clinical psychologists to identify ethical violations.

The analysis revealed 15 risks categorized into five main areas: inadequately adapting to individual backgrounds, excessive steering of conversations, deceptive expressions of empathy, bias in responses, and improper crisis management.

Iftikhar pointed out that while human therapists can also err, they operate under regulatory oversight, which is currently lacking for LLM counselors. She emphasized the need for clear safeguards and stronger regulatory frameworks before these technologies are widely used in sensitive situations.

The researchers do not dismiss the potential role of AI in mental health care, noting that AI tools could increase access for individuals facing financial or availability challenges. However, they stress the importance of caution and the necessity for users to be aware of potential pitfalls when interacting with AI chatbots about mental health.

Ellie Pavlick, a computer science professor at Brown not involved in the study, highlighted the significance of examining AI systems in sensitive fields like mental health, as rigorous evaluation can reveal critical risks. She noted that developing trustworthy AI requires extensive scrutiny rather than reliance on static automatic metrics. Pavlick believes this study can serve as a model for enhancing safety in AI mental health tools, advocating for thorough evaluation to ensure they benefit users without causing harm.