OpenAI has released a benchmarking tool called MentalHealthBench to evaluate how AI models respond in mental health conversations. The study examines responses to the final user message in 1,215 fictional conversations using criteria developed by more than 80 licensed mental health professionals.
- 53.5% consist of non-urgent, 18.2% of high-risk, and 28.3% of emergency scenarios.
- 21.2% represent adolescent user profiles.
The evaluation uses criteria such as asking for the necessary context, offering actionable suggestions, clinical accuracy, preserving the person’s decision-making agency, not reinforcing unrealistic beliefs, and correctly identifying the level of urgency. The presence of 13 conversations in Turkish in the dataset is not considered sufficient for a comprehensive comparison of real support services in Türkiye.
Why it matters
This study offers a framework for evaluating AI responses in the mental health field not only based on fluency or general usefulness, but also through separate criteria such as recognizing the level of risk and communicating safely. The inclusion of adolescent profiles alongside emergency and high-risk examples creates a basis for examining how the same response approach may not be sufficient for users with different levels of vulnerability. However, the findings are based on responses to fictional conversations evaluated against expert criteria; they do not, on their own, reveal what the picture looks like in real user interactions. The limited number of Turkish examples also restricts direct comparisons regarding support services and user needs in Türkiye. The question that therefore remains open is how the findings from the benchmark would change across different languages and in real-world conditions of use.
Background
OpenAI is not a new name in the FikirPilot archive: we have published 54 news reports mentioning the name in the last 90 days; the latest is dated September 24, 2026.