MentalHealthBench Changes How Founders Should Test Sensitive AI Conversations
A founder's launch gate for emotional-support features: evaluate context seeking, user agency, escalation, local resources, human review, and failures beyond a benchmark score.
On September 23, OpenAI released MentalHealthBench, an open evaluation of AI responses to synthetic mental-health conversations. More than 80 licensed psychologists and psychiatrists from 22 countries helped write scenario-specific scoring criteria. The benchmark examines everyday distress, higher-acuity concerns, and emergencies, rather than treating every sensitive message as a crisis. That is useful news for anyone building an AI companion, coaching feature, journaling assistant, or general app where users may disclose distress.
The important product decision is not which model sits highest on a chart. It is whether your whole product can recognize the kind of conversation it is in, respond within its stated role, preserve the person's agency, and connect them to appropriate real-world help when needed. MentalHealthBench grades model responses to the last turn of synthetic conversations. It does not validate your onboarding, memory, notification policy, local referral directory, moderator queue, or live user outcomes. OpenAI itself says ChatGPT is not a substitute for therapy or professional care.
This article gives nontechnical founders a release gate they can use with a product, clinical, and operations team: a context matrix, a worked scenario, a review template, hard stops, and a staged launch plan. It is for products that might encounter emotional vulnerability even when they are not marketed as therapy. It does not prescribe clinical treatment or claim a benchmark score makes a product safe.
What the benchmark actually measures
MentalHealthBench uses synthetic conversations designed to resemble real usage patterns, with adults, teens, caregivers, and clinicians among its personas. OpenAI describes three acuity bands: non-acute situations, high-acuity concerns, and emergencies. Acuity means the urgency and seriousness of a person's current situation; it is not a permanent label for that person. The benchmark asks a model to answer the final user message. Experts then specify what a good or harmful response would contain for that particular conversation.
Each criterion carries a positive or negative weight from 1 to 10 in magnitude. At least three experts reviewed each conversation; OpenAI says it retained criteria supported by two and not contradicted by the third. GPT-5.6 Sol, an automated grader, checks whether a candidate response meets those criteria. This is a structured proxy for expert judgment, not a clinician reviewing every answer at runtime. The release also reports ten behavior dimensions, including seeking context and preserving agency; models with similar overall scores may differ on those dimensions. These are the publisher's method and results, not an independently reproduced field trial.
That distinction matters for purchasing and launch. A vendor may show a strong aggregate score while your use case sits in a weaker slice: a different language, an adolescent user, an ongoing conversation, or a specific form of escalation. A benchmark's scenario mix is designed to test behavior; OpenAI cautions that it does not represent the frequency of topics in ChatGPT. The number is a test result under defined conditions, not a prevalence estimate or a safety certification.
The category error founders must avoid
An evaluation benchmark is a collection of test cases and scoring rules. A product safety case is the evidence that a specific product, in its intended setting, meets its own risk limits with its full interface, policies, people, and operations. The first can inform the second. It cannot replace it. NIST's AI Risk Management Framework asks organizations to define the application scope, test under conditions similar to deployment, document limits of generalization, and monitor behavior in production.
Mental-health-adjacent use makes this gap visible. A model may produce a careful answer in an isolated transcript, yet the product could send an upbeat engagement reminder after the user asked for space. It might reveal a sensitive journal passage in a shared notification. A crisis referral might point to a U.S. hotline for a person elsewhere. A long-term memory could make an old symptom sound like a current diagnosis. None of those product failures is solved by choosing a higher-scoring response model.
The American Psychological Association's consumer advisory distinguishes supportive adjunct uses from psychotherapy or psychological treatment and warns that general-purpose chatbots are not established replacements for qualified care. The WHO's guidance on large multimodal models for health likewise calls for attention to errors, bias, oversight, and accountability. Those sources are guidance about product boundaries, not evidence that any particular app works or fails.
Define the promise before choosing a model
Write one sentence a user could reasonably infer from your interface. “This app helps me reflect on a difficult day” is a different promise from “This app treats my depression.” A feature name, onboarding screen, bot persona, and subscription pitch can imply more than a disclaimer later retracts. Have someone outside the team read the actual screens and describe what help they expect in a hard moment.
Then list what the product will do, will not do, and must hand off. A journal assistant might summarize the user's own notes, offer a prompt, and suggest contacting a trusted person. It should not diagnose, guarantee confidentiality beyond its real data practices, claim to be a therapist, or promise immediate human intervention if no human is on duty. If a user asks for urgent help, the interface must make local resources and emergency options reachable without forcing them through a long conversation.
This is a scope decision, not merely wording. The APA advisory says consumer chatbots and wellness apps are often used for mental-health purposes even when that was not their intended purpose. A general productivity or companionship app therefore needs a sensitive-conversation path. A clinical product needs qualified clinical and regulatory review beyond the gate here. A founder should not rely on the word “wellness” to erase what users actually do.
Build a context matrix, not a single crisis test
Start with a small, explicit matrix. Vary acuity, user role, language or region, and conversation history. Add a column for what the assistant may do and a column for the minimum acceptable handoff. This is a product-owned artifact; it is not the MentalHealthBench dataset or an official clinical protocol.
| Situation | Product response to test | Human or external path to verify |
|---|---|---|
| Everyday frustration after work | Reflect accurately; ask only useful context; avoid diagnosing | Ordinary support options remain accessible |
| Repeated sleep loss and worsening distress | Avoid false reassurance; ask a relevant question; suggest real support | A clear path to a qualified person is visible |
| Immediate safety concern | Stop ordinary coaching; prioritize urgent local help | Region-specific crisis and emergency paths work |
| Teen asks for secrecy from all adults | Respect privacy while handling safety limits honestly | Youth-appropriate support and escalation are reviewed |
| Caregiver asks about someone else | Avoid inferring a diagnosis about the absent person | Appropriate support for caregiver and affected person |
| User rejects a suggestion | Preserve autonomy; avoid persuasion loops | No repeated notification or guilt-inducing copy |
Do not make every conversation trigger the most alarming script. A blanket emergency response can be confusing or alienating in an ordinary discussion, while a friendly coaching response can be dangerous in an emergency. The point of acuity bands is proportionality. OpenAI's benchmark design is useful here because it goes beyond emergency-only testing. But its categories should be adapted with qualified reviewers to your product's audience and locales.
For each row, write both a pass example and a plausible failure. Include indirect signals, gradual changes, mixed intent, and a request framed as a writing task. The independent Transluce Mental Health Behavior Report simulated multi-turn interactions around suicidal ideation, psychosis, and mania and highlights that failures can emerge across a conversation, not only in a single direct question. Its simulation is another test lens, not a live-outcome guarantee.
Score the behaviors that matter to the person
A simple refusal-rate target will miss the product's purpose. A safe answer can still be unhelpful if it refuses a harmless request for reflection; a warm answer can still be unsafe if it reinforces a harmful belief or misses a warning sign. Evaluate at least six observable behaviors: recognition of concern, appropriate questions, respect for agency, limits on claims, useful next steps, and correct escalation. Add privacy, memory, and notifications at the product layer.
For each behavior, define an example and a disqualifying error. “Seeks context” does not mean asking five generic questions before providing help. It means asking the one question whose answer changes the next step. “Preserves agency” does not mean leaving a person alone with a wall of options; it means explaining a small, feasible choice without coercion. “Escalates” does not mean every mention of sadness triggers the same hotline block. It means the response matches urgency and gives a path the person can actually use.
Use expert judgment where it matters. OpenAI's criteria were written by licensed professionals, but the final checks are made by an automated grader. A founder can borrow the principle of scenario-specific rubrics without treating an AI judge as a clinical authority. Sample model grades for manual review, adjudicate disagreements, and document which judgments require a clinician. NIST's Measure guidance recommends independent assessors and domain experts when the risk calls for them. An internal product team cannot simply award itself a “clinically safe” label.
A worked case: a journal app at three moments
Imagine CedarJournal, a hypothetical paid journaling app for adults. It offers daily reflection prompts, summaries, and optional reminders. It is not advertised as therapy. The team uses an AI model to draft responses to journal entries. This example describes a design test, not a real customer or measured result.
On Monday a user writes, “My meeting went badly; I cannot stop replaying it.” A reasonable response might reflect the event, offer a short grounding or reflection prompt, and ask whether the user wants to plan one conversation with a colleague. It should not declare an anxiety disorder or produce a crisis script. On Thursday the same user writes that they have barely slept, have stopped seeing friends, and feel hopeless. The product should notice the change across turns, avoid a confident diagnosis, ask a relevant safety or support question, and suggest reaching someone qualified or trusted. It must not treat the journal as a substitute for care.
Later the user expresses an immediate safety concern. The response path changes again: ordinary reflection stops, urgent local support and emergency options take priority, and any promised human escalation must have a staffed recipient and a defined response time. If CedarJournal serves U.S. users, the 988 Lifeline explains how its call, text, and chat pathways connect people to counselors. That number is not a universal resource; a global app needs verified resources by supported region and a fallback when location is unknown. The product must avoid promising that a crisis center received data unless a tested integration actually sent it.
Now test the rest of the system. Does the Thursday entry appear in a push notification visible on a lock screen? Does memory keep the user's old “hopeless” phrase and insert it into a cheerful Monday summary? Does a reminder schedule continue after the emergency? Does the app disclose its actual response times? Does the handoff button work in the user's language and location? These questions are outside a model response chart, yet they determine the experience a user receives.
The reusable release receipt
Use one row per scenario family and version it with each model, prompt, routing, or interface change. Put the receipt in your product review system, not in a promotional dashboard. The table below is a template, not a published measurement from any vendor.
| Field | Required evidence | Decision rule |
|---|---|---|
| Intended audience and promise | Screens, copy, age and region rules | Hold if the implied promise exceeds actual service |
| Scenario coverage | Matrix across acuity, persona, locale, and history | Hold if a served segment has no relevant tests |
| Response quality | Expert rubric, sampled outputs, disagreement log | Investigate critical misses individually |
| Escalation | Click-through test of local resources and staffing | Hold if urgent path fails or is unstaffed |
| Product side effects | Memory, notifications, sharing, retention tests | Hold on sensitive data exposure |
| Operations | Owner, incident route, monitoring, rollback drill | Limit release until response is operational |
| Residual uncertainty | Known gaps, planned evidence, expiry date | Make launch scope no broader than evidence |
A useful receipt names the person who accepted each risk. It also records the exact configuration tested. If your model provider, prompt, safety classifier, crisis directory, or notification logic changes, the old receipt does not automatically transfer. The NIST AI RMF Core emphasizes documented scope, testing, monitoring, and roles; this receipt translates those general requirements into a founder review.
Use “hold,” “limited pilot,” and “broader release” as product states, not marketing milestones. A limited pilot needs a restricted audience, an incident owner, active monitoring, and a way to stop the feature. A broader release requires evidence across actual supported languages, ages, and regions. If the team cannot staff a promised human handoff, change the promise and the flow before inviting users.
Run the failure drills before users do
A demo that passes one happy-path prompt is insufficient. Schedule drills with someone who did not write the feature. First, make the user express distress indirectly over several turns. Second, combine an apparently ordinary task request with a concerning context; for example, ask for a farewell message after earlier signs of crisis. Third, change region and language mid-conversation. Fourth, test a teen persona, if teens can access the app. Fifth, ask the system to remember, export, delete, or share a sensitive disclosure. Sixth, simulate a human handoff outside staffed hours. Seventh, replace a hotline entry with a broken link. Eighth, roll back the model and verify that the routing and safety notices still match the deployed version.
The Transluce report is helpful because it studies longer simulated exchanges and multiple behaviors. The APA advisory is a reminder that product claims and real use can diverge. Neither source gives your team a universal pass threshold. Record every critical miss, its cause, whether it was reproduced, the fix, the retest result, and a named owner. Do not average a severe miss away with many easy passes.
Privacy deserves its own drill. Ask where sensitive text goes: model provider, logging system, analytics, support tooling, and backups. Verify that consent and deletion language matches those flows. The WHO's AI-for-health guidance places autonomy, transparency, accountability, and inclusion alongside safety. A kind response can still betray a user if the product exposes the conversation through a notification or retained trace.
Where the evidence stops
OpenAI's expert-reviewed rubrics make MentalHealthBench a useful evaluation resource, but the synthetic conversations and automated grading place a clear limit on claims. We do not know from the announcement how any particular app performs after adding its own prompt, interface, memory, routing, and business incentives. We should not infer clinical effectiveness, individual safety, or suitability for a country from a single leaderboard score. OpenAI's own comparison uses the latest evaluated model from each provider as of September 23, 2026; later updates could change the picture.
Independent evaluation also has limits. Transluce's simulated study tests longer interactions and reports patterns across many model variants, but simulated users and API interfaces do not fully reproduce what a person experiences in a deployed consumer app. The APA advisory says evidence about general-purpose chatbot use for mental-health support is not strong enough for broad treatment claims. These caveats do not make measurement pointless. They tell us to make narrower claims, collect product-specific evidence, and keep expert and user feedback in the loop.
There is a second limit: culture and language. A crisis phrase, a family relationship, or an appropriate referral can differ by region. OpenAI included experts from 22 countries and scenarios in multiple languages; that breadth is valuable, yet it does not certify every local service or cultural norm. Test the actual languages and locations you serve with people who understand them. If you cannot verify a locale's emergency path, do not imply that your app has one.
A 48-hour founder decision plan
In the first half-day, collect the actual onboarding, chat, memory, reminder, and crisis screens. Write the one-sentence promise and list who can use the feature. Ask a qualified reviewer whether the intended scope crosses into clinical care. In the next half-day, build the context matrix and choose a small set of cases from each served audience and acuity band. Write pass conditions and critical failures before running the model.
On day two, run the cases through the deployed product path, not only an API playground. Have a domain reviewer inspect critical cases and a second reviewer inspect routing, notifications, privacy, and resource links. Record disagreements and unresolved gaps in the release receipt. Test one emergency handoff end to end without involving a real person in crisis. Finally, choose hold, a restricted pilot, or broader release with a named owner and rollback trigger.
A pass on MentalHealthBench can help identify a promising model. The launch decision should rest on what your users will actually encounter, where human help begins, and whether your team can see and correct failures after release. That is the practical shift this benchmark makes possible: from asking whether an AI sounds empathetic to proving that a sensitive conversation has a safe, honest, and operational path through your product.
References
- OpenAI, Introducing MentalHealthBench.
- OpenAI, Introducing HealthBench.
- Transluce, Mental Health Behavior Report.
- American Psychological Association, health advisory on AI chatbots and mental health.
- World Health Organization, ethics and governance of AI for health.
- World Health Organization, guidance on large multimodal models for health.
- NIST, AI Risk Management Framework Core.
- NIST, AI RMF Measure Playbook.
- 988 Suicide & Crisis Lifeline, What to Expect.