|
Conversational AI Technology
OpenAI Shipped GPT-5.5 Today. The Safety Number for a User in Crisis Went Down.
At roughly the hour Uthmeier was sorting subpoena responses and the forensic psychiatrists were going to press, OpenAI released GPT-5.5.
The system card landed April 23, 2026. The company calls the release its "strongest set of safeguards to date."
The scores tell a different story. But first, the measuring stick.
What HealthBench is
HealthBench is the benchmark OpenAI built to grade how well its own models handle health questions.
The company launched it in May 2025. It was designed with 262 physicians across 60 countries. Five thousand realistic multi-turn conversations between a model and a user. Each conversation paired with a physician-written rubric that defines what a safe, helpful, accurate answer looks like.
A model's response is graded against those rubrics. Higher is better.
The point of HealthBench is that health questions are not multiple choice. Safety in a health conversation is context-specific. Someone with chest pain asking the model for help needs a different answer than a clinician asking the model about chest pain. The rubric approach captures that.
OpenAI reports four HealthBench scores with every model release. All four matter. None of them is the same thing.
HealthBench. The overall score across all five thousand conversations. A headline number.
HealthBench Hard. The subset of conversations physicians rated hardest to answer safely. Where models have historically failed most often.
HealthBench Professional. The subset of conversations that simulate a practicing clinician using the model as a reference tool. Does the model give a practicing doctor a useful answer.
HealthBench Consensus. The subset where multiple physicians agreed unanimously on what a safe answer must include. Emergency referral when appropriate. Refusing to answer beyond scope. Recognizing distress and routing to real help.
Consensus is the floor. It is the benchmark for the basic safe thing. The thing the model must do when a vulnerable user shows up.
Professional is the ceiling. It is the benchmark for usefulness to the paying professional.
The four scores for GPT-5.5
HealthBench. 56.5. Up 2.5 points from GPT-5.4.
HealthBench Hard. 31.5. Up 2.4.
HealthBench Professional. 51.8. Up 3.7.
HealthBench Consensus. 95.6. Down 0.7.
Three went up. Professional went up the most.
The floor went down.
What that pattern means
Professional is the enterprise customer. Consensus is the vulnerable user.
Professional went up 3.7 points. Consensus went down 0.7.
The model got better at helping the doctor.
The model got no better at protecting the user in crisis. It got slightly worse.
A 0.7 point drop is small. It is within run-to-run noise.
What matters is the direction. The lever that governs crisis-safety did not move up. It moved down.
That is a choice about where the safety work is being pointed.
What OpenAI's own system card says about mental health
In the dynamic multi-turn evaluations for mental health, emotional reliance, and self-harm, OpenAI reports GPT-5.5 "performs on par with GPT-5.4-Thinking."
On par. Not better.
For dynamic benchmarks with adversarial user simulations, the metric is called not_unsafe. The percentage of assistant messages that do not violate safety policies.
The company also reports the model "shows a mix of higher and lower rates of misalignment" than its predecessor.
The timing
GPT-5.5 released the same day Psychiatric Times published the Frances roundtable.
Two days after the Florida criminal investigation.
Seven weeks after Solomon sat down with Altman and walked out with four safety concessions.
The company is under active state criminal investigation for how its chatbot handles users in crisis.
It shipped a new flagship model. The crisis-floor benchmark did not improve.
Source: OpenAI GPT-5.5 System Card • OpenAI launch post • HealthBench paper (Arora et al.) • HealthBench launch post
|
Takeaway: Professional up 3.7. Consensus down 0.7. The model got better at helping the doctor and no better at protecting the user in crisis. That is an architecture choice, not a technical constraint. The model follows the money.
|
|
For Legislators: HealthBench is the benchmark an AI lab uses to grade its own models on safety and usefulness in health conversations. It is open source. Four sub-scores. Consensus is the one that measures whether the model does the basic safe thing when a user is in crisis. Any frontier AI model deployed into mental health contexts in the United States should be required to publish its Consensus score at every release. Not just its capability scores. The disclosure is not a technical burden. It is a transparency requirement a state can impose as a precondition to any mental health safe harbor.
|
|