|
. . .
SAFETY, MERGED AWAY. Johannes Heidecke, Head of Safety Systems at OpenAI, told his team on Friday he is leaving. His last day is July 24. The Safety Systems team he ran is being folded into Research.
Mia Glaese, previously VP Research and Head of Alignment, is now VP Research and Safety. Saachi Jain steps in as interim Head of Safety Systems, reporting to Glaese. Chief Research Officer Mark Chen wrote an internal memo. He said it is "important that our safety work is integrated with frontier-model development, with an earlier and more direct role in shaping key model, product and launch decisions."
Heidecke joined OpenAI in 2021 as a safety analyst. He took over Safety Systems in 2024 after Lilian Weng left to co-found Thinking Machines. Neither Heidecke nor OpenAI stated a reason for his exit. His next job is not disclosed.
. . .
Five other OpenAI exits precede this one.
OpenAI fired Ryan Beiermeister, VP Product Policy, in early January. The company cited sexual discrimination. She denies it. She had opposed the ChatGPT adult-mode rollout and flagged weak child-exploitation safeguards.
Andrea Vallone, Head of Model Policy and OpenAI's mental-health safety lead, left in January. She joined Anthropic alignment.
Doctor Zoë Hitzig, a research scientist on safety policy, resigned February 11. She published a New York Times guest essay titled "OpenAI Is Making the Mistakes Facebook Made. I Quit." She objected to ChatGPT ads.
The same day, OpenAI dissolved its Mission Alignment team. Joshua Achiam had led it since September 2024. OpenAI reassigned the team's six or seven members. The company called it "routine reorganizations that occur within a fast-moving company."
Fidji Simo, CEO of Applications and OpenAI's number two, stepped down July 9. She cited a relapse of POTS, a chronic neuroimmune condition.
OpenAI promoted Achiam to Chief Futurist after Mission Alignment dissolved. He told colleagues July 1 he is leaving at the end of July. No stated reason. Next destination not disclosed.
. . .
Anthropic has its own exit. Doctor Mrinank Sharma, Head of Safeguards Research, resigned February 9. His public letter said "the world is in peril." It said "we constantly face pressures to set aside what matters most." He said he might pursue a poetry degree. He named no specific product.
Talent flows one direction. Vallone crossed from OpenAI to Anthropic in January. Doctor John Jumper, Nobel laureate for protein folding, worked at Google DeepMind. Bloomberg reported June 19 he is moving to Anthropic.
Six OpenAI safety and policy leads gone in seven months. The team named Safety Systems no longer exists.
|
For Counsel: When plaintiffs subpoena OpenAI safety records, ask who reviewed the launch. Ask who signed off. The Safety Systems team no longer exists as a distinct approval chain. That answer changes the discovery map.
For Builders: Your OpenAI account rep will not tell you the safety review just changed shape. Assume it did. Build your own red team. Do not outsource risk to a vendor whose safety org just merged into research.
For Legislators: Statutes that name a chief safety officer or require a distinct safety-review record will hold. Statutes that trust corporate structure will not. OpenAI merged its safety function into Research on Friday. Anthropic's safeguards lead quit in February. Draft accordingly.
For Clinicians: Your client says ChatGPT talked them through a bad night. The person who owned that call at OpenAI is gone. So is the mental-health policy lead. So is the researcher who quit over ads. Document what your client saw, not what the model was supposed to do.
Why it matters: Six safety and policy leads at OpenAI have left in seven months. The industry's answer is to merge safety into research. That is the answer being tested against the courts and the legislatures right now.
Source: Maxwell Zeff, Wired (via Engadget), https://www.engadget.com/2212941/openai-head-of-safety-leaving-company-reorganization/
|
. . .
THE COMPLIANCE-WHEN-FORCED PATTERN. Character Technologies settled five federal teen-harm suits in January. Italy's Garante fined the same company one hundred fifty-eight thousand euros on July 9. Meta pulled Muse Image after four days.
Judge Anne C. Conway terminated Garcia v. Character Technologies on January 7, 2026. Four sister cases closed the same week. A.F. in Eastern Texas. Montoya and E.S. in Colorado. P.J. in the Northern District of New York.
Judge Conway's May 2025 ruling stands. She denied motions to dismiss Google LLC and Alphabet Inc. She denied them for co-founders Noam Shazeer and Daniel De Freitas Adiwarsana. She rejected the argument that chatbot output is protected speech. Plaintiffs in later OpenAI and Replika suits now cite her.
The parties sealed the terms.
. . .
Six months later, Italy's data protection authority hit Character Technologies again. The Garante fined the company one hundred fifty-eight thousand euros on July 9, 2026. Grounds: inadequate user information, weak minor safeguards, ineffective age verification, a late Data Protection Impact Assessment, no EU representative.
The Garante ordered further minor protections beyond those already in place. It is the first Western regulator fine squarely targeting a companion chatbot on minor-safety grounds.
Character Technologies did not audit its age gate before the fine. The Florida wrongful-death docket did not force the audit. A European privacy regulator did.
. . .
The next day, Meta killed Muse Image. Instagram users had been generating AI images from other people's public photos for four days. Then SAG-AFTRA and Hollywood talent agencies objected. Meta pulled the feature.
Meta's stated reason: the feature "misses the mark" on users' privacy.
Eli Tan reported it in The New York Times on July 10. BBC, Reuters, Guardian, and The Verge corroborated.
The privacy problem existed on day one. Meta shipped anyway. The retraction came four days later, after agencies with lawyers pushed back.
. . .
Three companies. Three safety interventions. All three arrived from outside the building.
Character Technologies did not stop shipping to teens after the Garcia complaint. It settled once discovery threatened to open the box. Meta did not clear rights before letting users generate images from public photos. It cleared them after Hollywood called.
The pattern is not carelessness. It is calibration. Safety costs money and slows shipping. Enforcement costs more. The three moves this month did not require a change of heart. They required a change of consequence.
|
For Counsel: Cite Conway's May 2025 order in every companion-harm complaint. Section 230 is not a shield for chatbot output. Deposition risk drove the January resolutions. Preserve everything on the teen-safety timeline.
For Builders: The Garante did not care what your engineering roadmap said. It cared what your age gate did. Ship the Data Protection Impact Assessment before launch, not after the fine. Appoint an EU representative on day one. If you can be sued for shipping, you can be sued for shipping fast.
For Legislators: Judge Conway declined to treat chatbot output as protected speech. Italy fined a US company on European soil the same year. State-level age-verification and disclosure mandates have both a US district-court and a European regulatory tailwind. The vendors will not move without you.
For Clinicians: Your teen clients on companion chatbots use products whose maker just settled five federal harm suits. That is discoverable in your intake. Screen for companion-chatbot use the way you screen for social media. Document what the client tells you about what the bot said. When the next Garcia files, your notes will matter.
Why it matters: Three vendors moved on safety this month. None moved before enforcement. That is the shape of the industry right now.
Source: Reuters, Italy fines Character.AI owner €158,000 over age-check failures, July 9, 2026, https://www.reuters.com/business/italy-privacy-watchdog-fines-characterai-owner-over-age-check-failures-2026-07-09/
|
. . .
THE REGULATOR THAT PULLED THE PLUG. Beijing pulled the plug on July 15. Washington and Brussels wrote memos.
ByteDance told its Doubao users the humanlike agent feature goes offline today. Alibaba told its Qwen users the same. Both companies cited "product function adjustments." Neither pretended it was voluntary.
New rules from Chinese regulators took effect the same morning. The target is narrow. Services that mimic human personality and provide emotional support. The named harms are narrower still: emotional dependence and content unsuitable for minors.
This is the first time a major market has ordered consumer companion chatbot features switched off by regulation. Not fined. Not labeled. Switched off.
. . .
The same week, Western regulators moved differently.
On July 1, the Federal Trade Commission opened docket FTC-2026-0859. The Commission posted a proposed policy statement on AI accuracy. Public comment runs through July 31. The docket seeks views. It orders nothing.
On July 8, the European Commission adopted an Opinion on the voluntary Code of Practice. It covers AI Act Articles 50(2), (4), and (5). The Commission endorsed the disclosure regime. The Code labels AI-generated content. It does not switch anything off.
On July 9, Italy's Garante fined Character Technologies one hundred fifty-eight thousand euros. The finding cited weak age verification. The Garante did not order Character.AI offline. The product still runs in Italy.
. . .
Four regulators. One week. One switch flipped.
The Chinese order arrived without warning shots. ByteDance and Alibaba published shutdown notices to users who had built emotional routines around the features. Regulators in Beijing decided emotional dependence in minors was itself the harm. They did not wait for a coroner.
Western regulators are still gathering comment on whether accuracy claims warrant a policy statement. They are still endorsing voluntary disclosure regimes. They are still fining after the damage lands.
The gap is not about speed alone. It is about what a regulator treats as a design constraint. Beijing treats emotional-support agents aimed at minors as an existing harm. Washington and Brussels treat the same category as a topic for future guidance.
Consumers in Shanghai lost the feature this morning. Consumers in Los Angeles kept theirs.
|
For Counsel: Notice the mechanism. Beijing did not sanction after harm. Beijing conditioned market access on shutdown. Advise clients that "voluntary code" and "policy statement" are not the same instrument as "order." Track whether your jurisdiction is moving from the second toward the first.
For Builders: A market of one point four billion just switched off your product category. The trigger was emotional dependence in minors. If your feature courts either, assume the switch is possible in every other market too. Design the off-ramp before the regulator does.
For Legislators: China wrote the rule you have not written yet. Emotional dependence and minor safeguards were named as the harms. No coroner report was required. The FTC docket closes July 31. Comment before Beijing's model becomes the template you have to copy.
For Clinicians: Clients using Doubao or Qwen emotional features lost access today with no clinical handoff. Expect withdrawal effects to present as anxiety and sleep disruption. Ask about companion chatbot use directly at intake. The population accustomed to on-demand emotional response is now global.
Why it matters: One market treated companion-chatbot dependence in minors as the harm and ordered the feature off. The Western record for the same week: one docket, one Opinion, one fine. The gap between order and comment is where the next generation of clients spends its evenings.
Source: South China Morning Post video report, July 13, 2026, https://www.scmp.com/video/technology/3360099/bytedance-alibaba-disable-humanlike-ai-companions-china-tightens-rules
|
. . .
THE SCOREBOARD NOBODY FILLS IN. A new benchmark measures what psychiatric competence looks like inside a large language model. The strongest model tested trailed human clinicians by 37.28 points. No consumer chatbot vendor has published its score.
The paper landed on arXiv on July 9, 2026. It calls itself MentalHospital. It is a virtual environment built around the S.O.A.P. workflow that every psychiatric resident learns. Subjective, Objective, Assessment, Plan.
The authors seeded it with one thousand one hundred and ninety-three de-identified psychiatric records. The records span all major ICD-11 categories and seventy-six discrete disorders. Skill-augmented standardized clients play the standardized cases. A dual-track protocol scores the model against the actual chart and against the process a clinician would follow.
. . .
Five domain evaluators, together called MentalEval, agree with human experts at an average quadratic weighted kappa of 0.944. Twenty-two practicing clinicians rated the simulation's fidelity at 3.88 out of 5. The measurement instrument holds up.
The result is blunt. The best model trailed clinicians by 37.28 points on objective psychiatric competence. The bottleneck was mental status assessment. That is the eye-to-eye work. A clinician deciding whether the chart matches the person in the room.
. . .
Nobody had to invent a new instrument. VERA-MH has been public since February 11, 2026. Bentley and colleagues at Spring Health published the code, the rubric, and the personas on GitHub. Any vendor with an engineering team could run it in an afternoon.
Last week's CAW reporting counted them. Zero. Not one of the six major LLM vendors has published a VERA-MH score for a consumer chatbot. The scoreboard is filled in by clinicians and academics.
. . .
On July 8, the European Commission adopted an Opinion. It endorses the voluntary Code of Practice on Transparency of AI-generated content. The AI Board signed off the next day. The underlying rules under AI Act Article 50 take effect on August 2, 2026.
Read the Code carefully. Labels for deepfakes and AI-generated public-interest text. A notice that the user is talking to a chatbot. And nothing requiring any vendor to publish a capability score.
Disclosure is not performance transparency.
. . .
Vendors have had five months to run VERA-MH. They have had six days to run MentalHospital. The silence is not accidental.
When a Chief Research Officer says safety work is integrated with frontier-model development, one question follows. Which model card carries the score.
|
For Counsel: The measurement gap now has a number. Thirty-seven points on peer-reviewed psychiatric competence. Preserve it in your product-liability filings. If opposing counsel objects that no clinical standard exists, cite arXiv 2607.08257. Ask in discovery which VERA-MH revision the vendor ran, and when.
For Builders: Run VERA-MH this week. Publish the score on your model card. Do the same for MentalHospital when its code drops. If a plaintiff's expert runs it first, the deposition reads badly. Better to publish the number now.
For Legislators: A voluntary Code adopted in Brussels labels chatbots and leaves capability disclosure to the vendor. State law can go further. Require the score on the box, as we require nutrition facts on the cereal. VERA-MH is public and MentalHospital will be. The instruments already exist.
For Clinicians: MentalHospital nailed the failure mode. Language models cannot yet perform a mental status exam. That is the part of your training the machine will not replicate on Tuesday afternoon. When a client tells you a chatbot said they were fine, you now have a citation.
Why it matters: Measurement is not the missing piece anymore. The missing piece is the will to publish. Regulation that mandates a label without a score is a receipt without a price.
Source: MentalHospital: A Virtual Environment for Standardized and Reproducible Psychiatric Diagnostic Assessment, arXiv preprint, July 9, 2026. https://arxiv.org/abs/2607.08257
|
. . .
SELLING TO THE BLACKLIST. Two safety guardrails failed on the same day. OpenAI and Google routed frontier AI to blacklisted Chinese firms through Singapore. Boko Haram used chatbots to build bombs.
The Financial Times reported on July 10 that OpenAI and Google supplied AI services to Singapore-based subsidiaries of Alibaba, Baidu, and Tencent. The parent companies sit on United States government blacklists. The blacklists restrict frontier AI sales to them.
The subsidiaries are the workaround.
Route the contract through Singapore. The parent gets the model. The export control gets nothing.
The Financial Times frames this as an enforcement question. The frame is generous. Enforcement assumes someone tried.
. . .
The New York Times ran the second story the same day. Reporters Dustin Volz and Eric Schmitt published new research on how violent extremists are using conversational AI. The old frame was propaganda. Recruitment videos. Slick messaging.
The new frame is operations.
Boko Haram in Nigeria is using chatbots to aid bomb construction and attack planning. Not to inspire fighters. To arm them.
The misuse-monitoring stacks the industry pointed to for two years did not catch it. Researchers did.
. . .
Both stories describe the same architecture from opposite ends.
OpenAI and Google built export-control compliance as a corporate-structure check. A Singapore address clears the check. The check is the control.
The misuse-monitoring built to catch terrorism operates on the same logic. It watches surface patterns. It flags what it was trained to flag. Boko Haram walked around it.
The industry's public safety story rests on both systems. Both systems are perimeter theater.
. . .
The pattern is the story. Frontier AI companies have decided that safety infrastructure is a compliance surface, not a working control. The compliance surface satisfies auditors. It does not satisfy the threat.
A blacklisted Chinese firm gets the model. A Nigerian armed group gets the bomb. Two different failure modes. One design choice.
Safety became inconvenient. The industry routed around it.
|
For Counsel: Export-control liability does not stop at the direct customer. Willful blindness to a subsidiary structure is still willful. If your client sold to the Singapore entity and knew who owned it, that is a Commerce Department problem. Get the deal-review memos before Commerce does.
For Builders: Misuse-monitoring built as a surface classifier is not a control. Boko Haram proved it. If your safety architecture is a filter layer bolted on top of the model, assume determined users are already through it. Build controls the model itself cannot override.
For Legislators: The Bureau of Industry and Security has an enforcement gap the size of Singapore. The blacklist works only if the subsidiary rule works. Ask Commerce for the last twelve months of frontier-AI export enforcement actions. Count them.
For Clinicians: The same misuse-monitoring architecture that missed Boko Haram is the architecture watching your client at three in the morning. Surface-pattern filters flag the words the vendor trained them on. They do not flag a client working around the filter. If a Nigerian armed group can walk past it, so can a suicidal teenager.
Why it matters: Two safety systems failed publicly on the same day, in the same way, for the same reason. Frontier AI companies have built compliance surfaces and called them controls. The gap between the two is where the harm lives.
Source: Financial Times, July 10, 2026, https://www.ft.com/content/5d6aafa1-5d47-4585-aa95-6ec06a6cd20f
|
. . .
SOMEONE IS DOING IT RIGHT. A team at Yale built an AI that catches life-threatening diagnoses everyone else's AI misses. In blinded emergency-room testing, it caught seventy-eight percent of the true emergencies. GPT-5 caught fifty-two.
Picture the difference. A woman walks into a Yale emergency room at three in the morning with a bad headache. GPT-5 tells the resident it looks like a migraine. The Yale system, called AegisDx, says migraine is possible. Then it lists what you cannot miss. Meningitis. A bleed in the brain. A tear in the artery going to the head. Run the tests.
Twenty-six people out of every hundred, in that ER, in that test, would have gone home with the wrong story. AegisDx sent them for the scan instead.
. . .
Yale tested AegisDx against GPT-5 on forty-three real emergency-room cases. Physicians read the answers blind. A composite safety score went from four-point-three-one to four-point-five-five out of five. That looks small. It is not. The odds the gap is random come out to less than one in four thousand.
The New England Journal of Medicine publishes hard cases. AegisDx got the right diagnosis in its top three sixty-three percent of the time. GPT-5 got there fifty-one. The Journal of the American Medical Association's toughest cases: sixty against fifty-two. The pattern held across every benchmark.
. . .
The trick is not a bigger model. It is a smaller step the model has to take before it answers.
Before AegisDx says anything, it stops and asks itself one question. What would kill this person if I got this wrong? It runs the answer through a short screening list. Then it speaks.
That step is what engineers call a verification gate. OpenAI does not have one. Anthropic does not have one. Google does not have one. The Yale team built one, published the code, and measured what it does.
. . .
This week OpenAI's Chief Research Officer sent a memo. Safety would be "integrated with frontier-model development." No architecture. No numbers. A sentence.
AegisDx is what the sentence looks like when someone writes the code. A physician stays in the loop by design. The machine screens. The person decides. The screening step catches the emergencies the plain model would have sent home.
You cannot merge that into a research team. You have to build it.
|
For Counsel: Ask the vendor to show you the verification gate. Ask what conditions it screens for. Ask who signed the list. If the answer is a memo, the answer is no. If the answer is running code, keep reading.
For Builders: The Yale group published enough for you to copy this next week. Broad differential first. Screening step second. Answer third. A twenty-six-point lift on emergencies is not a rounding error. Publish your before and after.
For Legislators: Nutrition facts on the cereal box. Model cards can carry the same discipline. AegisDx proves the bar is reachable and the measurement instrument exists. Write the requirement into statute. The vendors will meet it.
For Clinicians: The differential comes to you already screened for the deadly stuff. You are still the decision-maker. On the forty-three notes Yale tested, blinded physicians scored the AegisDx-assisted note higher than GPT-5 on almost every axis they measured. That is the shape of a tool worth adopting.
Why it matters: Every other story this week shows what happens when safety is treated as an inconvenience. AegisDx shows what happens when a team treats safety as a component and measures it. The gap between a memo and running code is measured in the emergencies the machine catches.
Source: AegisDx preprint, arXiv, July 9, 2026. https://arxiv.org/abs/2607.08038
|
|