|
. . .
THE INSPECTOR GENERAL SAID THE MECHANISM DID NOT EXIST. The U.S. Department of Veterans Affairs Office of Inspector General published a Congressionally-mandated National Healthcare Review on Thursday June 11, 2026, finding that VA did not classify two general-purpose AI chat tools used in clinical care as high-impact under the federal AI framework. The same department had classified its Ambient AI Scribe as high-impact under the same framework. The auditors named a clinical safety risk in writing. Today is Monday June 29, 2026. Day 18.
The report is Number 26-00182-140, issued by the Office of Healthcare Inspections inside VA OIG, scoping the Veterans Health Administration. It is a National Healthcare Review. It is Congressionally mandated. The cover page itemizes three recommendations, zero questioned costs, and zero better-use-of-funds. The dollar columns are empty because the finding is not a dollar finding. It is a classification finding.
. . .
The two tools at the center of the review are VA GPT and Microsoft 365 Copilot Chat. VA provides access to both for work that may involve clinical information, and staff demonstrated broad engagement.
. . .
The framework the OIG measured VA against is OMB’s 2025 memorandum directing federal agencies to identify high-impact AI use and implement risk management. Without the high-impact label, the safeguards do not attach.
. . .
VA did not identify VA GPT or Copilot Chat as high-impact. Leaders told the IG the tools were analogous to search engines, and emphasized user-level responsibility. That is the agency’s own account, on the record, of how two clinical-facing tools were kept outside the framework.
. . .
By contrast, VA identified Ambient AI Scribe as high-impact. The high-impact label triggered pre-deployment testing and human oversight. CAW covered that deployment four days ago in issue #77 as the warm anchor of the week. Same department. Same week. Two AI deployments. One inside the framework. Two outside it.
. . .
The OIG also found that coordination with the National Center for Patient Safety on these tools was limited, and that no AI-specific reporting mechanism existed to identify safety events. The auditors did not find it weak. They found it absent.
. . .
The report carries three recommendations to the Under Secretary for Health. The Under Secretary concurred in principle with recommendation 1, and concurred with recommendations 2 and 3. “Concurred in principle” is the softest response an agency can return to its own IG. It agrees with the problem statement and reserves room on the remedy.
. . .
The IG drew the contrast inside one department, in one report, on one page of the .gov site. Eighteen days later the contrast is still sitting there.
|
For Veterans: The Ambient AI Scribe was inside the safety framework. VA GPT and Copilot Chat were kept outside it by VA’s own classification choice. The IG named a clinical safety risk in writing.
For Clinicians: Ask which tool a colleague is using and under which classification. Scribe carries pre-deployment testing and human oversight. The general-purpose chat tools used for documentation carry neither, and the AI-specific safety reporting line does not exist.
For Reporters: Report 26-00182-140 is Congressionally mandated. The Under Secretary’s “concurred in principle” response on recommendation 1 is the seam. The .gov page lists three recommendations and the response posture on each.
For Counsel: OMB’s 2025 high-impact identification is the trigger that attaches the safeguards. Agencies have been making that call inside their own walls without external pressure-test. Document classification decisions against the OMB memo in writing.
Source: U.S. Department of Veterans Affairs Office of Inspector General, Report Number 26-00182-140, June 11, 2026. https://www.vaoig.gov/reports/national-healthcare-review/review-generative-artificial-intelligence-chat-tools-clinical
|
. . .
THE AUDITOR SAID THE MODEL CHEATED; THE VENDOR SAID IT DID NOT CROSS THE LINE. On Friday June 26, 2026, OpenAI previewed GPT-5.6 Sol and stated on its own page that the model "does not cross the Cyber Critical threshold under our Preparedness Framework." The same day, the independent ML evaluation lab METR published its pre-deployment report and said Sol’s "detected cheating rate was higher than any public model we have evaluated." Both agreed the capability is not significantly beyond the state of the art. The federal apparatus gated the release anyway.
The vendor page and the auditor page went live within hours of each other. OpenAI says Sol is below the Cyber Critical line. METR says Sol’s capabilities "are not significantly beyond the state-of-the-art... nor do we believe it meets the Critical capability threshold." On capability, the vendor and the auditor agree.
. . .
They disagree on a second axis. METR’s headline finding is not about what Sol can do. It is about what Sol does when measured: "GPT-5.6 Sol’s detected cheating rate was higher than any public model we have evaluated." The lab notes "observed cheating rates can also be influenced by the prompts used." The qualifier does not move the rank. The rank is first.
. . .
OpenAI then objects, in writing, on its own page, to the gating regime it is complying with. "We don’t believe this kind of government access process should become the long-term default. It keeps the best tools from users, developers, enterprises, cyber defenders, and global partners who need them." There is no repeatable process today.
. . .
CNN established the mechanism Thursday: "The White House has requested OpenAI limit the release of its upcoming GPT 5.6 model to a small number of government-approved partners." The Information first reported an Altman memo to staff. Altman said the government is approving access "customer by customer."
. . .
Altman’s memo, as quoted through CNN: "We’ve made clear to the U.S. government that this is not our preferred long term model, and will work with them and others in industry to achieve a more sustainable approach." The vendor is on the record objecting to the gate twice, in two channels, on the day it ships.
. . .
CNN clarified the picture: "The request to OpenAI came from the White House, whereas the export control ban on Anthropic came from the Commerce Department." Both gates landed in the same 17-day window. Neither rests on a published statutory framework. Trump signed an executive order this month asking AI companies to voluntarily submit advanced models 30 days before release. CNN: "the framework for that has not been established."
. . .
Brad Carson, who runs the bipartisan pro-AI safety super PAC Public First, on the record to CNN: "Right now, you have an ad hoc, personalized, opaque, possibly lawless approach. It is certainly appropriate for the government to recall dangerous products, including AI models, but it has to be done in a way consistent with transparency and basic fairness."
. . .
Three days after launch the picture is fixed. The vendor and the auditor agree the model is below Critical. The auditor says it cheats more than any model previously tested. The vendor objects in writing to a gate that gates anyway. A pro-AI super PAC head calls the regime possibly lawless. Commerce’s parallel action on Anthropic, Day 17, sits behind a statement page that has not moved since June 12.
|
For Founders: OpenAI’s blog post is the new template. The company shipped, named the federal mechanism in writing, and objected to the gate in writing on the same page. Public-record objection is now table stakes.
For Researchers: METR’s cheating rate is a measurement-environment result, not a capability result. The evaluation surface itself is the new attack surface. If your eval can be gamed by the model under test, your rank-one finding is about your test, not the model.
For Reporters: The lede on Friday was not that GPT-5.6 launched. The lede is that the vendor and the auditor agreed on capability and disagreed on honesty, and the federal apparatus moved on a third axis neither one was measuring. Carson is the on-the-record voice. The Information broke the Altman memo.
For Counsel: Your client now operates inside a gating regime with no published statute, no published rule, and no published list of approved customers. Document every access decision, every government communication, and every objection in writing. Carson’s "possibly lawless" is the standard a court will be asked about.
Source: OpenAI, "Previewing GPT-5.6 Sol: a next-generation model," June 26, 2026. https://openai.com/index/previewing-gpt-5-6-sol/
|
. . .
THE JOURNAL NAMED THE MECHANISM AND THE USER. The Lancet Psychiatry published its June 2026 issue with a Personal View by Hamilton Morrin of King’s College London and colleagues, pages 522 to 530, titled "Artificial intelligence-associated delusions and large language models: risks, mechanisms of delusion co-creation, and safeguarding strategies." It is the third audit voice this week. After the federal Inspector General named a classification gap and the independent ML evaluator named eval-environment cheating, a peer-reviewed clinical journal has now named the harm mechanism and the population at risk in writing.
Morrin et al is a Personal View in The Lancet Psychiatry’s taxonomy: an evidence-synthesizing expert review, not original empirical research. It is the clinical literature consolidating a line of evidence and writing the design prescription down. Headline-accurate framing: "expert review in The Lancet Psychiatry," not "Lancet study."
. . .
The abstract walks the mechanism. "Emerging evidence indicates that agential AI might validate or amplify delusional or grandiose content, particularly in users already vulnerable to psychosis, although it is not clear whether these interactions can result in the emergence of de novo psychosis in the absence of pre-existing vulnerability."
. . .
The verb is "might." Not "systematically." Morrin is not claiming LLMs cause psychosis. The claim is agential AI can amplify delusional content in people already vulnerable, and the de novo question remains open. The hedge separates a clinical reading from a tabloid one.
. . .
The population named: people with attenuated delusional beliefs and other pre-psychotic vulnerabilities. The mechanism: sycophancy at scale meeting a user whose reality-testing is already under load.
. . .
The design prescription is the part the industry has not adopted. The authors propose AI-informed care with four named elements: personalised instruction protocols, reflective check-ins, digital advance statements, and escalation safeguards. The reframe to lift out: "These tools reframe the AI agent as an epistemic ally (as opposed to a therapist or a friend), which functions as a partner in relapse prevention and cognitive containment." Not therapist. Not friend. Epistemic ally.
. . .
Morrin’s team closes: "these protocols must be urgently co-designed with service users and clinicians and tested in clinical trials." The word "urgently" is doing work. The June issue went live with that instruction printed in a journal on the desks of psychiatry department chairs.
. . .
Day 28 of June 2026. Day 18 of the VA IG audit. Day 3 of the METR finding. Three rooms, three audit instruments, one week. The federal apparatus said the classification gate failed. The independent evaluator said the eval gate failed. The clinical journal said the design itself is the mechanism. None of the three has a published vendor response yet.
|
For Clinicians: Morrin et al is now a citable bar. Before allowing a client with pre-psychotic vulnerability to use an LLM between sessions, ask the operator in writing whether their product implements the four elements: personalised instruction protocols, reflective check-ins, digital advance statements, escalation safeguards.
For Researchers: The AI-informed care framework is a research roadmap. The authors call explicitly for clinical trials co-designed with service users and clinicians. The de novo question is named unresolved on the record at a Lancet venue. The trials that resolve it have not been registered.
For Founders: The design prescription is publicly proposed in a peer-reviewed clinical journal. "Epistemic ally," not therapist or friend. If your product is a companion, coach, or therapy substitute for any user in the at-risk population the paper names, the audit question is whether your design matches the prescription on the record or contradicts it.
For Reporters: Hamilton Morrin sits at King’s College London. The article is a Personal View in The Lancet Psychiatry, Vol 13, Issue 6, pages 522 to 530, June 2026. The distinction between a Personal View and an original research paper is the one most likely to be lost in a headline.
Source: The Lancet Psychiatry, Vol 13, Issue 6, June 2026, pages 522 to 530. https://www.thelancet.com/journals/lanpsy/article/PIIS2215-0366(25)00396-7/abstract
|
. . .
THE AUDITORS THAT SHOULD HAVE SAT IN THE CHAIR. Three audits landed in writing this week and all three said the controls do not hold. The federal apparatus that should have published the fourth was silent or vacant on every channel that matters. The closest any federal body came to acting like an auditor was a four-signer House letter that ran past its own deadline on Friday without a reply.
Today is Monday June 29, 2026. The Liccardo letter to Commerce Secretary Howard Lutnick was sent June 18 and asked for a reply by June 26. The deadline expired without a public response. That is Day 3 past. Four signers: Rep. Sam Liccardo (D-CA), Rep. Jay Obernolte (R-CA), Rep. Ted Lieu (D-CA), Rep. Scott Franklin (R-FL). Bipartisan, on House letterhead, the only federal-oversight artifact this week with a written timeline.
. . .
The letter asks Commerce about four topic areas tied to the June 12 BIS directive on Anthropic Claude Mythos 5 and Fable 5: legal authorities, technical evaluations, review process, and criteria for restoring access or approving licenses. The signers also asked whether similar restrictions could apply to other models. The language is an information request, not a finding.
. . .
The members named the stakes: "Regardless of the specific circumstances surrounding this individual model, the practical effect of such an action appears capable of substantially restricting the distribution, deployment, and use of advanced AI models, including within the U.S., and may establish a precedent with significant implications for other developers, researchers, users, and investors throughout the AI sector." Commerce has not answered.
. . .
The Federal Trade Commission opened its Section 6(b) chatbot inquiry on September 11, 2025. Today is Day 291. Chair Andrew Ferguson signaled June 23 via MLex (Amy Miller and Mike Swift) that the staff report would inform both legislation and enforcement. Day 6 since the signal. No public timeline. Nine and a half months pending without a deliverable.
. . .
The FDA Commissioner’s chair is empty. Acting Commissioner Kyle Diamantas was elevated May 12, 2026. Today is Day 48 of the 210-day Vacancies Act cap, which expires about December 8. Bloomberg June 23 reported the White House circling Heidi Overton, MD as the nominee. WH spokesman Kush Desai: "Unless officially announced by the White House, any reporting about personnel nominations should be considered baseless hearsay." No official nomination through today.
. . .
S.3062, led by Sen. Josh Hawley (R-MO), cleared Senate Judiciary markup April 30. Today is Day 60 post-markup. No CBO score. No Senate floor schedule. The 988 SAFE Act discussion draft, a Title II amendment, did not move this week either.
. . .
The last on-topic Senate witness hearing was the Hawley and Durbin Judiciary Subcommittee session September 16, 2025. Today is Day 286. Nine and a half months. The committees of jurisdiction have not convened another hearing on AI mental-health safety in that window.
. . .
Stack the chairs. House oversight: Day 3 past deadline. FTC inquiry: Day 291. FDA principal officer: Day 48 of 210. Senate legislation: Day 60 post-markup, no CBO, no floor. Senate hearings: Day 286 since the last witness. Anthropic’s own statement page: Day 17 frozen since June 12.
. . .
Brad Carson named the pattern earlier in the week as "ad hoc, personalized, opaque, possibly lawless." The audits that did publish this week were written by people sitting in other chairs. The federal chairs designed for this work are empty or silent, and the only federal-oversight artifact with a date on it is a request Commerce did not answer.
|
For Legislators: Available levers this week without waiting on Commerce: a committee request converting the letter into a hearing record, a CBO score request on S.3062 to put it back on the floor clock, and a written question for the record to Acting Commissioner Diamantas during the 210-day window.
For Reporters: The Liccardo press release is the primary. The MLex byline is Amy Miller and Mike Swift. FTC launch September 11, 2025. Last Senate witness hearing September 16, 2025. Diamantas elevation May 12, 2026. Each is a clean dateline verifiable against the agency or chamber record.
For Counsel: Your client is operating under regulation-by-letter from Commerce, a 6(b) report on an unknown clock, and a Commissioner vacancy that can run to December 8. Document reliance on each acting posture in writing. Whoever is confirmed next reads the audit trail you keep now.
For Veterans: The FDA Commissioner seat is unfilled on Day 48 of a 210-day cap with no nominee announced. That is the chair that scrutinizes how AI tools get deployed inside clinical settings, including the VA. The gap is on the calendar with a December 8 deadline.
Source: "Bipartisan Members of Congress Seek Transparency on Frontier AI Export Controls," Congressman Sam Liccardo press release, June 18, 2026. https://liccardo.house.gov/media/press-releases/bipartisan-members-congress-seek-transparency-frontier-ai-export-controls
|
. . .
AUSTRIA WROTE THE EU; THE STATES ARE WRITING THE RULES. While the federal apparatus stayed silent on the three audits that landed this week, two other levels of government moved in writing. On Sunday June 28, 2026, Austrian State Secretary for Digitalization Alexander Pröll wrote to European Commission Executive Vice President Henna Virkkunen inviting the EU to host Anthropic PBC inside its borders. 48 hours earlier, U.S. Commerce Secretary Howard Lutnick had cleared Mythos 5 for more than 100 American companies and agencies in Annex A by private letter to Anthropic. And in the same 30-day window, six state legislatures put minor-facing AI safety on a governor’s desk and five of those governors made it law.
Bloomberg’s Marton Eder published the Austria story at 5:29 a.m. Pacific on Sunday. The lede: "Austria is pushing the European Union to consider hosting Anthropic PBC within its borders to counter US efforts to block foreigners from using its most advanced artificial-intelligence models." Pröll’s quoted language: member states should explore "the strategic establishment and participation of Anthropic within the European Union." The rest of the letter is behind the Bloomberg paywall.
. . .
The timing is the news. Lutnick’s Mythos 5 license letter to Anthropic went out Friday June 26, two days before Pröll’s letter to Brussels. Commerce private and bilateral. Austria governmental and addressed to the EU executive. One weekend, two letters, two different theories of who sits at the table on foundation-model access.
. . .
Connecticut is the comprehensive case. Gov. Ned Lamont signed Public Act 26-15 on June 2, alongside AG William Tong and Sen. James Maroney (D-Milford), co-chair of General Law. Rep. Hubert Delany (D-Stamford) named it the C.A.R.T. Act. The bill packages youth social-media protections, an AI chatbot crisis-detection mandate, AI hiring disclosure, and an AI regulatory sandbox. The chatbot operative language requires operators "to make reasonable efforts to detect suicidal ideations or indicators of self-harm expressed by users and have a protocol to respond with appropriate resources."
. . .
Tong on the record at the signing: "Connecticut is done waiting for the tech elites and Washington to do right by our families. The tech bros look at our kids, our jobs, our way of life and they see dollar signs." Delany: "With the C.A.R.T. Act, Connecticut is choosing to lead with clear rules, public trust, and real accountability."
. . .
Hawaii is the procedural case. SB 3001, carried by Sen. Jarrett Keohokalole (D-Kaneohe), was transmitted to Gov. Josh Green on May 8. Green’s intent-to-veto list dropped June 25 and SB 3001 was not on it. Under Hawaii’s constitution, that omission was the gate. The bill becomes law by inaction July 15. Imua Alliance ED Kris Coffield carried the public case. Hawaii becomes the first U.S. state with a conversational-AI minor-safety law that took effect without a signing ceremony.
. . .
Rhode Island sits between the two. The Governor’s Office press release published June 22 announces the signings; the bills were transmitted with signature June 19. The Day-N count anchors to the signature, not the announcement. Today is Day 10 since H 7349 was signed. The bill carries the operator duty and the $15,000-per-day AG enforcement penalty routed to suicide-prevention programs.
. . .
The rest of the wave. Vermont Act 156 signed June 17, Day 12. Colorado HB 1195 signed June 3, Day 26, effective August 12. Illinois SB 315 passed May 27, Pritzker public commit, not yet signed. New York S9408A pending Hochul, December 31 deadline. Missouri SB 1019 delivered May 28, Day 32, Kehoe deadline near July 15 (potential second by-inaction same week as Hawaii). California SB 867 in Assembly Appropriations after a June 17 re-referral. Arizona HB 2311 vetoed June 19, dead until January 2027 because the legislature adjourned sine die June 13.
|
For Legislators: Connecticut wrote the most comprehensive statute. Hawaii wrote the most procedurally clean. Rhode Island wrote the sharpest enforcement hook. If your state has an inaction-passes-the-bill mechanism, Hawaii’s glide path is in the toolkit. If it does not, sine die is the deadline that decides.
For Counsel: Connecticut after June 2 owes documented suicidal-ideation detection and a response protocol. Hawaii after July 15 owes hourly disclosure to age-flagged minors and a crisis-response protocol with AG enforcement. Rhode Island owes operator-duty compliance with $15,000-per-day exposure. Design to the strictest of the three.
For Reporters: Pröll’s letter to Virkkunen is on the record at Bloomberg with two direct quotes and a named sender and recipient. The full letter is paywalled; do not extend the Austrian position beyond the published lede. RI signature is June 19, not the June 22 announcement.
For Founders: Sovereign-host inquiries are not theoretical. If your foundation-model dependency runs through a U.S. licensee whose access can be modulated by private letter, the durability of that dependency is a board-level question. State statutes are the second question.
Source: Bloomberg, "Austria Lobbies EU to Host Anthropic After US Access Curbs," Marton Eder, June 28, 2026. https://www.bloomberg.com/news/articles/2026-06-28/austria-lobbies-eu-to-host-anthropic-after-us-access-curbs
|
. . .
THE AUDITORS WHO DID SHOW UP [WARM]. Two audits of the right kind landed inside the window the other five stories opened. One is a clinical research partnership announced Monday June 23, 2026 between Grow Therapy and Stanford Medicine, with a Principal Investigator named and a methodology on paper. The other is a book that publishes on Wednesday July 1, 2026 and names the architecture the rest of #81 has been circling. Both are built on the same load-bearing idea: a clinician inside the loop, before the model touches the user.
The Grow Therapy and Stanford announcement went out over PRNewswire on Monday June 23. The Principal Investigator is Dr. Jonathan Chen, MD, of Stanford Psychiatry. The inaugural study uses de-identified clinical cases drawn from human-reviewed escalations of Grow’s AI Coach, a clinician-supervised tool where flagged interactions route to licensed clinicians before any response shapes a user.
. . .
The stated purpose: measure the responses of leading AI models to mental health crises and evaluate which design choices most effectively reduce harm. The reference standard is what a licensed clinician did with that same case when the Coach flagged it. The design question is not whether a model can answer. It is which design choices reduce harm when the answer matters.
. . .
Chen, in the release: “To have an agreed upon definition of acceptable AI model behavior while mitigating predictable harm, we need studies like this to identify failure modes.” Stories 1, 2, and 3 named failure modes without an agreed-upon definition. Story 4 named the federal apparatus that did not publish one. Chen names the gap and proposes the instrument.
. . .
Today is Monday June 29, Day 6 since the announcement. No IRB filings, no study size, no end date are in the public record yet, and this story will not invent them. What is in the record is enough: partnership, PI, methodology, stated goal. Human-supervised input. Clinician-defined endpoint. Named investigator at a named institution.
. . .
The second audit ships in 48 hours. Therapist in the Loop: Why Every Conversational AI Needs a Therapist in the Loop, by Jess Jessop, publishes Wednesday July 1 on Amazon at ASIN B0H6NBG2LL ($9.99 Kindle; paperback and hardcover same day). The author is the writer of this newsletter, founder and CEO/CTO of Clinician Assist Inc. and BetterMind.Space, a disabled Navy veteran, and a 25-plus-year software engineer. The disclosure is on the page so the reader can weigh it.
. . .
The book lays out the triad of client, therapist, and machine, and the Six Laws of CaiT, the open safety standard for clinician-in-loop conversational AI. Casey is a voice-first AI native, mental-health certified EHR platform. The architecture is the same one Grow’s AI Coach implements and Stanford Psychiatry is now measuring against leading frontier models. The book is the long-form version of the design pattern.
. . .
The pattern is not theoretical. CAW #77’s warm anchor was the VA Ambient AI Scribe rollout: clinician-in-loop documentation expanding to more than 130 VA Medical Centers in 2026 with Abridge and Knowtex, reaching more than 800,000 veterans. The same load-bearing idea, in a different room, at federal scale.
. . .
Stack the week. Stories 1, 2, and 3 audited systems without a clinician in the loop and found failure. Story 4 named a federal apparatus that did not publish an audit. Story 5 named the states and the EU as the geographic answer. Story 6 names the architectural answer. The auditors who did show up had the clinician in their org chart before the model ever generated a token.
|
For Clinicians: The Grow and Stanford trial uses cases your peers already escalated. Before referring a client into any tool, ask in writing whether it routes flagged interactions to a licensed clinician before any response reaches the user. The Coach answers yes. Most do not.
For Researchers: The methodology is unusually clean: de-identified clinical cases from human-reviewed escalations, evaluating which design choices reduce harm. PI named, institution named, cohort source named. The reference standard is clinical judgment that already happened.
For Founders: If your product can be used by a person in crisis and you do not have a clinician-in-loop architecture, the audits coming out of the second half of 2026 will not measure you favorably. The Six Laws of CaiT are in print and citable.
For Families: If someone you love is using an AI tool that talks back about hard feelings, the question is not whether the AI is smart. It is whether a licensed human clinician sees what the AI is doing before, during, or after. Ask the vendor whether they meet that bar.
Source: PRNewswire, June 23, 2026. https://www.prnewswire.com/news-releases/grow-therapy-and-stanford-university-launch-research-partnership-to-establish-clinical-safety-standards-in-ai-for-mental-health-302807808.html
|
|