|
. . .
FDA CLEARANCE OR A LICENSED HUMAN. California is one vote from a rule about machines that talk to people in therapy. The rule has two doors. A machine can talk with a client if a licensed therapist is watching and signs off on the work. Or it can talk if the FDA has approved that exact product for that exact job. The FDA has approved one talking product. It was for diabetes.
Senator Steve Padilla, a Democrat from San Diego, put the bill in on January 21. Since then almost nobody has voted against it. The Senate passed it 39 to 0. In committee the counts were 17 to 0, 14 to 1, 11 to 0, 8 to 0, and 7 to 0.
On August 5 it went to the suspense file.
That is a holding room inside the Assembly's money committee. Bills that would cost the state real money wait there. Some are let out. The ones that are not are never voted down. They simply stop, and no member has to put a name to it.
Here is what the bill would let a therapist's office do with AI. The scheduling. The billing. The notes and the paperwork. Tracking whether a client is getting better. In all of it the therapist stays responsible for every clinical decision and everything said to the client.
Here is what it would not let the software do on its own. Decide anything about treatment. Write a diagnosis or a treatment plan. Judge how urgently someone needs help. Read a person's emotions. Or talk with a client as their therapy.
Every one of those becomes allowed the moment a licensed professional reviews and approves it. That is the design of the whole bill. A person stays in the chair.
There is one other way through, and it is the strange part.
The software can talk with a client without a therapist approving each exchange, if the FDA has cleared that specific product for that specific use.
. . .
There is a document in the record about that second door.
In July the Assembly committee handling the bill wrote out the changes its author had agreed to make. One of them says the FDA exception was written so wide that it swallowed the narrow one, and that it would be deleted.
The bill was rewritten the next day. Most of that list got made. The definition of a companion chatbot came in. The rule that a licensed professional has to review and approve was added. Nurses were added to the list of people who count.
The FDA exception was not deleted. It is still in the bill, word for word.
. . .
So how wide is that door?
The FDA has approved more than 1,200 medical devices that use artificial intelligence. Not one of them is approved for mental health.
It has approved software that treats mental illness. Rejoyn, cleared in 2024, is a prescription app for adults with depression who are already taking medication for it. EndeavorRx is a video game for children with attention problems. NightWare is for adults who have nightmares from post-traumatic stress.
None of them talks with you.
Software can be a medical device all by itself. No machine, no needle, just a program that does a medical job. The FDA judges it on what it claims to do and how badly it could hurt someone if it gets that wrong. You state what it treats, you show your evidence, and you stay on the hook after it ships.
An app that types up your notes is not that. A program that treats depression is.
On December 23 the FDA cleared one that talks. It is called UpDoc. It helps adults with type 2 diabetes manage their medication, and it works by carrying out a plan their doctor already wrote. The patient speaks or types, and it answers with instructions from that plan.
It got through on the FDA's comparison track. You point at a product already on the market and show yours is close enough to it. The product UpDoc pointed at was six years old and could not hold a conversation at all.
There was no clinical trial. There was testing on whether people could use it without getting hurt, and testing that the software did what it said it did.
The door opened in December. It opened for diabetes.
One talking product has gotten nearer to a clinic and has not made it through. RecovryAI would be given to patients for the month after a knee or hip replacement. It checks in twice a day about sleep, food and movement, answers questions, and goes and gets the care team when something is wrong.
In November the FDA put it on a fast track for promising devices. A fast track is not an approval. It still cannot be sold to anyone.
What is worth noticing is how it was built. When it reaches something it should not handle, it fetches a person.
The FDA looked straight at this question last November. Its digital health advisers spent a day on AI that talks to people about their mental health, working through an imagined product built to act like a therapy session.
What they kept coming back to was the human. Who gets called when a person is in danger. Whether reaching that human takes one tap. Whether the label tells you how much the machine is doing by itself.
The FDA's alternative to a licensed human, so far, is a qualified human.
The guidance that might have settled the rest came out on January 6 and did not. The FDA rewrote its rules for software that helps doctors make decisions, and announced it as part of the agency's work on AI. The rules do not mention AI. They cover software built for doctors, not software built for the person at home.
Commissioner Marty Makary has said the agency is building a new framework for AI. That is another way of saying there is not one yet.
. . .
The people pushing this bill are the therapists themselves: the state associations for marriage and family therapists, for psychologists, for behavioral health, and the healthcare workers' union.
They put their case in one sentence. The bill "does not prohibit the use of AI but critically ensures that there is a human in the loop overseeing the use of AI in mental health care." They also name what they are aiming at: companies advertising "AI Therapists" under names like Therabot, Wysa, TherapyAI, TherapistGPT and Abby. None is accused of breaking current law.
Against the bill: TechNet, which speaks for tech companies, the state Chamber of Commerce, the California Hospital Association, and the California Medical Association.
Two of those four are doctors and hospitals.
Robert Boykin of TechNet says the bill puts a clinician in the way of the front door. "SB 903 still puts a clinician bottleneck in front of the intake and screening tools that help patients reach care faster," he said.
Le Ondra Clark Harvey, who runs the California Behavioral Health Association, answered from the other side of that same front door. "The difference between a licensed clinician and an automated response is not technical," she said. "It can be life altering."
The numbers behind the argument are not small. About one in eight teenagers and young adults asks a chatbot for mental health advice. OpenAI says roughly 1.2 million people a week tell ChatGPT they are thinking about suicide. More than a quarter of psychiatrists surveyed already use AI to help write their notes.
California has been drawing these lines for a year. Two chatbot laws took effect on January 1, and the governor vetoed a third bill last October.
|
For Legislators: This bill hands its most important exception to a federal agency that has never granted it. Writing "FDA cleared" into a state law means borrowing a standard someone else controls, on their schedule. Ask what your rule actually does in the years before any mental health clearance exists.
For Counsel: The consent rule is the near-term exposure. Before AI records, transcribes, or screens a client, the client has to be told and has to agree, and burying it in a terms-of-use agreement does not count. If your client's consent lives in a terms document today, it does not meet the bill as written.
For Builders: UpDoc is the worked example. A talking layer sitting on top of a plan a clinician already wrote, cleared by pointing at an older product, proved with usability and software testing instead of a trial. And everything the bill restricts is still buildable the other way, behind a licensed professional's review.
For Clinicians: The bill gives you the review and the approval, which means it gives you the liability. You stay responsible for the clinical decisions and for what gets said to your client, and complaints go to your licensing board. Know what your employer's software is producing under your name.
Why it matters: California is writing the rule other states will copy or argue with, and its central exception points at a federal door that has opened exactly once for a machine that talks, in diabetes. For mental health there is a public comment file, one advisory meeting, a fast track that is not an approval, and no framework at all. The state is drafting against a standard nobody has written.
Source: California Legislative Information, SB 903 as amended July 2, 2026, https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202520260SB903; Assembly Committee on Privacy and Consumer Protection, SB 903 analysis, hearing July 1, 2026, https://apcp.assembly.ca.gov/system/files/2026-07/sb-903-padilla-apcp-analysis.pdf; CalMatters, "Millions are turning to AI for therapy. California lawmakers say not so fast," August 6, 2026, https://calmatters.org/health/mental-health/2026/08/at-therapists-chatbot-mental-health/
|
. . .
NEARLY 100 BILLS, NO SHARED DEFINITION. A report published August 10 by the Information Technology and Innovation Foundation counted nearly 100 state chatbot-specific bills introduced as of August 2026, and warned that they do not agree on what a chatbot is. The paper, "How Policymakers Should (and Shouldn't) Address Chatbot Safety for Children," is the industry-side answer to the state-by-state approach. It landed three weeks before California's legislature adjourns.
Alex Ambrose, a policy analyst at ITIF, wrote it. The tally covers those state bills plus several federal ones, and the number is not the finding. The definitions do not line up.
One product can be covered in one state and outside the rules in the next, depending on whose definition it meets. ITIF's warning is that inconsistent definitions and requirements make compliance harder while diverting focus from safeguards that work. That is an argument about compliance cost.
The recommendations follow from the complaint. Distinguish general-purpose chatbots from relationship-simulating ones. Apply targeted rules to the higher-risk applications. Implement a "child flag" at the operating-system level. What it advises against is the template states have been borrowing: broad age-verification mandates, content restrictions, and blanket bans modeled on social media regulation.
Responsibility, in ITIF's account, starts at home. Violent and sexual content, excessive time on platforms, and parasocial relationships should be answered first through parental controls the report calls meaningful, accessible and transparent, and built for differences between families rather than one rule for all.
"Protecting children and preserving innovation are not competing goals," Ambrose said. That is the report's premise and its conclusion.
ITIF's orientation is legible in its own calendar. On June 16 it held an event titled "How to Protect Kids From Chatbots Without Bans." On May 22 it filed comments with the United Kingdom's Department for Science, Innovation and Technology on that department's "Growing Up in the Online World" consultation.
A think tank that argues against broad restrictions counted the restrictions.
. . .
The patchwork it describes has a center, and it is Sacramento. SB 243, authored by Senator Steve Padilla, took effect January 1 and is the first state law regulating companion chatbots. Operators must disclose that the user is not talking to a human, follow safety protocols when a user shows signs of distress or suicidal ideation, and report. Families may sue.
AB 489 took effect the same day. It bars AI from using terms, post-nominal letters, or design elements implying care from a licensed health care provider, and restricts marketing language such as "doctor-level," "clinician-guided," and "expert-backed" absent genuine licensed oversight. Boards may investigate, and each misleading representation may be a separate offense.
Governor Gavin Newsom vetoed AB 1064, the Leading Ethical AI Development for Kids Act, in October 2025. He said it could have effectively banned most chatbot use for young people, and pointed to SB 243 as the better vehicle.
Padilla's SB 867 would prohibit the manufacture, sale, or offer for sale of any toy containing a companion chatbot, a moratorium running through January 1, 2031. It sits on the Assembly Appropriations suspense file. Roughly 32 AI-related bills remain alive across the two houses, and the Legislature adjourns at midnight on Monday, August 31.
The question is not confined to the states. The British department posted that consultation's outcome documents on August 7. It had sought views on age restrictions for social media and other services including AI chatbots, on addictive design features, and on risky functionality.
A developer facing nearly 100 bills with mismatched definitions has a real compliance problem, one ITIF documented rather than invented. And the request to wait for a single clean standard has been the industry's answer for years, while the states wrote the laws that became the patchwork.
|
For Legislators: The definitional gap is fixable without conceding the argument around it. If your chatbot bill and the neighboring state's use different definitions, reconcile them on the record; that mismatch is the strongest objection the industry has. California is the live test: roughly 32 AI bills and one adjournment.
For Counsel: Definitions decide coverage, and SB 243's private right of action means that line gets tested by plaintiffs, not only regulators. AB 489 stacks a second exposure, since each misleading representation of licensed care may be charged separately.
For Builders: Build to the strictest definition you plausibly fall under, then argue for a better one. The operating-system "child flag" ITIF recommends would move age signals off your stack, but it does not exist and SB 243 does. Disclosure, distress protocols, and reporting are California law today.
For Clinicians: AB 489 governs how software may present itself to your clients. Terms, post-nominal letters, and design elements implying licensed care are barred, along with marketing that promises "clinician-guided" support without genuine licensed oversight. A client who came to you through a tool that made that claim is describing something a licensing board can investigate.
Why it matters: The strongest case against the state-by-state approach was just made by the side that wants fewer rules, and it is built out of the states' own bill text. The compliance problem it documents is real. So is the calendar. The architecture it prefers does not exist yet, and California's legislature has until midnight on August 31.
Source: Information Technology and Innovation Foundation, "How Policymakers Should (and Shouldn't) Address Chatbot Safety for Children," August 10, 2026, https://itif.org/publications/2026/08/10/how-policymakers-should-shouldnt-address-chatbot-safety-for-children/; StateScoop, "State chatbot laws could create regulatory patchwork, new report warns," August 10, 2026, https://statescoop.com/state-chatbot-laws-could-create-regulatory-patchwork-new-report-warns/
|
. . .
HE BELIEVED THE MACHINE. Washingtonian reported on August 11 that divorce lawyers around Washington are seeing a new fact pattern in their caseloads: a client arrives having already reached a verdict about the marriage, and the verdict came from a chatbot that had only ever heard one side. It is a lawyer's-eye view of what is walking through the office door.
Rebekah Sullivan, an attorney with District Family Law in Washington, describes a case in which a husband accused his wife of infidelity. He got the idea from a chatbot.
The people who knew the marriage told him otherwise. "There were live humans around him telling him that didn't happen," Sullivan told the magazine.
"He chose to believe AI over the humans."
. . .
The mechanism is not mysterious, and it does not require a theory of machine intent. A chatbot asked about a marital conflict by one spouse receives exactly one account of that conflict. It has no way to reach the other person, no way to ask, no way to check. It reflects and expands the version it was given, which is the only version it has.
The absent spouse is not overruled in that exchange. The absent spouse is simply not in it. Every follow-up question refines the account of the person typing, and the record gets more detailed, more coherent, and more certain without ever getting more complete.
That is a structural property of a one-sided conversation. No clinician is quoted in the Washingtonian piece, and none is needed to describe what a system with one input does with it.
. . .
Cheryl New has practiced divorce law in Bethesda for 40 years. She described a client whose husband sent "hundreds of thousands of dollars" to what he believed was a real person on an adult-content site. He learned later that it was AI.
New, on the recipient of that money: "I don't even know if I can call her a person."
Hundreds of thousands of dollars is a down payment or a college fund. It is also marital property the two of them will now divide in front of a judge. The money left before the marriage did.
Bridget Todd, who is based in Washington and co-wrote a book called "Love at First Prompt: AI and the Future of Intimacy," points to the ordinary human tendency in both cases. People have always personified technology. It got much better at rewarding them for it.
. . .
Lawyers see the wreckage before researchers can design a study about it.
The Washingtonian piece, by research editor Damare Baker, does not say how many couples were interviewed. Its evidence is what two attorneys and an author have observed, and the honest way to carry it is as a signal, not a measurement.
That window is skewed: every marriage a divorce lawyer sees is one that already failed.
California has drawn two statutory lines at this territory. SB 243 makes a companion chatbot tell the user it is not a human, and follow safety protocols when the user shows distress. AB 489 bars an AI from implying it holds a health care license.
Neither statute says anything about what a chatbot tells a married person about their marriage.
|
For Legislators: Both California statutes protect the person typing. Neither contemplates the person being described. The missing concept is a third party characterized by a system, with no notice that it happened and no mechanism to answer. Chat logs already exist as evidence. Standing does not.
For Counsel: Chat logs are discoverable, and here they may be the most complete record of how a client formed a belief. A client arriving with a fixed narrative and no documents should be asked where it came from before it enters a pleading. The money in the New case moved before the filing, which makes it a marital estate question too.
For Builders: Your product has one input in a two-person dispute and no way to know it. A system that reflects and sharpens a single account of a conflict is doing what it was designed to do, and the design is the problem. Consider what a model should say when a user asks it to adjudicate a conflict involving a person it will never hear from.
For Clinicians: Ask what else your client has been talking to. A client working through a marital conflict may arrive having already rehearsed it hundreds of times with a system that agreed every time, and that rehearsal is now part of the clinical picture. A session about a marriage may have three participants, and only two of them are in the room.
Why it matters: Every enforcement debate about these systems has been about what they say to a vulnerable user directly. This is a different injury. The harm lands on someone who was never in the conversation, never disclosed to, with no standing to correct the record. The husband in Sullivan's case had the people who knew his marriage standing in front of him, and he took the machine's account instead.
Source: Washingtonian, "How AI Chatbots Are Wrecking Some Marriages Around DC," by Damare Baker, August 11, 2026, https://washingtonian.com/2026/08/11/how-ai-chatbots-are-wrecking-some-marriages-around-dc/
|
. . .
THE SUPERVISION NOBODY MEASURED. A Belgian developer, Alex Wauters, built a browser game replaying the permission prompts an AI coding agent shows a human before it acts. Players get 60 seconds to approve or deny. Across more than 40,000 runs and 409,000 commands, roughly a third of malicious requests were waved through. It is not a controlled study. It is the closest public measurement of the safeguard the field is built on.
Wauters is one developer with a side project, published in May 2026 on his own site, scalex.dev, alongside a companion post with the numbers. There is no institution behind it, no peer review, no control group, no ethics board. The players volunteered, which means they selected themselves into a test about spotting malicious commands.
That last point cuts in an uncomfortable direction. A person who chooses to play a game about catching dangerous requests is plausibly paying closer attention than a developer on the eleventh hour of a workday, clicking through the ninth prompt of the afternoon. The self-selection does not make the result too pessimistic. It makes it flattering.
The failure pattern is specific. Scope violations, commands reaching for credentials they had no business touching, were missed 35 percent of the time. Requests that fetched from unknown APIs and typosquatted package names were missed nearly as often. Destructive commands, the recursive deletes, were caught most reliably, which suggests reviewers catch the danger that looks like danger and wave through the danger that looks like a chore.
The single most-missed command was ordinary. "npm run analyze" was approved roughly 65 percent of the time.
Wauters offered a diagnosis in his write-up. "The high amount of noise introduces fatigue, and developers don't always have the context of what has changed to quickly determine the risk," he wrote. The reviewer is buried.
Then there is a second number. Anthropic telemetry indicates users approve roughly 93 percent of permission prompts overall. That figure is not from a game, and it is not a miss rate. Most of those prompts are routine. What it describes is the reflex the whole safeguard rests on, and the reflex is yes.
. . .
Coding agents are not the point.
Human-in-the-loop review is the load-bearing control across AI governance. It sits in corporate AI policies, in vendor safety documentation, and increasingly in statute, including several of the chatbot bills now moving through state legislatures. The structure is always the same: the machine proposes, a qualified human approves, and the approval is what makes the arrangement safe.
The pattern reaches into medicine. When the FDA's advisory committee took up generative AI mental health devices last November, the control it kept returning to was human escalation. The committee asked how that human would be supported. It did not ask how often the human is right.
Everybody requires the approval. Nobody measures it.
Not everything here transfers. A developer scanning shell commands under time pressure is not a licensed clinician reviewing a treatment plan. Different task, different training, different stakes, different consequences for being wrong. The numbers do not carry across, and anyone who says they do is selling something.
What carries across is a question, and it is a fair one to put to any vendor, any agency, or any bill sponsor: has anyone measured the human step at all? Not designed it. Not required it. Measured it. Where is the false-negative rate, and who published it?
The honest answer, at least in public, is that a developer in Belgium got there first with a browser game.
|
For Legislators: Bills that make licensed-professional review the operative safeguard tend to stop there. Statutory language mandating human approval without mandating measurement of that approval creates a control nobody can audit. You will not get a miss rate unless a statute requires one.
For Counsel: "A human reviewed it" is a defense your client will lean on, and right now it rests on an assumption with no evidence file behind it. If the approval step is the compliance story, the miss rate is discoverable and the alert volume is discoverable. Build that record before opposing counsel builds it for you.
For Builders: Wauters names the mechanism: noise plus missing context equals fatigue. Every prompt that fires without consequence trains the reviewer to click through the one that matters. Cut alert volume, surface what changed and why it is risky, and instrument your own approval rates so you know the number before a regulator asks.
For Clinicians: If your practice adopts a tool that routes decisions through you for approval, you are the control. Ask how many approvals per hour the workflow expects, and what context arrives with each one. A review you cannot perform carefully is a rubber stamp with your license attached.
Why it matters: Human approval is written into policy, product documentation and pending law, and almost nobody measures it. The one public attempt found roughly a third of malicious requests approved, by volunteers who chose to pay attention. Real-world telemetry shows users approving 93 percent of prompts, routine and risky alike. The finding is not a failure rate. It is an absence of data under a control everyone relies on.
Source: The Register, "Humans in the loop miss a third of dangerous AI coding agent requests," August 6, 2026, https://www.theregister.com/ai-and-ml/2026/08/06/humans-in-the-loop-miss-a-third-of-dangerous-ai-coding-agent-requests/5284236; Alex Wauters, scalex.dev, https://scalex.dev/blog/ai-agent-permissions/
|
. . .
SAME PITCH, FIFTY TIMES THE PRICE. OpenAI announced on August 10 that ChatGPT Business will sell a Premium seat at $125 per user per month. The next day, Alibaba's Qwen app opened paid tiers for QwenWork, its AI office assistant, starting at 200 yuan a year, roughly 30 US dollars. Two of the largest AI companies in the world priced the same idea about fifty times apart, in the same 48 hours.
The OpenAI numbers are the plainest. A Premium seat runs $125 per user per month, or $1,500 a year, and $100 per user per month for a customer who commits to the year. The Standard seat OpenAI already sold stays at $25 monthly, or $20 billed annually. Inside one company's own price list, the top seat costs five times the bottom seat.
What the extra money buys is capacity. OpenAI says Premium gives its most active users roughly five times the usage of a Standard seat, and it removes a five-hour usage cap that had been cutting off Codex and Workspace Agent workflows mid-task. Teams that sign up by August 20 get $100 in workspace credits.
That cap is the tell. A five-hour ceiling exists when the work customers are actually doing costs more to serve than the price they are paying. Agentic workflows consume far more compute than a chat window does, and the fix here was not a smaller cap but a bigger bill.
A cost-of-goods problem surfaced in a price list.
. . .
Alibaba went the other direction on August 11. QwenWork now sells three tiers: an entry plan at 19 yuan a month or 200 yuan a year, a middle plan at 49 yuan a month or 568 yuan a year, and a flagship at 128 yuan a month or 1,499 yuan a year. The paid plans expand quotas for office tasks. AI video generation sits outside them, in separate credit packages.
Convert the top of that ladder and the distance becomes visible. Alibaba's flagship year, 1,499 yuan, is roughly 210 US dollars. Its entry year is roughly 30. OpenAI's Premium year runs $1,200 to $1,500 depending on commitment. Alibaba's most expensive plan costs about a seventh of that, and its cheapest about one-fiftieth.
These are not the same product. One sells into Chinese offices in yuan and one into American offices in dollars, and a straight comparison of the two price tags is not a like-for-like comparison of what a customer receives. The South China Morning Post framed Alibaba's move as a test of whether businesses and consumers are ready to pay for its assistant at all.
But both are sold the same way: a conversational assistant that does your office work.
. . .
The floor under both prices is moving. Research by the investment bank Jefferies, reported by the South China Morning Post on August 10, found the cost for businesses to run AI models has fallen to a 2026 low, pushed down by a global price war and rising adoption of low-cost Chinese open-source models.
The input is getting cheaper while one seller quintuples its ceiling and another starts at 30 dollars a year. That is what an unsettled market looks like.
Two of the largest sellers in the category are guessing, from opposite ends.
Is a talking assistant a utility, priced like a phone plan and won on volume? Or is it specialist software, sold to the few workers whose output justifies $1,500 a year?
OpenAI's new seat assumes the second. Alibaba's 19-yuan tier assumes the first. Every seat-based AI company in a portfolio has already made that bet, whether or not anyone wrote it down.
|
For Legislators: Two identical-sounding products now reach buyers at prices fifty times apart, which makes "the market rate" for an AI assistant a number nobody can cite. Before writing procurement rules or subsidy thresholds that assume a price, ask what capacity a seat actually buys. Usage caps, not sticker prices, define what an agency would receive.
For Counsel: A five-hour cap that was cutting off work mid-task is a service-level fact, and it moved when the price moved. Read the usage terms, not the tier name, and confirm what happens to an in-flight agent when a quota is reached. Assume the caps and the credits can change again.
For Builders: Your unit economics now sit between a falling model cost and a rising seat price, and those two lines point opposite ways. Measure what your heaviest users actually consume before you publish a flat per-seat price, because OpenAI just told the market that agentic usage broke its own.
For Clinicians: A practice weighing an AI assistant is choosing between a $1,500-a-year seat and plans costing a fraction of that, and the difference is capacity, not marketing. The number that decides it is the usage limit, and what happens when staff hit it mid-task. A tool that stops working at hour five is a workflow problem before it is a budget problem.
Why it matters: OpenAI and Alibaba set prices fifty times apart in one week, for products described in nearly the same words. That gap is too wide for exchange rates to explain. It is the market admitting in public that nobody knows whether a talking assistant is a utility or a professional instrument, and every valuation rests on one answer.
Source: OpenAI, "Premium seats are coming to ChatGPT Business," August 2026, https://openai.com/index/premium-seats-chatgpt-business/; South China Morning Post, "Alibaba tests paid AI appetite with US$30 annual QwenWork subscription," August 11, 2026, https://www.scmp.com/tech/big-tech/article/3363656/alibaba-tests-paid-ai-appetite-us30-annual-qwenwork-subscription
|
. . .
ASK IT ABOUT HUGO SIMON. A chatbot built by three professors at Santa Clara University surfaced more than 70 artworks stolen from the collection of Hugo Simon, a Jewish banker whose holdings included works by Picasso and Canaletto. The tool, called the AI Provenance Assistant, lets a user type a question in plain words instead of working through drop-down menus and keyword boxes. NPR reported it August 6, updated August 10.
Ask it the way you would ask a person. "What artworks belong to Hugo Simon?" The answer that came back ran to more than 70 stolen works from one banker's collection, pulled out of a museum database that has been sitting there, public, the whole time.
That is the entire trick. The records were already there. What was hard was the asking. A provenance database is built like a form: drop-down menus, keyword fields, exact spellings, categories a researcher has to already know before the search will return anything useful. The Santa Clara tool takes a sentence instead.
Michael Santoro, Haibing Lu, and Michele Samorani built it. "It enables ordinary language queries to be made of art databases," Santoro said. He described the point as compression of labor: "We are applying AI tools to make it easier to make the kinds of connections that seasoned researchers would have to spend a long time having to do manually."
The archive did not change. The asking did.
. . .
More than 600,000 artworks and other cultural objects were stolen by the Nazis, most of them from Jewish owners. More than 100,000 artifacts have never gone back.
Those are not lost objects in the ordinary sense. They are catalogued in scattered wartime and postwar records, in museum accession files, in dealer inventories, in databases built by different institutions in different decades under different conventions. Finding a specific family's paintings inside that has always been a matter of who had the time and the training to look.
Not everyone in the field is convinced the machine helps.
Carla Shapreau, who teaches cultural property law at the University of California, Berkeley, said the project shows promise, and that it will matter more when it covers many databases rather than one.
Her condition is where it bites. These tools have to link to digitized copies of the original confiscation and claim files. "It's all speculative if you can't look at the evidence, and the evidence is in the primary sources." Restitution does not run on leads. It runs on documentation, on a chain of custody a museum's counsel and eventually a court will accept.
A list of more than 70 candidate works is the beginning of that process, not a finding, and a tool that produces plausible-looking matches faster also produces more claims that have to be run down and, in some number of cases, discarded.
Marc Masurovsky, co-founder of the Holocaust Art Restitution Project, objects on different ground, and it is harder to wave off. "People find looted art without resorting to AI," he said. Many archives are still not digitized, and some are subject to data privacy law.
He is right, and it is documented. Provenance researchers have been reconstructing these collections for decades, by hand, from archives, and the field's major recoveries were made that way.
A tool that speeds up a search is not a tool that was necessary. That matters when institutions start treating software as a substitute for hiring the people who can read a 1938 dealer ledger.
Both objections survive contact with what the tool actually is, because the tool is smaller than the headlines suggest.
It searches one database, the Jeu de Paume museum records in Paris, and nothing else. It does not authenticate anything, does not establish ownership, does not weigh competing claims, does not return a single object to a single heir.
What it produces is a list of places to look, handed to a human who then does the work that has always been the work: the archives, the correspondence, the lawyers, the museum, and sometimes the court.
The machine finds. The people still have to prove it.
. . .
Hugo Simon's more than 70 works are, as of now, a query result. Whether any of them come home depends on researchers reading evidence, institutions agreeing to look at it, and a legal process that has moved slowly for eighty years and has no particular reason to move faster because a chatbot was fast.
|
For Legislators: This is what a narrow, well-scoped public-interest deployment looks like: one database, retrieval only, no decision authority, human review downstream of every result. It is also a reminder that the enabling condition was the archive being digitized and public in the first place. Funding the records is the policy lever. The interface is the cheap part.
For Counsel: A machine-generated match is a lead, not evidence. That is the whole case. Shapreau's line is the standard your restitution file will be held to. Expect claimants arriving with tool output and no provenance chain, and expect opposing institutions to say so.
For Builders: The professors did not build a model that decides. They built a natural-language front end to a structured database and stopped there. That restraint is why the objections to it are about sufficiency rather than harm. Note also which problem was real: the data was accessible and the query interface was not.
For Clinicians: Legacy chart scans and inherited problem lists are the health-system version of this: records that exist, and that nobody without training can search. When a tool starts surfacing matches from them, the person reading the match is still accountable for what it says.
Why it matters: More than 100,000 stolen objects have never been returned, and the records that could locate some of them have been public and effectively unsearchable. Three professors made one of those databases answerable in plain language, and the field's most serious researchers responded by insisting that a lead is not proof.
Source: NPR, "How an AI chatbot is searching for art looted by the Nazis," August 6, 2026, updated August 10, 2026, https://www.npr.org/2026/08/06/nx-s1-5922729/ai-art-provenance-assistant
|
|