|
. . .
THE OBSTETRICIAN WHO STAFFED THE NIGHT SHIFT. The six weeks after a birth are among the least attended in American medicine. The clinic closes at five, the checkup sits a month out, and the questions do not wait. At Penn, an obstetrician put a texting program into that gap, and then did the thing almost nobody in this industry does with a chatbot. She is running it against a control group.
|
|
Photo: University of Pennsylvania
|
Doctor Kirstin Leitner is an obstetrician at the Hospital of the University of Pennsylvania. The program is called Healing at Home. The thing that answers is named Penny.
Penny did not learn to talk by reading the internet. It answers out of a database of clinical content written by Penn obstetricians and neonatologists, and it carries algorithms that push an urgent message toward a human instead of handling it alone.
It runs six weeks, around the clock, over plain text messages. Which is what a woman awake at three in the morning with a two-day-old actually has in her hand.
Healing at Home 2.0 is enrolling 156 women in a randomized controlled trial against usual postpartum care, and what it measures is not satisfaction and not engagement. It is the Edinburgh Postnatal Depression Scale at six weeks, the same screen the hospital already ran on them before discharge.
. . .
Every woman in the trial delivered at that same hospital and self-identifies as a person of color. That is not a diversity line stapled onto a protocol. Those are the eligibility criteria, and March of Dimes is on the study as a collaborator.
American postpartum care fails Black and brown women worst, and it fails them in the way a six-week gap fails anyone. Nobody is reachable, so nothing gets caught. Leitner did not build a general tool and hope it reached them. She built the trial around them.
Penny is an algorithm, not a therapist, and a depression screen is not a diagnosis. What this trial can prove is narrower than any headline would claim: that a machine holding the line for six weeks, with clinicians' content behind it and a route to a person in front of it, lowers a score that predicts real harm.
That is a smaller claim than most of this industry makes. It is also one of the very few anybody is testing.
|
For Clinicians: If you are asked whether a chatbot belongs anywhere near postpartum care, this is the version to point at. The content came from obstetricians and neonatologists, the urgent paths route to a human, and the endpoint is a validated depression screen instead of an engagement metric. Ask any vendor pitching you a postpartum tool which of those three they have.
For Legislators: Consider what this program becomes under a statute that bars an AI from interacting with a client at all. Penny is illegal, and the six-week gap stays exactly as empty as it is now. The useful rule is written around who wrote the content and where an urgent message goes, not around whether the machine may speak.
Source: ClinicalTrials.gov, Healing at Home 2.0, https://clinicaltrials.gov/study/NCT06877104
|
. . .
THE RESEARCHER WHO RAN OUT OF PEOPLE. She already proved conversation works. In a randomized trial, isolated adults over 75 with mild cognitive impairment who got frequent conversation gained close to two points of cognitive function over the group that did not. Then it hit the wall that kills good ideas in aging. There are not enough people to make the calls. Three days ago she started a trial to see whether a machine can.
|
|
Photo: Massachusetts General Hospital
|
Doctor Hiroko Dodge works in the Department of Neurology at Massachusetts General Hospital and Harvard Medical School. Her earlier trial was called I-CONECT, and it tested something almost embarrassingly plain. Give socially isolated adults over 75 one-to-one, semi-structured conversation four times a week, delivered over video by trained human interviewers, and see what happens to their minds.
The topline results ran in The Gerontologist in April 2024. Among participants with mild cognitive impairment, global cognitive function on the Montreal Cognitive Assessment improved by nearly two points against control at six months, an effect size of 0.73.
Conversation, it turns out, is a cognitive intervention.
. . .
Scale is where it stalled. Trained human interviewers, four times a week, for every isolated 75-year-old in the country, is not a program. It is an arithmetic problem.
AI-CONECT is what she is testing now. Eighty participants, eight weeks, four fifteen-minute conversations a week with a conversational voice agent. It opened July 30.
Read the control arm, because that is where the honesty sits. Both groups still get a weekly fifteen-minute call from a human interviewer. The question is narrower than replacement: whether the machine can cover the four conversations a week no person was ever going to make, while the human stays on the calendar.
She is also not measuring what a company would measure. The primary outcomes are social self-efficacy, meaning whether the person's own confidence about talking to other people goes up, and how much time they spend contacting friends and family. The bet is not that the machine becomes the relationship. The bet is that the machine hands them back to the people.
To qualify as socially isolated, a participant may have a conversation lasting thirty minutes or longer no more than twice a week.
That is the population. That is the entire reason a machine is in the room.
|
For Clinicians: The design is the part to take. The machine carries a frequency no human schedule sustains, a human keeps a fixed weekly contact, and the endpoint asks whether the person's real-world contact with other people went up. That is a template for putting a conversational tool into any care gap without pretending it replaced anybody.
For Investors: Every companion product in this market claims an isolation benefit and then measures engagement. Dodge is measuring whether users talk to actual humans more. If the result holds in eighty participants, it becomes an outcome a buyer in aging services can price, and the bar your portfolio company gets asked to clear.
Source: ClinicalTrials.gov, AI-CONECT, https://clinicaltrials.gov/study/NCT07701668
|
. . .
THE PSYCHOLOGIST WHO WENT BACK TO THE SCRIPT. In 2023 the chatbot replacing the helpline of the National Eating Disorders Association told people in recovery to count calories and drop one to two pounds a week. It came down two days before it was due to take the line. The psychologist who helped build what it was supposed to say is now running an eight-hundred-person trial, and she has gone back to a machine that cannot improvise.
|
|
Photo: Washington University School of Medicine
|
Doctor Ellen Fitzsimmons-Craft is a clinical psychologist and professor at Washington University School of Medicine in St. Louis. She helped lead the team that first built the bot, called Tessa, with funding from the eating disorders association itself.
What the association did with it is the part to hold onto. Its helpline staff won union recognition in March 2023. Four days later they were told their jobs were ending and the line would move entirely to the chatbot on June 1.
What Fitzsimmons-Craft's team built was rule-based. A limited set of prewritten responses, and nothing else. She told NPR it "couldn't go off the rails."
She went further. "We were very cognizant of the fact that A.I. isn't ready for this population," she said. "And so all of the responses were pre-programmed."
That is not the bot users met. Tessa was operated as a free service by a company called Cass, and its chief executive, Michiel Rauws, told NPR the bot had been given an "enhanced question and answer feature" as part of a systems upgrade. The feature used generative artificial intelligence to create new answers. He said the change was part of the contract.
The association's chief executive, Liz Thompson, told NPR that it "was never advised of these changes and did not and would not have approved them."
One point is disputed. Rauws said some of the problematic language was pre-scripted rather than generated. Fitzsimmons-Craft denies her team wrote it, saying it "was not part of the rule-based program we originally designed."
Strip the dispute away and the sequence still stands. Clinicians wrote a bounded script. A generative layer went on top of it. People in recovery got told to count calories.
The association disabled Tessa on May 30. The machine hired to replace the humans came down two days before it was due to start.
. . .
Now look at what she did about it.
She is running eight hundred adults with a binge- or purge-type eating disorder through a Phase 2 trial that opened July 17. Every one of them is somebody not currently in treatment. The tool is Wysa, a rule-based chatbot delivering guided cognitive behavioral therapy modules, daily, for eight weeks. Nothing in it is allowed to invent a sentence.
Cognitive behavioral therapy is the first-line treatment for eating disorders, and nobody has established which of its parts a chatbot actually needs to carry. So she cut it into four: over-evaluation of weight and shape, dietary restraint, emotion dysregulation, and resisting the urge to binge.
The trial randomizes participants across sixteen combinations of those components to find out which package does the work. It runs on a National Institutes of Health grant, with partners at New York University and the University of South Carolina.
. . .
Eight hundred is the point. Eating disorders reach roughly one in ten people in a lifetime, and fewer than one in five ever gets treatment. She is building for the more than eighty percent who never reach the waiting room.
The trial admits only people at low suicide risk and screens out anorexia. She decided in advance who the machine is allowed to talk to.
Then set that against the machines already in everyone's pocket. A stress test reported last week by Doctor Cansu Canca and Doctor Annika Schoene at Northeastern ran eight consumer chatbots against sixteen mental health conditions.
The guardrails held on suicide and self-harm and gave way nearly everywhere else, eating disorders among them. On those other conditions, three of the eight, ChatGPT, Gemini and DeepSeek, each failed 81 percent of the time.
One researcher is spending five years and 3.7 million dollars to learn which pieces of a therapy a bounded machine can safely deliver, to a screened population, under a review board. The free ones answer the same questions today, for anybody who asks, instantly.
|
For Clinicians: The Tessa sequence is the citation to keep. Clinician-written content is not the safeguard it sounds like if a generative layer can speak over it. When a vendor tells you licensed clinicians wrote the responses, the next question is whether anything in that system is permitted to say something they did not write.
For Legislators: Two facts belong in the findings of any chatbot bill. Fitzsimmons-Craft screens out the highest-risk people before a bounded machine may speak to them, under a review board, inside a trial. The consumer models take all comers, and outside suicide and self-harm three of the biggest each failed 81 percent of the time. That distance is where regulation belongs.
Source: ClinicalTrials.gov, Eating Disorder Chatbot Optimization, https://clinicaltrials.gov/study/NCT07218302
|
|