Three Champions, and their revolutionary applications

Conversational AI Watch

Conversational AI Watch

The news that moves policy, portfolios, and patient safety.

By Jess Jessop  |  August 9, 2026  |  Issue #121

▶ WATCH🎧 QUICK LISTEN🎧 DEEP DIVE
Infographic: three champions testing supervised conversational AI: the HALO family-support chatbot trial inside a human helpline, the QuitBot smoking-cessation trials, and NIH-funded guideline-grounded cancer symptom chatbots.
Jess Jessop

JessJessop.Info

Jess's Sunday Reflection

Three Champions, and their revolutionary applications

A helpline scientist teaching a machine from forty thousand family calls, a quit coach who raced the government's own program, and a cancer doctor who measures the machine before it speaks.

Three Champions, and the safe, life changing, revolutionary applications they are building.

Every weekday this newsletter watches the machines, and the people who answer for them.

. . .

Sunday is different.

. . .

Sunday is the part of the week when we focus on the good work being done that isn’t making headlines. We celebrate people in the trenches doing the unglamorous hard work.

The three people on this page are the tip of an iceberg of good work and great things being built. They focus on what the machine will be allowed to say. They put the plan in front of a review board. They recruit a control group. They agree, in advance and in public, to be measured against it.

Then, and only then, does the machine get to talk to a human being who needed one.

The federal grant database and the trial registry are full of people working this way. No press release. No launch video. A number instead of a headline. That is where I went looking this week.

. . .

A researcher who runs the science at the nation's family addiction helpline, teaching a machine to answer the mother of an addict at two in the morning by feeding it forty thousand conversations her own counselors already had.

. . .

A psychologist in Seattle whose quit-smoking bot went head to head with the government's own program, a 1,647-person trial just completed, and who is now rebuilding it with the Native communities the tobacco industry chased hardest.

. . .

A cancer doctor in Boston who caught the world's most popular chatbot getting cancer treatment wrong a third of the time, and who just took federal money to build the version that has to cite its sources.

. . .

If you are a legislator or a staffer reading this, notice what none of these three asked for. Not a ban. Not a license to run loose. A protocol, a supervisor, and a scorecard someone else can check.

. . .

Test it first. Bound what it may say. Keep a human on the line. Publish the number, whatever it turns out to be.

. . .

A helpline scientist. A quit coach. A cancer doctor. Three different angles on the same case.

. . .

Here they are.

Reader Pulse

The people testing the machines before they touch you.

🔥  Glad someone checks
✏️  Sending to a clinician
💪  Trials move too slow
🤔  Which trial was which
💬  Naming a fourth

Forward to a colleague →  ·  Join the discussion →

. . .

THE RESEARCHER WHO ANSWERED THE OTHER PHONE. The person with the addiction has a hotline. The mother awake at two in the morning, phone in hand, listening for a car in the driveway, has one phone number. At Partnership to End Addiction, a researcher is building a chatbot for her the slow way: inside a human helpline, out of that helpline's own conversations, against benchmarks it must clear before it is allowed to grow.

Doctor Sarah Dauber of Partnership to End Addiction, contact PI of the HALO family-support chatbot project
Photo: Partnership to End Addiction

Doctor Sarah Dauber is Vice President of Research and Evaluation at Partnership to End Addiction, the national nonprofit formed when Partnership for Drug-Free Kids merged with the Center on Addiction. The Partnership runs a free national helpline for families of people with addiction, 855-DRUGFREE. Its helpline and digital family support programs reach more than twenty-three thousand families a year.

Dauber knows exactly who calls, because she counted. A study she co-authored in the Journal of Medical Internet Research analyzed 24,096 helpline client records from April 2011 to December 2021. 76.1% of the help-seekers were women. 68.9% were parents contacting the service about their own children. 17.8% asked for support in Spanish.

The people answering do not improvise. They work from motivational interviewing and Community Reinforcement and Family Training, both evidence-based approaches, over phone, text, email, and Facebook Messenger.

This spring, the National Institute on Drug Abuse decided that decade of conversations was worth building on.

The grant is an R61/R33, project 1R61DA064828-01, $453,750 in first-year funding, its first phase running March 2026 through February 2028. Dauber is contact PI. Her co-PI is Frederick Muench, the Partnership's former president and CEO, now its Senior Advisor. What they are building is called HALO, Helping Affected Loved Ones, a chatbot for the parents, partners, and caregivers the research literature files under a bloodless label: concerned significant others.

. . .

HALO does not start by reading the internet. The first phase runs natural language processing over more than forty thousand actual helpline interactions, the Partnership's own counselors doing the work the Partnership's own way. Alongside that come surveys and interviews with family members, with their loved ones, and with the peer coaches who will work next to the machine.

Then a three-month pilot with 60 families, in three arms: helpline alone, helpline plus HALO, helpline plus HALO plus peer support.

And then a wall. Only if the pilot clears predefined feasibility benchmarks does the second phase begin: a six-month factorial trial with 800 family members, testing the helpline, the chatbot, and peer support separately and in combination. A sub-sample of 200 of the people using substances gets assessed directly. The outcomes are family recovery capital, the family member's own functioning, and the loved one's substance use.

The grant states the aim plainly: "the first rigorously tested AI-human hybrid support model" for these families. Nobody has rigorously tested one before.

. . .

What Dauber is not doing is the design. She is not replacing her counselors; the bot learns from their ten years of conversations. She is not launching an app into a store; HALO sits inside the helpline that already exists. She is not skipping the peers; they are a trial arm. And she is not scaling on faith; the 800-person trial happens only if the 60-family pilot earns it.

The families themselves are the point. They carry high distress and get little evidence-based support, blocked by stigma, cost, and a treatment system built around the person using, not the people holding the household together. Family involvement improves treatment engagement and recovery.

That is who the phone belongs to.

For Clinicians: The helpline behind this project exists today, free, in English and Spanish, at 855-DRUGFREE; that is where a client's family member sends the questions. When a vendor pitches a family-support bot, hold it against HALO's design: trained on its own counselors' conversations, embedded in a staffed service, benchmarked against a control before scale. Ask which of those they have.

For Legislators: Consider this grant under a statute that flatly bars an AI from interacting with a person seeking help. HALO is illegal before its first benchmark is measured, the trial dies, and the family keeps what it has now. Write the rule around what the machine was trained on, where it sits, and what evidence it must clear, not whether it may speak.

Source: NIH RePORTER, HALO (Helping Affected Loved Ones) R61/R33, https://reporter.nih.gov/search/1R61DA064828-01/project-details

Comment on this story →  ·  Forward this →

. . .

THE PSYCHOLOGIST WHO RACED THE GOVERNMENT FIRST. He tested his quit-smoking chatbot against the government's own program before he tested it against anything else. Among pilot participants who finished the 42-day program, 63% were off cigarettes at three months, against 38.5% on the government's. The full-scale trial, 1,647 adults, finished in April. Instead of cashing out, he is rebuilding it with the communities commercial tobacco has hit hardest, and the new trial opens September 30.

Doctor Jonathan Bricker of Fred Hutchinson Cancer Center, creator of the QuitBot smoking-cessation chatbot
Photo: Fred Hutchinson Cancer Center

Doctor Jonathan Bricker is a psychologist at Fred Hutchinson Cancer Center in Seattle. His team built QuitBot over four years of user-centered design, and per the NIH grant application it is the first known smoking-cessation chatbot with conversation features supported by a large language model.

The program itself is plain. Forty-two days. Motivations to quit, coping skills for the triggers, relapse prevention. It is a free smartphone app in the app stores, with no clinic visit, no provider training, no integration into anybody's medical system. It is awake whenever the craving is.

Before scaling anything, Bricker ran it against the strongest comparator available, the National Cancer Institute's own SmokefreeTXT text-messaging program. The pilot randomized 404 adults, and the results appeared in JMIR mHealth and uHealth in 2024. Retention held at 96% at three months.

Among participants who viewed all 42 days of program content, 30-day abstinence came in at 63% (39 of 62) for QuitBot against 38.5% (45 of 117) for SmokefreeTXT, an odds ratio of 2.58, P=.005. A completer comparison, not the whole-trial answer.

Receipts first. Reach second.

. . .

The whole-trial answer is coming. The full-scale trial enrolled 1,647 adults starting in July 2022, followed them for twelve months, and completed on April 8, 2026. The results are not published yet. They will publish either way. Running against the government's program instead of a waitlist is what makes the answer worth having.

. . .

The next population is not the easiest market. It is the hardest-hit one. Per the NIH grant application, American Indian and Alaska Native people have the highest rates of commercial cigarette smoking of any racial or ethnic group in the country, six times the rate of smoking-related cancers compared with other groups, and half the odds of quitting.

Commercial cigarette smoking now accounts for half of all deaths in those communities nationwide.

Commercial is the operative word. Traditional and ceremonial tobacco are something else entirely, and the grant language holds that line.

The grant names two causes. Lack of access to cessation interventions, and lack of interventions proven to work for these populations. That reads like the job description of a free chatbot in a pocket, in communities where 68-78% of people own smartphones.

The new trial is called NAITIVE, Navigation and Artificial Intelligence Technology for Indigenous Virtual Education. It will randomize QuitBot, adapted for American Indian and Alaska Native adults nationwide, against SmokefreeTXT, 386 per arm. It opens September 30.

And it is not an app-store push. NAITIVE is Research Project 1 of a National Institute on Minority Health and Health Disparities center grant at Fred Hutch, the CANOE Partnership, Cancer Awareness, Navigation, Outreach, and Education, which carries a dedicated community engagement core. Bricker leads it, $591,836 in fiscal 2026.

The adaptation is being built with the communities, inside partnership infrastructure, before a single participant enrolls.

For Clinicians: The sequence is the part to take. Prove the tool against the strongest comparator, publish, then adapt it for the population with the worst access through partnership rather than a download link. The access math holds on its own: a free app, no referral, no training, available at any hour to a client the system was never going to reach. It is in the app stores now.

For Investors: Most cessation products cite engagement against no comparator. Bricker ran 404 people against the government's own program, published the completer caveat alongside the win, then enrolled 1,647 for the twelve-month answer. When those results land they set the evidence bar for the category. And NAITIVE shows where defensibility lives in underserved markets: not in the model, in community partnership a competitor cannot clone from an API.

Source: ClinicalTrials.gov, NAITIVE, https://clinicaltrials.gov/study/NCT06697496

Comment on this story →  ·  Forward this →

. . .

THE ONCOLOGIST WHO MEASURED THE MACHINE FIRST. In 2023 Doctor Danielle Bitterman's team put the world's most popular chatbot through the guidelines her field lives by, and one response in three carried treatment advice the guidelines do not support. She published that number. Then she published the harm rate of AI-drafted replies to patient messages. Now she has taken federal money to build the version that has to ground every answer in the actual clinical guideline.

Doctor Danielle Bitterman of Brigham and Dana-Farber in Boston, principal investigator of AI-CChaSE
Photo: Mass General Brigham

Doctor Danielle Bitterman is a radiation oncologist at Brigham and Women's Hospital and the Dana-Farber Cancer Institute in Boston, an assistant professor at Harvard Medical School, and faculty in the Artificial Intelligence in Medicine program at Mass General Brigham. In January 2025 the system made her its Clinical Lead for Data Science and AI.

Start with the first study. In August 2023 her team published a test in JAMA Oncology, Shan Chen as first author, Bitterman as corresponding author. They ran ChatGPT against the 2021 National Comprehensive Cancer Network guidelines, the reference oncologists actually treat from, for the three most common cancers: breast, prostate, and lung.

They wrote 26 descriptions of cancer diagnoses, ran each through four prompt templates, 104 prompts in all, and put every response in front of three board-certified oncologists scoring against five criteria.

About one response in three, 34%, contained at least one treatment recommendation not concordant with the guidelines. Roughly 13% recommended treatments that were not in the guidelines, or did not exist at all.

The dangerous number is the third one: 98% of responses also included at least one approach that did agree with the guidelines. The wrong advice never arrived alone. It sat mixed in with sound advice, and it was hard to spot.

. . .

Then she measured the other direction. Hospitals were starting to let AI draft doctors' replies to patient messages, so in 2024 her team at Mass General Brigham tested GPT-4 doing exactly that, on messages about cancer scenarios, in The Lancet Digital Health.

More than half the drafts, 58.3%, were acceptable to send without a single edit. But 7.1% of unedited drafts posed a risk to the patient, and 0.6% posed a risk of death.

Her sentence on that finding is the one to keep: "Keeping a human in the loop is an essential safety step when it comes to using AI in medicine, but it isn't a single solution."

Now look at what she is building.

The grant is called AI-CChaSE, AI-enabled Cancer CHatbots for Symptom Education. It is a National Cancer Institute R01, $819,628 in first-year funding for fiscal 2026, awarded to Brigham and Women's Hospital. Bitterman leads it with two co-principal investigators, Harry Hochheiser and Guergana Savova.

The premise comes straight from the abstract. People with cancer routinely live with chronic symptoms from both the disease and the treatment, devastating to physical, psychological, social, and financial wellbeing. Good symptom care is limited by access, education, and communication. Chatbots have helped in studies, but the safe ones were rules-based, inflexible, effort-heavy to build, and unable to adapt when guidelines change.

Aim one answers her own 2023 finding: develop methods that link the language model to the knowledge in the clinical practice guidelines themselves, so an answer about a symptom is grounded in the document, not the model's memory. Aim two: simplify guideline and health-record language for different readers without distorting it.

Aim three: study the needs, perceptions, concerns, and ethics questions of patients and clinicians, feed all of it into the design, and evaluate the resulting interface in empirical studies.

. . .

Hold the timeline. In 2023 she measured the consumer machine's failure rate on cancer treatment questions and published it. In 2024 she measured the harm rate of AI drafts, published it, and said the human in the loop is essential and not sufficient. In 2026 she took federal money to build the bounded version: tied to the guidelines, built with its users, tested before deployment.

The chatbot she measured in 2023 was already answering cancer questions for anyone who asked. Hers will not speak until it is tested.

For Clinicians: The 98% figure is the citation. A general-purpose chatbot is not obviously wrong; the wrong recommendation arrives braided into right ones, and it took board-certified oncologists to separate them. The drafts that look sendable are what erodes the reviewing habit, and 0.6% of unedited drafts posed a risk of death. Ask any vendor what their system is grounded in besides the model's memory.

For Legislators: Three numbers belong in any health chatbot bill's findings. The most popular chatbot gave cancer advice the guidelines do not support 34% of the time, and sounded partly right 98% of the time. AI drafts sent unedited carried a 7.1% risk rate. The federal government now funds the alternative: guideline-grounded, designed with its users, evaluated before deployment. That is the bar a statute can name.

Source: NIH RePORTER, AI-CChaSE, https://reporter.nih.gov/search/1R01CA311924-01/project-details

Comment on this story →  ·  Forward this →

Disclosure

Conversational AI Watch is published by Jess Jessop and sponsored by Clinician Assist Inc. The author has no commercial relationship with Doctor Sarah Dauber, Doctor Jonathan Bricker, or Doctor Danielle Bitterman. None of them was compensated, consulted, or shown this issue before publication.

Clinician Assist builds software that keeps a licensed clinician in charge of an AI tool. Every project in this issue leans the same way, toward bounded machines with humans close by, and readers should weigh that interest. The reporting draws on the federal grant database, the trial registry records, the published studies, and the researchers' own public statements. This newsletter is produced with an artificial intelligence model.

The fast road will be back on this page tomorrow. Another model, another pause, another investigation into what got loose.

Carry this morning's three into the week.

A chatbot for the family of the addict, learning from forty thousand conversations its own counselors answered. A quit coach that faced the government's program in a 1,647-person trial and is waiting on the answer. A cancer bot that is not allowed to speak until it can cite the guideline.

None of them shipped first. All of them are doing the hard, time consuming work that is necessary for safety.

Three names from the tip of the iceberg. The trenches are full of people building the same way and my hat’s off to them. If you are one of these builders; you have my undying respect and admiration, thank you for all of us!

See you tomorrow morning.

Today's Question

The chatbot that helps you at your worst moment: who gets to build it?

The clinicians treating you
The big AI labs
Anyone who proves it in a trial
Nobody. Keep machines out of it

One tap. Results on the other side.

The Book • Out Now

Therapist in the Loop book cover: a therapist and a client in armchairs joined by a glowing loop of light

Therapist in the Loop

by Jess Jessop

One billion people live with a mental health disorder. Most will never see a therapist. Into that gap has rushed a generation of chatbots that talk like clinicians and answer to no one.

The book lays out the architecture this newsletter tests against every statute and docket: client, therapist, and machine, governed by Six Laws offered as an open safety standard.

The machine can help. It cannot be left in charge.

Get the Book on Amazon →

Kindle, hardcover, and paperback

More On Our Radar

London promises the child-chatbot rules The UK government's response to its childhood-online consultation, posted in full Friday, commits to blocking under-18s from sexualized AI chatbot services, banning sexual role-play features on the rest, and mandating breaks in chatbot use for minors. Not a bill yet. The commitments are on paper. Source

The tutor that knows when to hold back The Allen Institute for AI released TutorMoments, a benchmark testing whether AI tutors know when to step in and when to let the student do the work. First finding: told only to tutor well, models over-help and rarely push the student to think. Source

Clinical AI, behind closed doors STAT reports federal health regulators invited tech companies, researchers, and lobbyists to closed-door meetings on boosting clinical AI adoption. The push to put more AI into care is being shaped in rooms the public cannot enter. Source

Brush Your Brain - The jingle

that started a movement

Watch on YouTube

This Issue

40,000 conversations, 1,647 smokers, a bot on a leash.

The 2am families
Filed all three trials
Funding isn't proof
Lost in the acronyms
Know one we missed

If you or someone you know is in crisis, call or text 988 (Suicide and Crisis Lifeline).

Jess Jessop is the Founder and CEO/CTO of Clinician Assist Inc. (BetterMind.Space), building the first voice-first AI-native mental health EHR with Casey Life and Peer AI Coach supervised by licensed therapists. A disabled veteran and 25-year AI/software engineering veteran, Jess brings lived experience as a mental health client to the mission of making daily mental health care as integrated as oral care.

ClinicianAssist.ai  |  BetterMind.Space  |  JessJessop.info

Subscribe  |  Archive  |  Unsubscribe