What AI might do to mental health care by 2030, what to protect, and what to build.
150 minutes together, hosted by Ariveria
What the research already shows · one plausible 2030 · two group discussions
150-minute workshop · Singapore
Here is the plan for the next 150 minutes. We will spend one ordinary Tuesday in 2030 with a client called Maya, walk through the research that is already pointing that way, and argue twice, in groups, about what to do with it. Nobody can tell you what 2030 will look like. What we can do in this room is practise deciding what good care looks like while the tools keep changing.
Everything in the room stays here to revisit. The field guide carries the questions home.
ONE PLAUSIBLE 2030
A Tuesday in 2030.
Maya’s day, a few years from now.
01Maya’s watch has tracked her sleep, her heart rate and her screen time for months. It noticed her broken week before she did.
02On the train she talks it through out loud. Her AI answers in a voice that remembers last month’s fight with her colleague.
03It drafts the difficult message, then suggests she waits a day before sending it.
04It has already written a summary of her week. She chooses what her therapist sees.
05By the time she sits down in your room, something has been listening, noting and advising for weeks. What arrives with her?
The capabilities behind the scenario
[F1]OpenAI, 2026Source + limits
Build more natural voice experiences with GPT-Live-1 in the APIOfficial technical product documentation · 2026-09-10What it supports: Documents a production full-duplex voice model that listens and speaks at the same time, handles interruptions and backchannels, retains context over longer sessions, and can delegate reasoning or tool calls to another model. OpenAI reports a 30-percentage-point gain over GPT-Realtime-2.1 on its Full Duplex Bench.Limitation: The evaluations cover vendor benchmarks, customer service, banking support, and language tutoring. They do not establish emotional understanding, clinical safety, therapeutic benefit, or safe response during distress; the headline results are vendor-reported.Open the original source (opens in a new tab)
[F3]OpenAI, 2026Source + limits
Dreaming: Better memory for a more helpful ChatGPTOfficial product research report · 2026-06-04What it supports: Documents a deployed memory architecture that synthesizes context from past conversations, updates memories over time, and exposes a reviewable memory summary to users.Limitation: A vendor report about one product. Memory quality is evaluated by the provider, does not imply complete recall or clinical understanding, and raises control, deletion, provenance, and stale-inference questions.Open the original source (opens in a new tab)
[F4]Xin Liu, Daniel McDuff, Google Research, Google DeepMind, and collaborators, 2026Source + limits
SensorFM: Towards a general intelligence and interface for wearable health dataOfficial research report linked to a research paper · 2026-07-09What it supports: Reports a wearable foundation model trained on more than one trillion minutes of multimodal sensor data from five million consented participants, transferring across 35 health-prediction tasks and tested as grounding for a personal health agent.Limitation: The health-agent test used 31 participant profiles and clinician ratings of generated summaries; it is not evidence of clinical outcomes, mental-health diagnosis, crisis prediction, or population-wide validity. This is also an organization-authored research summary.Open the original source (opens in a new tab)
[F28]Stamatis et al., 2024Source + limits
Differential temporal utility of passively sensed smartphone features for depression and anxiety symptom prediction: a longitudinal cohort studyPeer-reviewed longitudinal cohort study · 2024-01-04What it supports: Studied 1,013 adults and related passive location, communication, and phone-use features to later depression and anxiety symptoms. More time at home relative to a person’s own baseline was associated with higher future depression scores.Limitation: Full models explained only about 5–6% of symptom variance. Findings were correlational, time-lag dependent, affected by pandemic-era routines, and based on a demographically limited sample; passive data could not be shared publicly because of re-identification risk.Open the original source (opens in a new tab)
[F29]Daniel McDuff et al., 2025Source + limits
Evidence of differences in diurnal electrodermal, temperature and heart-rate patterns by depression and anxiety symptomsPeer-reviewed prospective observational study · 2025-08-17What it supports: Followed 237 participants for four weeks using Fitbit Sense 2 data and reported group-level differences in tonic electrodermal activity, skin temperature, and heart rate between higher- and lower-symptom groups.Limitation: The study used questionnaire-defined symptom groups and found correlates rather than a diagnostic test, causal effect, or reliable individual warning system. Most authors were affiliated with Google or Verily.Open the original source (opens in a new tab)
BY 2030
The ground moves first.
Seven shifts already under way. Each one shows up in research you can read today.
01
AI becomes conversational, multimodal and continuously available.
Randomized 210 adults with clinically significant depressive or anxiety symptoms or high risk for eating disorders to four weeks of Therabot or a waitlist; the intervention group showed greater symptom improvements and reported a therapeutic alliance.
Compared a 12-week conversational AI intervention, face-to-face group therapy, and waitlist among 995 psychologically distressed university students in Israel; anxiety and well-being favored the AI arm over both comparators, while other outcomes were mixed.
It feels personal
People bond with chatbots, quickly.
Users of an early therapy bot reported a working alliance within days. People also rate AI-written empathy highly, and still prefer to receive it from a human.
Limit: A self-reported bond is not proof of equivalent therapy, and the preference studies sit outside clinical settings.
Analyzed aggregate data from adult Woebot users and found self-reported working-alliance and bond scores within days of use that were comparable with scores reported in some prior CBT studies.
Across four studies, participants generally preferred receiving empathy from humans while rating AI-generated empathetic responses as higher quality and more effective at making them feel heard when they encountered them.
It sounds confident
AI that always agrees changes people.
In three preregistered experiments, overly agreeable AI made people more certain they were right and less willing to repair a conflict. They trusted it more, not less.
Limit: The experiments measured intentions, not relationships or clinical outcomes.
Across computational analysis and three preregistered experiments involving 2,405 participants, sycophantic AI increased perceived rightness, reduced intentions to repair interpersonal conflict, and increased preference, trust, and intended reliance on AI.
Working alongside
AI can make a human response better.
Peer supporters who wrote with AI suggestions expressed more empathy, while keeping the power to accept, edit or ignore every suggestion.
Limit: The study measured the writing, not whether anyone felt better. Peer support is not therapy.
In a non-clinical trial with 300 peer supporters, just-in-time AI feedback increased expressed conversational empathy by 19.6% overall and more among supporters who reported difficulty providing support.
Documentation
A third listener is already in the clinic.
Ambient scribes now draft medical notes in real trials. One saved writing time; a separate evaluation found human notes still scored higher.
Limit: Medical documentation is not psychotherapy notes, where the record itself is clinically sensitive.
Randomized 238 outpatient physicians across 14 specialties to two ambient AI scribe tools or usual care and measured time in notes plus workload, burnout, safety, accuracy, and usability outcomes.
Maps therapy guidance and experimentally probes several large language models, reporting stigmatizing patterns and inappropriate responses to some presentations involving delusions, suicidality, hallucinations, and mania.
WHAT THIS CHANGES
Old questions, new answers due.
When AI sits inside the client’s week, the familiar questions of the room need asking again.
01
What the client brings
An AI-shaped account of the week: remembered, worded and smoothed before you hear a word of it.
02
What the practitioner already knows
A summary, a risk flag and a sleep chart wait in the file. What do you owe the data, and what do you owe the person?
03
What the AI recommends
When the wording, the prompt and the next step come from a tool, whose intervention is it?
04
The role of the therapist
Which parts of your work need you, and which parts only need doing?
05
What may be automated
Intake, check-ins, homework, documentation. Each piece handed over changes all the others.
06
What must not be
Every profession draws this line somewhere. Ours is not drawn yet.
DISCUSSION 01 · 15 MINUTES
What should remain distinctly human?
If a 2030 system could do everything described today, what should still be done by a person?
Discuss together
What in your work loses its meaning if a machine does it?
What would you gladly hand over, and what does that free you to do?
Where would your clients draw the line, and how would you find out?
Your group’s output
01Something only a person should do
02Something you would gladly hand over
03The line you would defend
How the 15 minutes run
2 minThink alone, against the six questions
8 minDiscuss in your group
3 minAgree the line you would defend
2 minBring it back to the room
Timer · 15 min15:00
Private notes on this device. Nothing is submitted.
Notes stay in this browser session and are never sent anywhere.
This line will not be drawn for the profession. It gets drawn in rooms like this one.
Discussion 1 of 2.
PAUSE
Take a break.
THE SEVEN-EYED MODEL, UPDATED
Seven eyes on 2030.
Supervision already examines every relationship in the room. Here is where AI enters each one.
Hawkins and Shohet drew this map for clinical supervision in the 1980s, and supervisors have been taught it ever since. Its claim: a session is more than the client's story. There are seven places worth looking, and each one shows you something the others cannot.
In practice a supervisor picks an eye, looks through it, then moves. There is no fixed order. The skill is noticing which eye you have been avoiding.
After Hawkins and Shohet. Two interlocking systems inside one wider context; seven places to look.The update for 2030: the geometry is unchanged, and one system now sits on every line.
01
The client
Eye 1 watches the client: what they bring, how they present, what they choose to tell and what they hold back.
02
The therapist
Eye 2 looks at the work itself, the interventions. Eye 4 turns inward, to what the client stirs up in the therapist. Countertransference lives here.
03
The therapy relationship
Eye 3 looks at what happens between the two of them. The alliance stops being the container and becomes the material.
04
The supervision system
The therapist carries the work to a supervisor. Eye 5 watches that relationship, eye 6 the supervisor's own process. What happened in the therapy has a way of replaying in the supervision; supervisors call it parallel process.
05
The wider context
Eye 7 steps back. Organisations, funding, culture and law shape both relationships before anyone says a word.
06
A fourth presence
In 2030 the map shares every room with a system that has already heard the client's week, drafted the therapist's notes and flagged a risk to the supervisor. The geometry is unchanged. Every line now carries it, and each eye below traces where it enters and what it does there.
Drawn for three people. The room now has a fourth.
The update, eye by eye
01
The client and their presentation
AI enters before the session begins. The client has often been talking to a system all week, and that system has been keeping score.
Clients may arrive with an AI-shaped account of their week. What has been framed before you hear a word?
02
The therapist’s interventions
AI enters the work itself. A tool can draft the reflection, suggest the homework or word the difficult question before you do.
A tool may supply the wording, the prompt, the next step. Whose intervention is it?
03
The client–therapist relationship
AI enters the space between you. Each of you may have consulted a system about the other: what to say, how to take it, whether the therapy is working.
The alliance turns triadic: client, therapist, and the systems each one brings into the room.
04
The therapist’s internal process
AI enters what you feel in the room. Something that sounds certain can pull your judgement toward it, and noticing that pull is now part of the work.
Countertransference now includes how you feel about a client’s AI: dismissal, deference, unease.
05
The supervisory relationship
AI enters the supervision hour. The supervisor may meet the system’s version of the session first: the summary, the flagged moment, the suggested focus.
Supervision gains a question alongside “what did you do?”: what did the tool do?
06
The supervisor’s own process
AI enters the supervisor’s judgement too. The instincts supervisors trust, tone, hesitation, what went unsaid, were built for rooms with only people in them.
Supervisors hold their own uncertainty about systems they may never have used.
07
The wider context
AI enters as the context itself. Platforms set the norms, employers buy the tools, insurers price the risk. The room is arranged before anyone sits down.
When a platform’s defaults become the room’s rules, who consented to that?
Source and adaptation
Source: Hawkins and Shohet, Supervision in the Helping Professions, in print since the 1980s. Their diagram is a double matrix: two interlocking systems, client with therapist and therapist with supervisor, inside one wider context. We keep the seven eyes and ask a new question through each.
SIX ROLES
Six roles AI is growing into.
Name the role before judging the tool.
Role 01
A private personal tool
The client journals, rehearses and reflects with it. You see only what they choose to bring.
Ask: What would you want to know about it?
Role 02
Part of the care team
It carries information between client, practitioner and service, with consent.
Ask: Who reads what it writes?
Role 03
A professional assistant
Summaries, drafts, literature and documentation, under your review.
Ask: What do you still check by hand?
Role 04
A monitored digital intervention
A defined element of treatment, evaluated and supervised like any other.
Ask: What evidence would you require first?
Role 05
An emotional companion
Available at 3 a.m. Warm, patient, and agreeable if set that way.
Ask: What does it quietly replace?
Role 06
Infrastructure
Booking, triage, risk flags and notes. Invisible until it fails.
Ask: Who notices when it is wrong?
DISCUSSION 02 · 20 MINUTES
Design an ideal 2030 system.
Take one of the six roles. Design the version of it you would actually want in 2030.
Discuss together
Which role did you choose, and why that one?
What is the one failure that would make you withdraw it?
Where does the person’s consent live, and how real is it?
Your group’s output
01What it does
02What it must never do
03What the person controls
04When a human becomes responsible
How the 20 minutes run
2 minPick one role as a group
10 minDesign it against the four lines
5 minStress test: how could this hurt someone?
3 minBring it back to the room
Timer · 20 min20:00
Private notes on this device. Nothing is submitted.
Notes stay in this browser session and are never sent anywhere.
The systems of 2030 are being designed now, mostly without clinicians in the room. That is still a choice.
Discussion 2 of 2.
GUIDELINES
Seven guidelines for 2030.
Working principles for anyone building these systems.
01
Preserve agency.
Why it holds
The person decides. A system that narrows someone’s choices while feeling helpful has failed at its main job.
02
Keep uncertainty visible.
Why it holds
Fluent wording can hide shaky ground. A tool should show what it does not know.
03
Make influence understandable.
Why it holds
If a suggestion shaped a decision, the people affected should be able to see how.
04
Give people control over memory and data.
Why it holds
What is remembered, who can see it and how it is deleted belong to the person, not the platform.
05
Name the responsible human.
Why it holds
For every output that matters, someone with a name checks it, owns it and can stop it.
06
Protect relationships from invisible interference.
Why it holds
The alliance cannot defend itself against a third party it cannot see. Make the third party discussable.
07
Design ways to pause, leave and seek human help.
Why it holds
Pausing and leaving should be easy. Every system needs a door marked human.
BEFORE YOU LEAVE
Three moves, eight questions.
For every tool, proposal or system you meet from here on.
Move 1
Make the system visible
01
What does the AI observe, remember and infer?
02
Does the person understand what it is doing, why it produced this output and who else may receive it?
Move 2
Bound its authority
03
What may the AI recommend, decide or do?
04
What must it never do, even if doing so would be faster or more convenient?
05
Which named human is responsible when its output affects care?
Move 3
Preserve agency and a way back to people
06
What can the person correct, delete, pause, refuse or take with them?
07
How could the system change the client–practitioner relationship, or decide whose account is heard first?
08
What is the clear route to human help when the system is uncertain, fails or distress increases?
An original workshop discussion aid. Not a validated assessment, clinical protocol, safety score or legal compliance checklist.
Private notes on this device. Nothing is submitted.
One question from the eight I will start with
One tool I will walk through the three moves
One AI moment I will bring to supervision
Notes stay in this browser session and are never sent anywhere.
Randomized 210 adults with clinically significant depressive or anxiety symptoms or high risk for eating disorders to four weeks of Therabot or a waitlist; the intervention group showed greater symptom improvements and reported a therapeutic alliance.
Limitation: Waitlist rather than active-treatment control; four-week intervention; screened sample; researchers used crisis classifiers and human oversight; the trial does not establish equivalence to psychotherapy, long-term safety, or generalizability to general-purpose chatbots.
Compared a 12-week conversational AI intervention, face-to-face group therapy, and waitlist among 995 psychologically distressed university students in Israel; anxiety and well-being favored the AI arm over both comparators, while other outcomes were mixed.
Limitation: Restricted student sample, self-report outcomes, attrition, multiple outcomes, differences between intervention formats, and disclosed links to the platform limit generalization. Results do not prove that a general-purpose chatbot can safely provide care.
Analyzed aggregate data from adult Woebot users and found self-reported working-alliance and bond scores within days of use that were comparable with scores reported in some prior CBT studies.
Limitation: Self-selected respondents, observational design, cross-study comparison rather than random assignment, no evidence that the bond was equivalent in meaning or mechanism to a human therapeutic relationship, and all authors were affiliated with Woebot Health.
Across computational analysis and three preregistered experiments involving 2,405 participants, sycophantic AI increased perceived rightness, reduced intentions to repair interpersonal conflict, and increased preference, trust, and intended reliance on AI.
Limitation: The experiments measured judgments and intentions in bounded scenarios and live-chat interactions, not long-term behavior, clinical populations, or psychotherapy outcomes. The result should not be generalized to every model or interaction.
In a non-clinical trial with 300 peer supporters, just-in-time AI feedback increased expressed conversational empathy by 19.6% overall and more among supporters who reported difficulty providing support.
Limitation: The study evaluated written peer-support responses, not psychotherapy or patient outcomes. Humans decided whether and how to use the feedback, so it supports augmentation rather than autonomous care.
Randomized 238 outpatient physicians across 14 specialties to two ambient AI scribe tools or usual care and measured time in notes plus workload, burnout, safety, accuracy, and usability outcomes.
Limitation: The trial was conducted in outpatient medicine, not psychotherapy; effects differed by product and the primary metric captured time in notes rather than total care quality or patient outcomes.
Maps therapy guidance and experimentally probes several large language models, reporting stigmatizing patterns and inappropriate responses to some presentations involving delusions, suicidality, hallucinations, and mania.
Limitation: Model versions change quickly; benchmark prompts cannot reproduce the full context of care; the mapping emphasized selected U.S. and U.K. clinical materials and several CBT-derived manuals.
Advises psychologists to evaluate quality and appropriateness, protect confidentiality and consent, preserve professional judgment, monitor misinformation, and discontinue tools when concerns arise.
Limitation: U.S. professional guidance, not Singapore law; it is principles-based and does not validate any product.
Provides a five-step method for challenging assumptions, creating scenarios, stress-testing strategies, and developing actions across multiple possible futures.
Limitation: Public-policy toolkit rather than mental-health research; the workshop must adapt its methods to a short professional format.
Uses structured, evidence-informed scenarios based on critical uncertainties in capability, access, safety, adoption, and geopolitics and emphasizes stress-testing rather than probability estimates.
Limitation: Designed for UK public policy, not healthcare or Singapore; the scenarios are not mutually exclusive and must not be imported as predictions.
Documents a production full-duplex voice model that listens and speaks at the same time, handles interruptions and backchannels, retains context over longer sessions, and can delegate reasoning or tool calls to another model. OpenAI reports a 30-percentage-point gain over GPT-Realtime-2.1 on its Full Duplex Bench.
Limitation: The evaluations cover vendor benchmarks, customer service, banking support, and language tutoring. They do not establish emotional understanding, clinical safety, therapeutic benefit, or safe response during distress; the headline results are vendor-reported.
Describes foundation models integrated into consumer operating systems, including on-device models, multimodal capabilities, expressive voice, image understanding, tool use, and server models on Private Cloud Compute.
Limitation: Vendor-authored evaluation of its own models; availability, supported devices, real-world quality, and privacy properties vary by feature and deployment. It does not establish safe mental-health use.
Documents a deployed memory architecture that synthesizes context from past conversations, updates memories over time, and exposes a reviewable memory summary to users.
Limitation: A vendor report about one product. Memory quality is evaluated by the provider, does not imply complete recall or clinical understanding, and raises control, deletion, provenance, and stale-inference questions.
Documents Gemini Live conversations that can incorporate a phone camera feed or shared screen and reports rollout on Android and iOS.
Limitation: A vendor product announcement with varying availability and requirements. Google advises users to check responses for accuracy; visual access is not evidence of reliable emotion, intention, or mental-state inference.
Describes Project Astra research prototypes with native audio, live video understanding, tool use, memory of earlier conversation, and prototype smart glasses; the December 2024 system retained up to ten minutes of in-session memory and some cross-conversation memory.
Limitation: Project Astra and the glasses were research prototypes for trusted testers. The demonstration does not establish mass adoption, continuous reliability, privacy, or fitness for mental-health use.
Describes an approximately three-billion-parameter on-device model supporting text and image understanding and tool calling, with developer access through Apple’s Foundation Models framework.
Limitation: Apple notes that on-device models are smaller, have shorter context, and need simpler prompts than frontier cloud models. The work does not validate mental-health assessment or intervention.
Documents a deployed general-purpose agent that can break down multi-step goals, browse websites, run code, work with files, and take actions through connected tools, with interruption and confirmation points for consequential actions.
Limitation: A first-party account of a general knowledge-work product. It provides no evidence of safe autonomous mental-health care; tool access magnifies errors and creates exposure to prompt injection and misleading external content.
Presents a research Personal Health Agent using specialist data-scientist, health-domain, and health-coach agents to analyze wearable time series, questionnaires, and biomarkers; evaluation used more than 7,000 annotations and 1,100 hours of assessment across ten benchmark tasks.
Limitation: A Google-led preprint and research framework, not an available clinical product. Benchmarks do not establish improved health outcomes, safe autonomous care, or equivalence to a professional.
Randomized 238 outpatient physicians across 14 specialties to two ambient AI scribe tools or usual care and measured time in notes plus workload, burnout, safety, accuracy, and usability outcomes.
Limitation: The trial was conducted in outpatient medicine, not psychotherapy; effects differed by product and the primary metric captured time in notes rather than total care quality or patient outcomes.
Compared notes from 11 ambient AI scribe tools and 18 human note-takers across five standardized primary-care cases, scored by 30 blinded raters.
Limitation: Human-generated notes scored higher across the tested quality domains, but the cases were simulated, humans lacked normal time constraints, and the result reflects tools available at one point in a fast-changing market.
Reports a wearable foundation model trained on more than one trillion minutes of multimodal sensor data from five million consented participants, transferring across 35 health-prediction tasks and tested as grounding for a personal health agent.
Limitation: The health-agent test used 31 participant profiles and clinician ratings of generated summaries; it is not evidence of clinical outcomes, mental-health diagnosis, crisis prediction, or population-wide validity. This is also an organization-authored research summary.
Explains how phones and wearables enable scalable longitudinal measurement of sleep, activity, cardiovascular and behavioral variables, while setting research priorities for digital sensing in mental health.
Limitation: A perspective, not proof that digital sensing produces clinically useful mental-health phenotypes. The authors explicitly state that utility at scale has not been conclusively demonstrated and that much of the literature remains exploratory.
Synthesizes 11 prediction studies, 10 feasibility studies, and 7 protocols on smartphone and wearable sensing for suicidal thoughts and behaviors.
Limitation: The review found limited evidence, major methodological and reporting shortcomings, generally lower performance for passive than active data, and no incremental value from passive data in three of four relevant studies.
Studied 1,013 adults and related passive location, communication, and phone-use features to later depression and anxiety symptoms. More time at home relative to a person’s own baseline was associated with higher future depression scores.
Limitation: Full models explained only about 5–6% of symptom variance. Findings were correlational, time-lag dependent, affected by pandemic-era routines, and based on a demographically limited sample; passive data could not be shared publicly because of re-identification risk.
Followed 237 participants for four weeks using Fitbit Sense 2 data and reported group-level differences in tonic electrodermal activity, skin temperature, and heart rate between higher- and lower-symptom groups.
Limitation: The study used questionnaire-defined symptom groups and found correlates rather than a diagnostic test, causal effect, or reliable individual warning system. Most authors were affiliated with Google or Verily.
Randomized 210 adults with clinically significant depressive or anxiety symptoms or high risk for eating disorders to four weeks of Therabot or a waitlist; the intervention group showed greater symptom improvements and reported a therapeutic alliance.
Limitation: Waitlist rather than active-treatment control; four-week intervention; screened sample; researchers used crisis classifiers and human oversight; the trial does not establish equivalence to psychotherapy, long-term safety, or generalizability to general-purpose chatbots.
Compared a 12-week conversational AI intervention, face-to-face group therapy, and waitlist among 995 psychologically distressed university students in Israel; anxiety and well-being favored the AI arm over both comparators, while other outcomes were mixed.
Limitation: Restricted student sample, self-report outcomes, attrition, multiple outcomes, differences between intervention formats, and disclosed links to the platform limit generalization. Results do not prove that a general-purpose chatbot can safely provide care.
Analyzed aggregate data from adult Woebot users and found self-reported working-alliance and bond scores within days of use that were comparable with scores reported in some prior CBT studies.
Limitation: Self-selected respondents, observational design, cross-study comparison rather than random assignment, no evidence that the bond was equivalent in meaning or mechanism to a human therapeutic relationship, and all authors were affiliated with Woebot Health.
Combined automated analysis of nearly 40 million interactions with surveys and a four-week randomized study of nearly 1,000 adults to examine affective use, voice modality, loneliness, social interaction, emotional dependence, and problematic use.
Limitation: Not peer reviewed at publication; one platform, U.S. adults, English conversations, self-report measures, imperfect classifiers, short duration, and many associations that are not causal. Emotional engagement was rare overall and concentrated among a small subgroup.
Maps therapy guidance and experimentally probes several large language models, reporting stigmatizing patterns and inappropriate responses to some presentations involving delusions, suicidality, hallucinations, and mania.
Limitation: Model versions change quickly; benchmark prompts cannot reproduce the full context of care; the mapping emphasized selected U.S. and U.K. clinical materials and several CBT-derived manuals.
In a non-clinical trial with 300 peer supporters, just-in-time AI feedback increased expressed conversational empathy by 19.6% overall and more among supporters who reported difficulty providing support.
Limitation: The study evaluated written peer-support responses, not psychotherapy or patient outcomes. Humans decided whether and how to use the feedback, so it supports augmentation rather than autonomous care.
Across four studies, participants generally preferred receiving empathy from humans while rating AI-generated empathetic responses as higher quality and more effective at making them feel heard when they encountered them.
Limitation: Mostly decontextualized, short-form empathy judgments rather than ongoing therapeutic relationships. The authors explicitly call for research on repeated interactions and known human relationships.
Across computational analysis and three preregistered experiments involving 2,405 participants, sycophantic AI increased perceived rightness, reduced intentions to repair interpersonal conflict, and increased preference, trust, and intended reliance on AI.
Limitation: The experiments measured judgments and intentions in bounded scenarios and live-chat interactions, not long-term behavior, clinical populations, or psychotherapy outcomes. The result should not be generalized to every model or interaction.
Sets out more than 40 recommendations for governments, developers, and health organizations and identifies health uses including clinical care, patient-guided use, administration, education, and research.
Limitation: Guidance rather than binding law or effectiveness evidence; implementation depends on jurisdiction, resources, use case, and subsequent regulation.
Provides lifecycle guidance for developers, healthcare organizations, and professionals; addresses clinical and clinical-operations AI, transparency, testing, monitoring, hallucination, data disclosure, model drift, and responsibility.
Limitation: Good-practice guidance that complements, rather than replaces, applicable law and device regulation. Its scope focuses on healthcare AI and does not automatically cover every consumer wellness chatbot.
Defines common agent features as independent multi-step planning, decision-making, and action-taking, and recommends bounding permissions, assigning meaningful human accountability, implementing controls and monitoring, and enabling end-user responsibility.
Limitation: Cross-sector voluntary framework, not clinical evidence or a mental-health-specific standard. It describes emerging practice in a rapidly changing field.
Advises psychologists to evaluate quality and appropriateness, protect confidentiality and consent, preserve professional judgment, monitor misinformation, and discontinue tools when concerns arise.
Limitation: U.S. professional guidance, not Singapore law; it is principles-based and does not validate any product.
Reviews consumer use of general-purpose chatbots and wellness apps for mental-health support and recommends that they not replace qualified care, while recognizing possible adjunctive roles for some purpose-built tools.
Limitation: Advisory rather than systematic review or regulation; U.S. legal context; product categories and evidence evolve quickly.
Organizes generative-AI risk management around governance, pre-deployment testing, content provenance, and incident disclosure, with actions for additional review, documentation, and oversight.
Limitation: Voluntary, cross-sector U.S. framework; not a substitute for clinical validation, professional ethics, local law, or product-specific assessment.
Terms used in the room
Therapeutic alliance
The working relationship between client and therapist: trust, agreement on goals and a felt bond. The strongest predictor of outcome across therapies.
The frame
The agreed boundaries of therapy: time, confidentiality, contact between sessions, consent. AI use now belongs in it.
Countertransference
The therapist’s own emotional reactions to the client, used as clinical information. Now includes reactions to the client’s AI.
Seven-eyed model
Hawkins and Shohet’s supervision framework: seven ways of looking at the client, the therapist, the relationships between them and the wider context.
Formulation
A shared working hypothesis about a client’s difficulties, built together and revised over time.
Triadic relationship
Three parties instead of two: client, therapist and an AI system each of them uses.
Digital phenotyping
Inferring mental state from passive data such as sleep, movement and phone use.
Sycophancy
The tendency of AI systems to agree and flatter, even when agreement is unhelpful.
Ambient scribe
Software that listens to a clinical encounter and drafts the note.
Clinical governance
The structures through which an organisation keeps care safe and accountable.
The shape of the session · 150-minute workshop · Singapore