Read the studies and guidance behind the discussion. Findings, advice and original workshop prompts are labelled separately.
E1 Sycophantic AI decreases prosocial intentions and promotes dependence
Myra Cheng, Cinoo Lee, Pranav Khadpe, Sunny Yu, Dyllan Han, Dan Jurafsky · Science / American Association for the Advancement of Science · 26 March 2026 · Model evaluation and three preregistered experiments · Research
A response that clients prefer can still narrow reflection. Ask what the advice encourages them to believe and do, alongside how supported it makes them feel.
What this does not establish: Intentions and ratings are not actual relationship outcomes or psychiatric dependency. Do not generalize tested responses to every reassurance exchange.
Read the original ↗ · Public abstract ↗
Study context
- Population
- 2,405 experimental participants; a separate evaluation tested 11 models.
- Intervention
- Exposure to sycophantic, overly agreeing responses.
- Comparator
- Less sycophantic responses in experiments; human responses in the separate model evaluation.
- Duration
- Experimental interactions with subsequent judgments; no longitudinal clinical follow-up established here.
- Outcome
- Greater perceived rightness, lower conflict-repair intentions, and greater trust/preference for sycophantic responses.
Final publication: 3 experiments, N=2,405. Do not substitute the older preprint's 2 experiments/N=1,604.
Final abstract does not provide disclosure details; no conflict assertion is made.
Final publisher/PubMed abstracts verified. Full-text arm breakdowns are not used.
E2 Expressing stigma and inappropriate responses prevents LLMs from safely replacing mental health providers
Jared Moore, Declan Grabb, William Agnew, Kevin Klyman, Stevie Chancellor, Desmond C. Ong, Nick Haber · ACM Conference on Fairness, Accountability, and Transparency · June 2025 · Therapy-guidance mapping and constructed-response experiments · Research
Check what a fluent reply assumes, omits and encourages. A polished tone does not establish that the response is appropriate to the person's situation.
What this does not establish: Constructed tests do not estimate real-world harm prevalence. The authors' philosophical account of therapeutic alliance is distinct from their measured findings.
Read the original ↗ · Open-access full text ↗
Study context
- Population
- Selected general models and therapy bots; a comparison involved 16 US therapists.
- Intervention
- Mental-health scenarios probing stigma and appropriateness of responses.
- Comparator
- Responses across model versions, contexts, symptom presentations and a therapist comparison.
- Duration
- Response experiments; no patient treatment period.
- Outcome
- Stigma and inappropriate replies appeared in tested systems, including responses to delusions and suicidal ideation.
Authors limit the work to systems resembling those tested in April 2025, not arbitrary future AI.
No conflict claim made; conference-hosted paper consulted.
Full text verified; paper states CC BY-SA 4.0. Link only in this kit.
E3 Evidence of Human-Level Bonds Established With a Digital Conversational Agent: Cross-sectional, Retrospective Observational Study
Alison Darcy, Jade Daniels, David Salinger, Paul Wicks, Athena Robinson · JMIR Formative Research / JMIR Publications · 11 May 2021 · Retrospective cross-sectional observational study · Research
A client's felt connection with a chatbot can be meaningful to them. Explore what that connection provides without treating a bond rating as proof of equivalent therapy.
What this does not establish: Self-selected respondents and developer-led analysis. No causal symptom outcome or randomized therapist equivalence test.
Read the original ↗ · Open-access full text ↗
Study context
- Population
- 36,070 adult respondents, selected from 177,212 eligible Woebot registrants.
- Intervention
- Naturalistic use of structured CBT-based Woebot; no randomized intervention.
- Comparator
- Descriptive comparisons with separate published CBT studies; no concurrent therapist arm.
- Duration
- Questionnaires completed within five days of first app use.
- Outcome
- Mean adapted WAI-SR bond subscale 3.84/5 (SD 1.0); abstract rounds this to 3.8.
Earlier structured chatbot; not a study of all generative AI systems.
Authors disclosed Woebot employment, stock options or fees.
Open publisher PDF verified.
E4 Human–AI collaboration enables more empathic conversations in text-based peer-to-peer mental health support
Ashish Sharma, Inna W. Lin, Adam S. Miner, David C. Atkins, Tim Althoff · Nature Machine Intelligence / Springer Nature · 23 January 2023 · Nonclinical randomized study · Research
AI can help people improve supportive writing. Keep the human able to accept, edit or reject suggestions, and distinguish a better draft from a demonstrated improvement in care.
What this does not establish: No recipient symptom or long-term alliance endpoint. Harm-related posts were filtered; peer supporters were not professional therapists.
Read the original ↗ · Open-access full text ↗
Study context
- Population
- 300 TalkLife peer supporters; 150 participants per condition.
- Intervention
- HAILEY's editable suggestions and feedback while composing responses.
- Comparator
- Writing without AI feedback; both arms received initial empathy training.
- Duration
- Each participant wrote replies to ten posts in one study task outside the live platform.
- Outcome
- Expressed-empathy evaluation favored assistance; reported overall relative increase 19.6%.
Human and automated assessments concerned expressed empathy, not the original support seeker's experience.
No no-conflict assertion; disclosure section not relied on for teaching.
Publisher abstract and indexed author-hosted manuscript verified; direct author PDF retrieval may be intermittent.
E5 Randomized Trial of a Generative AI Chatbot for Mental Health Treatment
Michael V. Heinz, Daniel M. Mackin, Brianna M. Trudeau, Sukanya Bhattacharya, Yinzhou Wang, Haley A. Banta, Abi D. Jewett, Abigail J. Salzhauer, Tess Z. Griffin, Nicholas C. Jacobson · NEJM AI / Massachusetts Medical Society · 27 March 2025 · Randomized controlled trial · Research
A promising trial supports a claim about the tested intervention, participants and comparison. It does not automatically validate another chatbot or establish equivalence to psychotherapy.
What this does not establish: Screening-defined symptoms; no randomized psychotherapy comparator. No inference to other products, acute crisis care or longer-term outcomes.
Read the original ↗ · Public abstract ↗
Study context
- Population
- 210 adults screened for clinically significant depression/anxiety symptoms or high feeding/eating-disorder risk.
- Intervention
- Expert-fine-tuned Therabot, n=106.
- Comparator
- Waitlist without app access during the study, n=104.
- Duration
- Four-week intervention; primary symptom-change assessments at weeks four and eight.
- Outcome
- Greater improvement on the three studied symptom domains versus waitlist. Publisher abstract reports four-week depression-score changes −6.13 versus −2.63; subgroup scores, not all-participant recovery rates.
Keep the effect-size and percentage-recovery headlines out of the presentation.
Dartmouth-funded developer research; detailed disclosure forms not evaluated here.
Publisher abstract reverified; full-text access may require subscription. Numerical endpoint optional and excluded from classroom copy.
E6 Efficacy of a Conversational AI Agent for Psychiatric Symptoms and Digital Therapeutic Alliance: A Randomized Clinical Trial
Anat Shoshani, Bar Gurfinkel, Ariel Kor, Yael Ben-Haim, Or Kanarek, Romi Segev, Or Shafir, Romi Arbel · JAMA Network Open / American Medical Association · 14 April 2026 · Three-arm randomized clinical trial · Research
Some newer trials include active comparisons. Examine the particular format and outcomes before treating a result against group therapy as evidence about every form of psychotherapy.
What this does not establish: Self-report; restricted student sample; roughly one-third follow-up attrition.
Read the original ↗ · Open-access full text ↗
Study context
- Population
- 995 distressed Hebrew-speaking Israeli university students, ages 18–35; severe disorder, acute risk and current treatment excluded.
- Intervention
- Kai conversational platform, n=336, with escalation allowing clinician intervention.
- Comparator
- Psychologist-led group therapy, n=331; waitlist, n=328.
- Duration
- Twelve weeks; group therapy weekly for 90 minutes; three-month follow-up.
- Outcome
- Primary measures: GAD-7, PHQ-9, brief PTSD checklist, WHO-5 and life satisfaction. Anxiety/well-being favored AI over both comparators; depression favored AI over waitlist, not group therapy after multiplicity adjustment. No PTSD group difference.
Do not claim depression superiority versus group therapy from the unadjusted confidence interval.
Shoshani: fees/options; Gurfinkel: fees/employment at KAI.AI.
Open full text verified; link, do not republish figures.
E7 Ambient AI Scribes in Clinical Practice: A Randomized Trial
Paul J. Lukac, William Turner, Sitaram Vangala, Aaron T. Chin, Joshua Khalili, Ya-Chen Tina Shih, Catherine Sarkisian, Eric M. Cheng, John N. Mafi · NEJM AI / Massachusetts Medical Society · 26 November 2025 · Pragmatic randomized trial · Research
Test the actual workflow, including the review step. A reduction in one writing metric does not establish better notes or an equivalent reduction in total administrative work.
What this does not establish: One system; brief trial; not psychotherapy-specific. Metric excludes editing inside the scribe application; registration completed after commencement.
Read the original ↗ · Open-access full text ↗
Study context
- Population
- 238 UCLA outpatient physicians across 14 specialties; English-only visits.
- Intervention
- DAX, n=79, or Nabla, n=79.
- Comparator
- Usual care, n=80.
- Duration
- 4 November 2024–3 January 2025; primary comparison used the second intervention month versus baseline.
- Outcome
- Primary endpoint: EHR time-in-note. Nabla −9.5% versus control (95% CI −17.2% to −1.8%); DAX −1.7% (−9.4% to +5.9%), nonsignificant. Occasional important inaccuracies were reported.
Well-being endpoints were secondary.
UCLA-funded; supplementary author disclosures linked in manuscript, not audited here.
Open author manuscript verified.
E8 Rapid Evaluation of Artificial Intelligence Technology Used for Ambient Dictation in Primary Care: Comparing the Quality of Documentation of Artificial Intelligence-Generated and Human-Produced Clinical Notes
Ashok Reddy, Eric Gunnink, Chelle L. Wheat, Scott Pawlikowski, Chína M. Payne, Scott Wiltz, Terrence L. Hubert, Susan Kirsh, Evan Carey, Donna Hill, Karin M. Nelson · Annals of Internal Medicine / American College of Physicians · 17 April 2026 · Cross-sectional simulated-case evaluation · Research
Review a generated note for its clinical meaning, not just its readability. A time-saving tool still needs a separate check for omissions, attribution and unsupported statements.
What this does not establish: Simulation; humans lacked normal time constraints. Not a comparison of every real-world, clinician-edited AI note.
Read the original ↗ · Public abstract ↗
Study context
- Population
- Five standardized primary-care audio cases; 11 AI scribes, 18 human note takers, 30 blinded raters.
- Intervention
- AI-produced encounter notes.
- Comparator
- Human-produced notes from the same audio cases.
- Duration
- One evaluation of standardized cases; no clinical follow-up.
- Outcome
- Modified documentation-quality instrument: 10 domains, maximum 50. Human notes scored higher overall across all five cases.
Not psychotherapy-specific; complements E7's different endpoint.
VHA funding reported; detailed author disclosures not verified from abstract.
Primary PubMed abstract verified; no open full-text URL confirmed.
E9 Delivering Cognitive Behavior Therapy to Young Adults With Symptoms of Depression and Anxiety Using a Fully Automated Conversational Agent (Woebot): A Randomized Controlled Trial
Kathleen Kara Fitzpatrick, Alison Darcy, Molly Vierhile · JMIR Mental Health / JMIR Publications · 6 June 2017 · Unblinded randomized controlled trial · Research
Evidence follows the studied intervention, not the word chatbot. Before applying an older finding, check the tool's design, intended use and the people involved in the study.
What this does not establish: Small, short, unblinded trial; unequal attrition. Earlier structured chatbot, not a modern open-ended generative model.
Read the original ↗ · Open-access full text ↗
Study context
- Population
- 70 young adults recruited through a US university community with self-identified depression/anxiety symptoms.
- Intervention
- CBT-oriented Woebot, n=34.
- Comparator
- NIMH information ebook, n=36.
- Duration
- Two weeks; no longer-term follow-up.
- Outcome
- PHQ-9 favored Woebot in intention-to-treat analysis (P=.01). Both groups' GAD-7 improved among completers; this does not establish chatbot-specific anxiety benefit.
Useful historical example, not a current product recommendation.
Darcy founded Woebot Labs; company funded participant incentives.
Open publisher full text verified.
P1 Artificial Intelligence in Healthcare Guidelines (AIHGle 2.0)
Singapore Ministry of Health and Health Sciences Authority · March 2026 · Official healthcare AI guidance · Singapore guidance
Sets out responsibilities for AI developers, deploying organisations and healthcare professionals, including input/output review and contextual patient communication.
What this does not establish: Guidance accompanies applicable laws, codes and organisational policy; it is not a universal statutory checklist or product endorsement.
Read the original ↗
Study context
Discuss what responsibilities remain when AI contributes to professional work.
Official 42-page PDF opened directly on 13 September 2026. Relevant printed pages: 7, 28–32. Published March 2026; launch date 10 March 2026 supported by official HSA material.
P2 Data security requirements apply to AI tools that process patient data
Singapore Ministry of Health · 4 August 2026 · Official parliamentary answer · Singapore guidance
MOH states that data-security requirements apply to AI processing patient data on cloud services or on premises and separately describes additional public-healthcare practices.
What this does not establish: Public-healthcare non-retention commitments cannot be assumed for private practices or consumer accounts. The answer does not determine a particular organisation's compliance.
Read the original ↗
Study context
Explain why hosting location or a privacy setting alone does not resolve data governance.
Official page opened directly on 13 September 2026; answer paragraphs 1–2 checked.
P3 Digital Health
Singapore Health Sciences Authority · Date not stated · Official regulatory overview · Singapore guidance
Explains that intended medical purposes generally bring digital health tools within medical-device regulation and links to relevant software-device and AI guidance.
What this does not establish: Does not establish any named product's registration, exemption, effectiveness or suitability. Singapore requirements should not be replaced with US assumptions.
Read the original ↗
Study context
Ask what the product claims to do and which intended use is being evaluated.
Current official page opened directly on 13 September 2026. No specific publication date recorded; site footer date is not treated as publication date.
P4 Advisory Guidelines for the Healthcare Sector
Singapore Personal Data Protection Commission · 20 September 2023 · Official sector guidance on PDPA interpretation · Singapore guidance
Addresses healthcare data collection, use and disclosure, including consent and exceptions, protection, retention, transfers, breach notification and accountability.
What this does not establish: Consent alone does not determine lawful or appropriate use. Applicable duties depend on the organisation, activity and arrangements; this workshop does not provide legal determinations.
Read the original ↗
Study context
Map what information goes where and identify questions for the responsible organisation or data-protection lead.
Revision date and relevant topics checked against indexed text from PDPC's official PDF and landing page on 13 September 2026. Direct reads returned sparse text; full direct PDF extraction was unavailable.
P5 Ethics and governance of artificial intelligence for health: Guidance on large multi-modal models
World Health Organization · 2024 · International health-governance guidance · Professional resources
Describes generative AI applications in health and risks including inaccurate output, bias, automation bias and cybersecurity concerns.
What this does not establish: Potential applications are not evidence that a specific tool is effective, safe or suitable. WHO guidance is not Singapore legislation.
Read the original ↗
Study context
Introduce why fluent output still needs verification for its intended purpose.
Official publication page and WHO's 18 January 2024 summary were checked on 13 September 2026. Summary URL: https://www.who.int/news/item/18-01-2024-who-releases-ai-ethics-and-governance-guidance-for-large-multi-modal-models
P6 The App Evaluation Model
American Psychiatric Association · Date not stated · Professional app evaluation framework · Professional resources
Six steps: Background; Access; Privacy and Security; Clinical Foundation; Usability; Integration toward Patient-Centered Goals.
What this does not establish: Structured inquiry does not certify safety or legal compliance. The workshop's five-question aid is original and is not a validated version of this model.
Read the original ↗
Study context
Show a professional framework for asking product, evidence, privacy, usability and care-fit questions.
Official page re-opened directly on 13 September 2026. Full step headings checked; navigation abbreviates the last as Data Integration. No publication date displayed.
P7 Use of generative AI chatbots and wellness applications for mental health
American Psychological Association · November 2025 · Professional health advisory · Professional resources
Recommends against replacing qualified care with chatbots or wellness apps and distinguishes mental-health-specific tools from general-purpose chatbot evidence.
What this does not establish: US professional advice is not Singapore law. Cite individual studies for numerical effects and do not generalise all tools' capabilities or risks.
Read the original ↗
Study context
Separate feeling helped, research evidence and suitability for a particular person and setting.
Relevant text and November 2025 creation date checked through indexed official APA text on 13 September 2026. Direct page reads returned only a shell.
P8 Discussing AI use in therapy
American Psychological Association · 16 June 2026 · Professional practice article with expert commentary · Professional resources
Encourages discussing clients' AI use, understanding its role, exploring concerns collaboratively and maintaining ongoing dialogue.
What this does not establish: Expert commentary is not a validated clinical protocol or evidence of treatment effects. US psychologist survey figures do not establish Singapore client prevalence.
Read the original ↗
Study context
Rehearse a curious opening conversation about AI use using a fictional vignette.
Official-domain indexed article text confirms Zara Abrams as author and Date created: June 16, 2026. Rechecked 13 September 2026. Direct access, including the .html version, returned a one-line shell; no full direct-page read.