Overview

This lecture builds a basic understanding of what AI is and how it works, surveys the opportunities AI creates in healthcare, and then works through the ethical challenges this raises, using AI medical scribes as the running case study for applying a critical-constructive approach. It opens with a cartoon satirising AI sycophancy (a robot surgeon “correcting” itself to agree with a patient’s factually wrong statement) and closes with one on liability (a human doctor shrugging off blame that has shifted to an AI radiology device), framing the two poles of the lecture: AI’s fallibility, and who answers for it.

What AI Is and What Types Are Used in Healthcare

  • AI: “technologies with the ability to perform tasks that would otherwise require human intelligence”, with strong “information-processing capacities” - good at processing large amounts of different types of data (statistical systems).
  • Symbolic/rule-centric AI (GOFAI): works like a decision tree/flow chart, reasoning through information as an expert system. Transparent and interpretable.
  • Neural networks (machine learning/deep learning): interconnected networks of simple units. Black box phenomenon - programmers know the network’s architecture but not precisely what happens in its intermediate layers or how it reaches a decision.
  • Machine learning is still a statistical system: the underlying task is usually pattern recognition, using identified patterns to explain data and predict future data. Supervised learning uses training data with specified variables; unsupervised learning has categories that are not predefined. Correlation is not causation.
  • Large Language Models (LLMs): neural networks trained on large amounts of text for next-word prediction; a form of generative AI and foundation models, often deployed as chatbots.
  • Criticism of the LLM industry (framed as “rockets vs bicycles”, referencing Karen Hao’s Empire of AI): exploitative, resource-intensive, monopolises knowledge production, driven by large corporations (Silicon Valley), and prone to AGI scaremongering.

Opportunities of AI in Healthcare

Five categories, each with supporting evidence/examples:

  1. Clinical documentation and administrative support - generating patient summaries and classification lists; booking/cancelling appointments via chatbots, voice assistants, intelligent scheduling; AI scribes (e.g. Nabla Co-pilot, Heidi Health, iMedX) listen to a consultation, transcribe it, and produce an organised note (demo: audio encrypted during transcription and not stored, transcript shared with privately hosted AI models; output must be reviewed before use).
  2. Risk prediction - processing large volumes of EHR data to predict risk of illness, post-operative complications, and death, e.g. predicting chronic kidney disease risk/decline (Bouderhem et al. 2024), NLP identifying suicidal ideation for suicide prevention (Arowosegbe et al. 2023), ChatGPT performing comparably to traditional methods for cardiovascular disease risk prediction (Han et al. 2024).
  3. Diagnostic and triage support - AI chatbots/VR doctors triaging symptoms; deep learning for diabetic retinopathy diagnosis from images validated in real-world settings (Topol 2019; Romero-Aroca et al. 2020), extending screening to remote/rural areas or areas with staff shortages.
  4. Precision medicine and therapeutic support - generative AI predicting which treatment best suits an individual from genetic information, pathology reports, imaging, lifestyle and other data (Reddy et al. 2024), e.g. enhancing cancer treatment or optimising drug dosing for epilepsy (Bouderhem et al. 2024).
  5. Self-management and patient empowerment - chatbots, self-monitoring and risk prediction tools, and technologies supporting people with disabilities (WHO 2021); patient education content at different reading levels and in multiple languages, improving health literacy (Reddy et al. 2024); virtual coaching giving individualised advice from medical literature and patient data for chronic condition management (Topol 2019).

Current examples cited in primary care: Heidi (AI scribe) and OpenEvidence (medical evidence search engine).

Ethical Challenges: The Full Landscape

Full list of clinical-ethical challenges raised by (gen)AI use: overhyping/lack of evidence, transparency/explainability, data and algorithmic bias, hallucinations, equity, risk of group harm, Māori data sovereignty and governance, impact on clinical reasoning, impact on the patient-provider relationship, automation bias, deskilling, alert fatigue, accessibility, data security (breaches, hacking), privacy/confidentiality/data sharing, training data quality, data leakage/model drift, AI sycophancy, AI slop, informed consent, accountability/liability, (lack of) regulatory frameworks, and lack of appropriate evaluation methods.

The lecture focuses on six of these in depth: overhyping/lack of evidence, transparency/explainability, data and algorithmic bias, hallucinations, impact on clinical reasoning, and automation bias.

Hype and the Evidence Gap

  • Media hype: headlines claim AI is “starting to beat doctors” and “outperforms doctors in clinical reasoning tests”, but accompanying commentary (BMJ) urges caution about how these systems will perform in real clinical decision-making.
  • Marketing hype (Heidi advertisement): promises restored eye contact, warmer care, getting home on time; claims solo practitioners save up to 2 hours/day on documentation, some customers cut charting time by 70% and recouped over $10,000 in 12 weeks.
  • Modality Partnership (largest NHS GP super-partnership) pilot of Heidi across 47 GPs and 2,800+ consultations reported: 51% drop in documentation time during appointments, 61% decrease in after-hours admin, 58% reduction in documentation-related stress, 78% reporting better patient rapport, 78% reporting reduced cognitive load. Scaled to 200+ clinicians across 53 surgeries; Heidi is now used by 1 in 2 UK GPs, across 15 NHS trusts, supporting 1.5 million NHS appointments/month.
  • Political rollout (New Zealand): AI scribe live in every Emergency Department nationwide, 1,250 clinicians using it, 1,000 further licences being added for mental health teams.
  • Political critique: in an RNZ interview, GPs told John Campbell that Heidi is useful but does not increase the number of patients seen in a fixed appointment slot; Labour’s Ayesha Verrall countered that GPs reported it frees up roughly 3 appointments/day, which formed the basis of Labour’s calculations.
  • Rigorous evidence is more modest than the marketing suggests:
    • JAMIA Open simulation study (9 primary care physicians, 4 simulated encounters each): documentation occupied 36.3% of the encounter without an AI scribe vs 11.2% with one - a 69.1% reduction in documentation time.
    • JAMA multisite study (8,581 clinicians, 1,809 AI scribe adopters): adoption associated with 13.4 fewer minutes of EHR time, 16.0 fewer minutes of documentation time, and 0.49 additional weekly visits; effects greatest for primary care specialists, advanced practice clinicians, female clinicians, and heavier users; EHR time outside scheduled hours did not change significantly.
    • NEJM AI RCT of ambient AI (71,487 notes, 38% AI-generated): significant reduction in work exhaustion/interpersonal disengagement, a nonsignificant increase in professional fulfilment, documentation time down 0.36 hours/day, improved diagnostic billing codes, documentation quality maintained (PDSQI-9 scores 3.97-4.99/5), no software drift detected. [flag: funding statement at the bottom of this slide is cut off]
    • NEJM AI RCT comparing two scribes, DAX and Nabla (24,696 and 23,653 visits respectively): Nabla gave a significant 9.5% decrease in time-in-note vs control; DAX showed no significant change. Both showed improvements on burnout/task-load/work-exhaustion scores (Mini-Z, PTL, PFI-WE) vs control. One mild adverse event reported; clinically significant inaccuracies were noted “occasionally” for both scribes, flagged as requiring ongoing vigilance. [flag: funding statement at the bottom of this slide is cut off]
  • Why hype is a problem: technologies are introduced without robust evidence (lacking in test settings, in the real world, or of poor quality); this evidence gap guides policy decisions; regulation for software as a medical device is thin on both evidence and ethics; there is a risk of overpromising and underdelivering.
  • Six tensions between bioethical values and AI innovation pressures: do no harm vs risk; confidentiality vs data sharing; bioethical principles vs profit-making; slow innovation vs being first on the market; realistic outcomes vs great expectations; equity and inclusion vs early adopters moving to wide use.

Explainability and the Black Box Problem

  • With deep learning/neural networks, when a mistake is made it is very difficult or impossible to understand and investigate why, and to rectify it.
  • One alternative is to use only symbolic AI (decision trees) in medicine, but this risks decreased efficiency and misses out on the potential benefits of neural networks and LLMs.
  • The resulting value judgement: is explainability - the ability to understand and correct errors - more important than using more advanced AI that could potentially save more lives and improve wellbeing for more people?
  • High-quality studies and evidence about outputs may offset the explainability issue, but this leaves open whether correction of errors is actually possible. (Kerasidou 2021)

Data and Algorithmic Bias

  • Incomplete or biased training datasets can lead to unfair outcomes: biased datasets may perpetuate systemic inequities based on race, gender identity and other demographic characteristics, and may also limit AI’s performance as a diagnostic/treatment tool through lack of generalisability. This produces “data gaps” that particularly affect Māori and Indigenous populations, with a risk of group harm. (Kerasidou 2021; Murphy et al. 2021)
  • Bright side: AI can also be used to identify and correct existing bias.
  • New Zealand case study, “Could AI solve bias?”: the Equity Adjustor Tool prioritised Māori and Pacific elective surgery patients on waitlists ahead of others with identical clinical factors, weighting six areas - clinical specialty, clinical priority, time on waitlist, geographic (isolated) location, ethnicity (Māori/Pacific), and deprivation level. It was developed with Māori health, Pacific health and surgical services leadership, and rolled out at Te Toka Tumai Auckland, the wider Northern region, and (via a similar tool) the Southern district. In August 2024 its use was stopped, despite an independent review recommending it continue - illustrating how a bias-correcting algorithm can be politically contested and withdrawn even with positive evidence behind it.

Hallucinations, Automation Bias and Impact on Clinical Reasoning

  • Hallucinations: AI scribes can generate incorrect or misleading content. Example (Dr Medlicott, quoted in an article by Emily Cavana): an AI scribe documented that a young transgender patient wanted to join the military - a hallucination prompting the question of what bias produced it. The same clinician found the scribe (Nabla) ignored details about pad/tampon use when checking a patient with menorrhagia, raising a question of possible sexism in the tool. On the day this was observed the clinician had time to notice and correct the errors, prompting the question of what happens on a busier day.
  • Automation bias / impact on clinical reasoning: the same clinician described using her own note-taking to structure her consultations - narrating to the patient as she types, which both summarises and checks what she has heard, and does the same for the management plan. She contrasted this with her own past habits of sometimes forgetting to write notes at all, or returning to half-finished notes months later unsure what she had been thinking - illustrating the human failure mode an AI scribe might fix, but at the possible cost of losing the reasoning-structuring value of manual note-taking.

AI Scribes: Weighing Benefits Against Risks

Potential benefits:

  • Lets the GP focus on the conversation with the patient, improving patient care and the patient-doctor relationship.
  • Saves note-taking time, reducing unpaid work.
  • Reduces cognitive load.
  • Improved accuracy and detail, less likely to miss important information in the note.
  • Produces structured notes that facilitate sharing with other providers and with the patient (limitation: does not add coded information to the record, e.g. classification lists or medication lists).

Clinical-ethical risks and challenges:

  • Introduced without robust real-world evidence, treated as a low-risk “admin” tool, leaving open questions about benefits vs harms.
  • Unanswered questions: note quality; hallucination/error rate; impact on clinical reasoning; negative effect on recall; whether longer notes add reading time; what actually counts as “good use” (currently more time is gained with poor use than with careful use).
  • Likely to be most useful to some doctors, for some consultations, rather than universally beneficial.
  • Draws on the lecturer’s own publication, “The elephant in the room: a postphenomenological view on the electronic health record and its impact on the clinical encounter” (Moerenhout, Fischer & Devisch, 2020), on how documentation technology reshapes the clinical encounter.

Summary of Main Ethical Considerations

  1. AI hallucinations/bias: AI scribes may generate incorrect or misleading information, producing errors in the notes.
  2. Privacy/confidentiality: what level of data protection is in place, including the risk of reidentification.
  3. Responsibility/liability: who is responsible for AI-generated errors (the liability cartoon closing the lecture shows a human doctor shrugging off blame for a missed diagnosis onto the AI - “comes with the job”).
  4. Māori data sovereignty: cultural safety questions - has the AI scribe been adjusted to the local context, trained in te reo Māori, and does it meet Māori data sovereignty principles?
  5. Hidden work/added burden: time-saving benefits may be lost if the system needs a high level of scrutiny, or if reading the resulting notes takes longer.
  6. Lack of evidence: most AI scribes have not undergone robust studies of their performance, error rate, etc.
  7. Impact on clinical reasoning/recall: where note-taking is used to structure diagnostic reasoning, what happens to that clinical skill when an AI scribe takes it over?

Conclusion: AI Is Not a Miracle Cure

  • We need to make sure the tools we develop support our values and priorities: what do we actually want, and how can AI tools support the ethical principles underpinning healthcare?
  • There are currently no clear guidelines or robust regulatory frameworks, so clinicians need to conduct their own critical analysis before use.
  • Avoid hype: a critical-constructive assessment of AI technology helps improve its design and implementation - ask critical questions.
  • Education, including patient engagement, is needed.
  • A call for “slow innovation”: not “move fast and break things”, but move slowly enough to avoid harm, while making sure the benefits are not missed.

Self-test

  1. Give the two definitions of artificial intelligence presented in the lecture.
  2. Distinguish symbolic/rule-centric AI (GOFAI) from neural-network-based AI, including what the “black box” problem refers to.
  3. What are large language models, and what four criticisms does the lecture raise against the LLM industry?
  4. Explain the difference between supervised and unsupervised machine learning, and why the lecture stresses that correlation is not causation.
  5. List the five categories of opportunity for AI in healthcare, with one concrete example for each.
  6. Compare the marketing claims made for AI scribes (e.g. the Heidi advertisement and the Modality Partnership pilot) with the effect sizes found in the four peer-reviewed studies cited (JAMIA Open, JAMA, and the two NEJM AI trials).
  7. Explain why the lecturer treats “hype” itself as an ethical problem, not just an exaggeration.
  8. Describe the six tensions the lecture identifies between bioethical values and AI innovation/market pressures.
  9. Explain the explainability/black box problem and the value judgement it forces between explainability and more advanced but opaque AI.
  10. Describe how data and algorithmic bias arise in AI systems, and explain what the New Zealand Equity Adjustor Tool case illustrates about using algorithms to counter bias.
  11. Describe the two hallucination examples from the AI scribe (Nabla) case reported by Dr Medlicott, and explain what bias each one might suggest.
  12. Explain what automation bias and impact on clinical reasoning mean in the context of AI scribes, using the clinician’s own note-taking habits as illustration.
  13. List the potential benefits of AI scribes and the clinical-ethical risks/unanswered questions raised against them.
  14. List the seven main ethical considerations in the lecture’s closing summary.
  15. What does the lecturer mean by a “call for slow innovation”, and what critical-constructive approach does she recommend before adopting AI scribes?
  16. Integrative: a GP is deciding whether to adopt an AI scribe. Using concepts from this lecture, outline the main opportunity, the main ethical risk, and what “good use” of the tool would look like in practice.

Answers