Peer-reviewed  ·  BMC Health Services Research

Ambient documentation, evaluated at scale in a Gold Coast hospital

What a large-scale, independently conducted evaluation of ambient documentation found, and what it means for teams implementing these tools.

Published March 2026Reviewed by the Lyrebird Clinical & Research team
01   The study

Measured in routine practice, across nineteen specialties

Clinical documentation has become one of the most significant threats to sustainable healthcare delivery. Clinicians spend a large share of the working day on documentation rather than direct patient care, contributing to burnout and reshaping the therapeutic encounter.

Ambient AI documentation is one proposed response, and it is moving from pilots into routine care. Gold Coast Hospital and Health Service, one of Australia's largest public health services, published peer-reviewed findings from a 16-week evaluation of ambient documentation across 19 outpatient specialties and 7,499 consultations.

This article summarises those findings and extracts practical learnings for teams considering or implementing these tools.

Setting
Tertiary public hospital outpatient clinics at GCHHS, across 19 specialties including paediatrics, orthopaedics, cardiology and mental health.
Scale
100 clinicians · 7,499 consultations · 16 weeks (Jul to Dec 2024).
What was assessed
Tool performance (quality, utility, reliability) and impact on clinician and patient experience in routine practice.
Methods
Mixed methods: staff and patient surveys, interviews, scribe outputs and EMR review, using the validated PDQI-9 and ROUGE.
02   Key outcomes

Benefits across four domains, with reliability considerations

Efficiency and workflow

  • 84% of clinicians reported a positive impact on efficiency
  • 58% of ambient-generated content was accepted without modification
  • Interviews described relief from the "dread" of post-clinic documentation

Note quality

  • Ambient notes scored 37.06/40 on the PDQI-9 versus 34.56/40 for clinician-written notes (blinded, matched pairs)
  • 4.83/5 for freedom from hallucination; 5.0/5 for freedom from bias
  • Rated more thorough, better organised and more useful

Patient experience

  • 68% of patients reported their clinician spent more time speaking directly with them
  • 59% reported the technology had a positive effect on their visit
  • Clinicians described more therapeutic conversations and better eye contact

Reliability and safety

  • 47% of clinicians reported observing a hallucination at least once (self-reported, 43% response rate)
  • 16% observed potential bias in outputs
  • Signals that quality monitoring and safe review matter in implementation

Collectively, the results suggest benefits in clinician experience and workflow, note quality and patient experience, alongside implementation considerations that require attention.

Note quality, measured on a validated instrument
37.06 ambient  vs  34.56 clinician

On the PDQI-9, a validated measure of clinical note quality, ambient-generated notes scored higher than clinician-written notes in a blinded assessment of matched pairs, while being produced faster and with less effort.

03   Implementation lessons

Six lessons for putting ambient documentation to work safely

These reflect Lyrebird's interpretation of the GCHHS findings, informed by implementation experience.

1

Interpret impact relative to baseline documentation quality

The right evaluation question is comparative: does the tool improve note quality and reduce effort compared to your baseline, with acceptable risks and appropriate safeguards? In this evaluation, ambient notes scored higher on the PDQI-9 than standard clinician notes (37.06 vs 34.56) while being produced faster.

"Perfect notes" is not the only benchmark. The more useful question is whether a tool improves quality and reduces burden relative to what actually happens today, under real time pressure.

What good looks like

Teams make the baseline explicit, including the pressures shaping it, before judging the tool against it.

2

Design for safe review, not zero edits

Ambient documentation works best as a draft that reduces effort while keeping clinical judgement firmly in the loop. In the evaluation, 58% of outputs were accepted without modification on average; in other cases clinicians amended before finalising. This mirrors how documentation is safely produced today.

What good looks like

Key facts are easy to verify and correct, especially medications, numbers, diagnoses, laterality, allergies and safety-critical negatives. Safe review feels built in, not bolted on.

3

Build a quality assurance loop to make errors hard to miss

Two findings look hard to reconcile at first. In surveys, 47% of respondents reported observing a hallucination at least once during the trial. In a separate structured assessment of matched note pairs, outputs were rated highly on freedom from hallucination (4.83/5). They measure different things: whether a clinician ever noticed something concerning across many consultations, versus the rated quality of a reviewed sample.

Structured assessment · rated qualityFrontline survey · ever noticed
Structured assessment
4.83/5

Freedom from hallucination, on a reviewed sample of matched note pairs.

Frontline signal
47%

Of clinicians noticed at least one concerning instance across the trial. A frontline safety signal worth capturing.

What good looks like

Clinicians can flag a concern in seconds, and a structured process triages reports, investigates patterns and closes the loop back to the people who raised them.

4

Create shared norms for consistent, safe use

What ends up in the record is shaped by the tool and by the everyday habits around it. The authors raise a longer-term risk: as trust grows, complacency may increase and inaccuracies could be perpetuated. Making safe-use defaults explicit means safety does not rely on each person reinventing a checking approach.

What good looks like

Consent before recording, the whole note reviewed before sign-off, any inconsistency triggering a pause, and issues flagged immediately through a standard pathway.

5

Value is contextual: expect different starting points

The evaluation showed meaningful variation between specialties. The authors highlight orthopaedics, where participation was lower, likely reflecting notes written by junior doctors with a preference for brevity. The more clinicians invested in template customisation and familiarisation, the more they got out of the tool.

What good looks like

Mixed early experiences are treated as feedback, not failure: where it saves time, where it improves notes, and where it needs adjustment to match how that clinic documents.

6

Look beyond time saved: measure the patient experience

It is easy to focus on efficiency, but the findings suggest the impact shows up in the room. 68% of patients said their clinician spent more time speaking directly with them, and 59% felt the technology had a positive effect on their visit. Clinicians described more direct conversation and better eye contact, including during sensitive discussions.

What good looks like

Implementation tracks clinician attention, patient rapport and health literacy as clinical outcomes, not just minutes saved.

04   Evaluating vendors

How to look past the demo to day-to-day clinic reality

Ambient documentation is a fast-moving category. These questions focus on how a tool performs in routine practice, and how a vendor manages quality over time.

  • Quality definitions and detection

    How do you define and classify quality issues (capture problems, mis-structuring, hallucinations, bias)? How are they detected automatically and via user reporting? Can a clinician flag an issue in seconds?

  • Review workflow

    Does the interface make it easy to verify key details (medications, numbers, diagnoses, procedures)? Are edits obvious, trackable and fast, or easy to miss?

  • Monitoring and governance

    Can you track quality trends by specialty, template and model version? How do you test updates before release, and what happens when an update makes something worse?

  • Feedback loops

    What happens after a clinician flags an issue? How quickly do you respond, and what reporting do customers receive?

  • Transparency and partnership

    Will you share quality metrics and respond to independent evaluation? What does support look like after go-live?

05   Open questions

What this evaluation adds, and what it leaves open

This is early evidence from a single health service over 16 weeks. It contributes real-world data on impact, usability and the reliability issues that show up in routine outpatient workflows.

?

Do the benefits persist over longer time horizons, and does complacency develop?

?

How do outcomes compare across vendors and settings?

?

What implementation strategies work best by specialty?

?

How do we measure and mitigate bias reliably?

06   Since the evaluation

Where Lyrebird stands

Lyrebird is a clinician-led company, and our clinical leadership is involved at every level of the business. The trial covered July to December 2024, and the platform has evolved significantly since.

We understand that human judgement remains the gold standard for assessing the quality of a clinical note. Our framework for clinical note quality evaluation reflects that, combining blinded head-to-head comparison with internal clinicians, structured categorisation of issues such as hallucinations, and automated evaluation tools.

This analysis was prepared by the clinical and research leadership team at Lyrebird Health, who are committed to objective interpretation of research findings and transparent discussion of both benefits and limitations.

Purpose-built clinical AI

Clinical AI, held to a clinical standard.

Lyrebird is the clinical AI platform for Australian clinicians. Ambient scribing is one core feature: it documents the consult and the work around it, to a standard you can measure.