What a large-scale, independently conducted evaluation of ambient documentation found, and what it means for teams implementing these tools.
Clinical documentation has become one of the most significant threats to sustainable healthcare delivery. Clinicians spend a large share of the working day on documentation rather than direct patient care, contributing to burnout and reshaping the therapeutic encounter.
Ambient AI documentation is one proposed response, and it is moving from pilots into routine care. Gold Coast Hospital and Health Service, one of Australia's largest public health services, published peer-reviewed findings from a 16-week evaluation of ambient documentation across 19 outpatient specialties and 7,499 consultations.
This article summarises those findings and extracts practical learnings for teams considering or implementing these tools.
Collectively, the results suggest benefits in clinician experience and workflow, note quality and patient experience, alongside implementation considerations that require attention.
On the PDQI-9, a validated measure of clinical note quality, ambient-generated notes scored higher than clinician-written notes in a blinded assessment of matched pairs, while being produced faster and with less effort.
These reflect Lyrebird's interpretation of the GCHHS findings, informed by implementation experience.
The right evaluation question is comparative: does the tool improve note quality and reduce effort compared to your baseline, with acceptable risks and appropriate safeguards? In this evaluation, ambient notes scored higher on the PDQI-9 than standard clinician notes (37.06 vs 34.56) while being produced faster.
"Perfect notes" is not the only benchmark. The more useful question is whether a tool improves quality and reduces burden relative to what actually happens today, under real time pressure.
Teams make the baseline explicit, including the pressures shaping it, before judging the tool against it.
Ambient documentation works best as a draft that reduces effort while keeping clinical judgement firmly in the loop. In the evaluation, 58% of outputs were accepted without modification on average; in other cases clinicians amended before finalising. This mirrors how documentation is safely produced today.
Key facts are easy to verify and correct, especially medications, numbers, diagnoses, laterality, allergies and safety-critical negatives. Safe review feels built in, not bolted on.
Two findings look hard to reconcile at first. In surveys, 47% of respondents reported observing a hallucination at least once during the trial. In a separate structured assessment of matched note pairs, outputs were rated highly on freedom from hallucination (4.83/5). They measure different things: whether a clinician ever noticed something concerning across many consultations, versus the rated quality of a reviewed sample.
Freedom from hallucination, on a reviewed sample of matched note pairs.
Of clinicians noticed at least one concerning instance across the trial. A frontline safety signal worth capturing.
Clinicians can flag a concern in seconds, and a structured process triages reports, investigates patterns and closes the loop back to the people who raised them.
What ends up in the record is shaped by the tool and by the everyday habits around it. The authors raise a longer-term risk: as trust grows, complacency may increase and inaccuracies could be perpetuated. Making safe-use defaults explicit means safety does not rely on each person reinventing a checking approach.
Consent before recording, the whole note reviewed before sign-off, any inconsistency triggering a pause, and issues flagged immediately through a standard pathway.
The evaluation showed meaningful variation between specialties. The authors highlight orthopaedics, where participation was lower, likely reflecting notes written by junior doctors with a preference for brevity. The more clinicians invested in template customisation and familiarisation, the more they got out of the tool.
Mixed early experiences are treated as feedback, not failure: where it saves time, where it improves notes, and where it needs adjustment to match how that clinic documents.
It is easy to focus on efficiency, but the findings suggest the impact shows up in the room. 68% of patients said their clinician spent more time speaking directly with them, and 59% felt the technology had a positive effect on their visit. Clinicians described more direct conversation and better eye contact, including during sensitive discussions.
Implementation tracks clinician attention, patient rapport and health literacy as clinical outcomes, not just minutes saved.
Ambient documentation is a fast-moving category. These questions focus on how a tool performs in routine practice, and how a vendor manages quality over time.
How do you define and classify quality issues (capture problems, mis-structuring, hallucinations, bias)? How are they detected automatically and via user reporting? Can a clinician flag an issue in seconds?
Does the interface make it easy to verify key details (medications, numbers, diagnoses, procedures)? Are edits obvious, trackable and fast, or easy to miss?
Can you track quality trends by specialty, template and model version? How do you test updates before release, and what happens when an update makes something worse?
What happens after a clinician flags an issue? How quickly do you respond, and what reporting do customers receive?
Will you share quality metrics and respond to independent evaluation? What does support look like after go-live?
This is early evidence from a single health service over 16 weeks. It contributes real-world data on impact, usability and the reliability issues that show up in routine outpatient workflows.
Do the benefits persist over longer time horizons, and does complacency develop?
How do outcomes compare across vendors and settings?
What implementation strategies work best by specialty?
How do we measure and mitigate bias reliably?
Lyrebird is a clinician-led company, and our clinical leadership is involved at every level of the business. The trial covered July to December 2024, and the platform has evolved significantly since.
We understand that human judgement remains the gold standard for assessing the quality of a clinical note. Our framework for clinical note quality evaluation reflects that, combining blinded head-to-head comparison with internal clinicians, structured categorisation of issues such as hallucinations, and automated evaluation tools.
This analysis was prepared by the clinical and research leadership team at Lyrebird Health, who are committed to objective interpretation of research findings and transparent discussion of both benefits and limitations.
Lyrebird is the clinical AI platform for Australian clinicians. Ambient scribing is one core feature: it documents the consult and the work around it, to a standard you can measure.