Skip to main navigation Skip to search Skip to main content

Enhancing clinical documentation with voice processing and large language models: a study on the LAOS system

  • Yupeng Xu (Co-first Author)
  • , Huixun Jia (Co-first Author)
  • , Maolin Wang (Co-first Author)
  • , Jie Feng (Co-first Author)
  • , Xun Xu
  • , Haiyan Wang
  • , Jieqiong Chen
  • , Zheng Zheng
  • , Xiaoyan Yang
  • , Yue Shen
  • , Jian Wang
  • , Chenyi Zhuang
  • , Peng Wei
  • , Ruocheng Guo
  • , Xiangyu Zhao
  • , Junxiang Fan*
  • , Xiaodong Sun*
  • *Corresponding author for this work

Research output: Journal Publications and ReviewsRGC 21 - Publication in refereed journalpeer-review

6 Downloads (CityUHK Scholars)

Abstract

The growing volume of Electronic Health Records (EHRs) has enhanced patient care quality but significantly increased the cognitive workload on clinicians, particularly in ophthalmology where specialists handle 1.6 times more patient consultations than other specialties. This study introduces the “LLM-based Auxiliary Ophthalmic System (LAOS),” an integrated framework leveraging Large Language Models (LLMs) and audio processing to improve clinical documentation accuracy and efficiency. LAOS combines voice recognition with Retrieval-Augmented Generation (RAG) and Low-Rank Adaptation (LoRA) to convert clinical conversations into structured documentation while dynamically retrieving relevant medical knowledge. The system was evaluated across three critical documentation tasks: Admission Reports, Surgery Records, and Discharge Summaries. Through both quantitative metrics (BLEU, ROUGE-L, BERT Score) and clinical validation by board-certified physicians, LAOS demonstrated significant improvements in documentation completeness, accuracy, and efficiency. While challenges remain in balancing comprehensiveness with conciseness, this research highlights the potential of speech-enabled LLM systems to alleviate physician burnout, enhance documentation quality, and improve healthcare delivery. © The Author(s) 2025.
Original languageEnglish
Article number798
Number of pages15
Journalnpj Digital Medicine
Volume8
Online published28 Nov 2025
DOIs
Publication statusPublished - 2025

Bibliographical note

Full text of this publication does not contain sufficient affiliation information. With consent from the author(s) concerned, the Research Unit(s) information for this record is based on the existing academic department affiliation of the author(s).

Funding

This study was funded by the National Natural Science Foundation of China (82388101, U22A20311), National Key R&D Program (2022YFC2502800), Shanghai Municipal Education Commission(2023ZKZD18), Science and Technology Commission of Shanghai Municipality(23J41900200), National Clinical Key Specialty Construction Project(10000015Z155080000004). The funder played no role in study design, data collection, analysis and interpretation of data, or the writing of this manuscript.We also want to thank all the first-year ophthalmologists participated in this study.

Publisher's Copyright Statement

  • This full text is made available under CC-BY-NC-ND 4.0. https://creativecommons.org/licenses/by-nc-nd/4.0/

Fingerprint

Dive into the research topics of 'Enhancing clinical documentation with voice processing and large language models: a study on the LAOS system'. Together they form a unique fingerprint.

Cite this