Clinical note generation from doctor-patient conversations using parameter-efficient fine-tuning large language models: Comparative study

bracu.type.groupResearch Publications
datacite.rightsOpen Access
dc.contributor.authorAhmed, Saib
dc.contributor.authorSadeque, Farig Yousuf
dc.contributor.departmentDepartment of Computer Science and Engineering
dc.date.accessioned2026-09-16T10:57:28Z
dc.date.available2026-09-16T10:57:28Z
dc.date.issued2026-01-01
dc.description.abstractBackground: Clinical note documentation is a vital yet time-intensive task in health care. While advancements in natural language processing have transformed many domains, generating accurate summaries of doctor-patient conversations remains underexplored due to the limited availability of open-source datasets. Large language models (LLMs), with their training on vast datasets, present a promising solution to this challenge. Objective: Precision in clinical summarization is crucial, as it directly impacts patient care and safety. This study aimed to evaluate the effectiveness of parameter-efficient, fine-tuned, decoder-only LLMs for clinical note generation from doctor-patient conversations. We focus on assessing medical accuracy, robustness, and the feasibility of parameter-efficient fine-tuning (PEFT) approaches under practical resource constraints. Methods: We used the Medical Training Summarization Dialog dataset containing 1700 doctor-patient conversations paired with clinical notes. Several decoder-only LLMs, including Mistral, Meditron, and Llama, were fine-tuned using PEFT techniques to reduce computational and memory overhead. Evaluation was performed using standard automatic metrics, including the Recall-Oriented Understudy for Gisting Evaluation score and bidirectional encoder representations from transformers score, to assess content overlap and semantic similarity between generated and reference clinical notes. In addition, an expert physician assessed the LLM-generated notes for medical accuracy, completeness, concision, relevance, and clinical coherence and readability. Results: Model performance was evaluated using the Recall-Oriented Understudy for Gisting Evaluation score and bidirectional encoder representations from transformers scores, demonstrating that Meditron-7B and Llama3-8B achieved state-of-the-art results among open-source, parameter-efficient, fine-tuned models, with Mistral-7B also performing competitively. The findings indicate that decoder-only LLMs, particularly Llama variants, outperform traditional models. Moreover, fine-tuning with higher quantization has the potential to further enhance performance. Human expert evaluation further indicated that Llama3-8B and Mistral-7B produced clinically coherent and accurate summaries, with Meditron-7B and Llama3-3B also performing reliably across evaluation criteria. The findings suggest that higher quantization during fine-tuning may improve efficiency without substantially compromising performance. Conclusions: This study underscores the potential of the PEFT of decoder-only LLMs to transform clinical workflows by streamlining medical documentation, thereby enabling health care professionals to dedicate more time to patient care. These models offer a scalable and resource-efficient alternative to traditional architectures and have the potential to streamline clinical documentation workflows.
dc.description.versionPublished
dc.format.extent11 pages
dc.identifier.citationAhmed S, Yousuf Sadeque F Clinical Note Generation From Doctor-Patient Conversations Using Parameter-Efficient Fine-Tuning Large Language Models: Comparative Study JMIR Med Inform 2026;14:e82545 URL: https://medinform.jmir.org/2026/1/e82545 DOI: 10.2196/82545
dc.identifier.doi10.2196/82545
dc.identifier.issn22919694
dc.identifier.other2-s2.0-105042247563
dc.identifier.urihttps://hdl.handle.net/10361/30012
dc.language.isoen_US
dc.publisherJMIR Publications Inc.
dc.relation.hasversion10.2196/82545
dc.relation.ispartofJmir Medical Informatics
dc.relation.ispartofseriesJmir Medical Informatics
dc.relation.journalJMIR Medical Informatics
dc.relation.urihttps://medinform.jmir.org/2026/1/e82545
dc.subjectBERTScore
dc.subjectBidirectional encoder representations
dc.subjectClinical natural language processing
dc.subjectClinical NLP
dc.subjectDecoder-only
dc.subjectDialogue2Note
dc.subjectLlama
dc.subjectMeditron
dc.subjectMistral
dc.subjectRecall-Oriented understudy for gisting evaluation
dc.subjectROUGE score
dc.subjectTransformer
dc.subject.lcshMedical records--Data processing.
dc.subject.lcshNatural language processing (Computer science).
dc.subject.lcshMedical informatics.
dc.subject.lcshPhysician and patient.
dc.titleClinical note generation from doctor-patient conversations using parameter-efficient fine-tuning large language models: Comparative study
dc.typeArticle
oaire.citation.volume14
person.affiliation.nameBRAC University
person.affiliation.nameBRAC University
person.identifier.scopus-author-id60697713300
person.identifier.scopus-author-id55843529500

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Clinical Note Generation From Doctor-Patient Conversations Using Parameter-Efficient Fine-Tuning Large Language Models Comparative Study.pdf
Size:
251.31 KB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: