Clinical note generation from doctor-patient conversations using parameter-efficient fine-tuning large language models: Comparative study
| bracu.type.group | Research Publications | |
| datacite.rights | Open Access | |
| dc.contributor.author | Ahmed, Saib | |
| dc.contributor.author | Sadeque, Farig Yousuf | |
| dc.contributor.department | Department of Computer Science and Engineering | |
| dc.date.accessioned | 2026-09-16T10:57:28Z | |
| dc.date.available | 2026-09-16T10:57:28Z | |
| dc.date.issued | 2026-01-01 | |
| dc.description.abstract | Background: Clinical note documentation is a vital yet time-intensive task in health care. While advancements in natural language processing have transformed many domains, generating accurate summaries of doctor-patient conversations remains underexplored due to the limited availability of open-source datasets. Large language models (LLMs), with their training on vast datasets, present a promising solution to this challenge. Objective: Precision in clinical summarization is crucial, as it directly impacts patient care and safety. This study aimed to evaluate the effectiveness of parameter-efficient, fine-tuned, decoder-only LLMs for clinical note generation from doctor-patient conversations. We focus on assessing medical accuracy, robustness, and the feasibility of parameter-efficient fine-tuning (PEFT) approaches under practical resource constraints. Methods: We used the Medical Training Summarization Dialog dataset containing 1700 doctor-patient conversations paired with clinical notes. Several decoder-only LLMs, including Mistral, Meditron, and Llama, were fine-tuned using PEFT techniques to reduce computational and memory overhead. Evaluation was performed using standard automatic metrics, including the Recall-Oriented Understudy for Gisting Evaluation score and bidirectional encoder representations from transformers score, to assess content overlap and semantic similarity between generated and reference clinical notes. In addition, an expert physician assessed the LLM-generated notes for medical accuracy, completeness, concision, relevance, and clinical coherence and readability. Results: Model performance was evaluated using the Recall-Oriented Understudy for Gisting Evaluation score and bidirectional encoder representations from transformers scores, demonstrating that Meditron-7B and Llama3-8B achieved state-of-the-art results among open-source, parameter-efficient, fine-tuned models, with Mistral-7B also performing competitively. The findings indicate that decoder-only LLMs, particularly Llama variants, outperform traditional models. Moreover, fine-tuning with higher quantization has the potential to further enhance performance. Human expert evaluation further indicated that Llama3-8B and Mistral-7B produced clinically coherent and accurate summaries, with Meditron-7B and Llama3-3B also performing reliably across evaluation criteria. The findings suggest that higher quantization during fine-tuning may improve efficiency without substantially compromising performance. Conclusions: This study underscores the potential of the PEFT of decoder-only LLMs to transform clinical workflows by streamlining medical documentation, thereby enabling health care professionals to dedicate more time to patient care. These models offer a scalable and resource-efficient alternative to traditional architectures and have the potential to streamline clinical documentation workflows. | |
| dc.description.version | Published | |
| dc.format.extent | 11 pages | |
| dc.identifier.citation | Ahmed S, Yousuf Sadeque F Clinical Note Generation From Doctor-Patient Conversations Using Parameter-Efficient Fine-Tuning Large Language Models: Comparative Study JMIR Med Inform 2026;14:e82545 URL: https://medinform.jmir.org/2026/1/e82545 DOI: 10.2196/82545 | |
| dc.identifier.doi | 10.2196/82545 | |
| dc.identifier.issn | 22919694 | |
| dc.identifier.other | 2-s2.0-105042247563 | |
| dc.identifier.uri | https://hdl.handle.net/10361/30012 | |
| dc.language.iso | en_US | |
| dc.publisher | JMIR Publications Inc. | |
| dc.relation.hasversion | 10.2196/82545 | |
| dc.relation.ispartof | Jmir Medical Informatics | |
| dc.relation.ispartofseries | Jmir Medical Informatics | |
| dc.relation.journal | JMIR Medical Informatics | |
| dc.relation.uri | https://medinform.jmir.org/2026/1/e82545 | |
| dc.subject | BERTScore | |
| dc.subject | Bidirectional encoder representations | |
| dc.subject | Clinical natural language processing | |
| dc.subject | Clinical NLP | |
| dc.subject | Decoder-only | |
| dc.subject | Dialogue2Note | |
| dc.subject | Llama | |
| dc.subject | Meditron | |
| dc.subject | Mistral | |
| dc.subject | Recall-Oriented understudy for gisting evaluation | |
| dc.subject | ROUGE score | |
| dc.subject | Transformer | |
| dc.subject.lcsh | Medical records--Data processing. | |
| dc.subject.lcsh | Natural language processing (Computer science). | |
| dc.subject.lcsh | Medical informatics. | |
| dc.subject.lcsh | Physician and patient. | |
| dc.title | Clinical note generation from doctor-patient conversations using parameter-efficient fine-tuning large language models: Comparative study | |
| dc.type | Article | |
| oaire.citation.volume | 14 | |
| person.affiliation.name | BRAC University | |
| person.affiliation.name | BRAC University | |
| person.identifier.scopus-author-id | 60697713300 | |
| person.identifier.scopus-author-id | 55843529500 |
Files
Original bundle
1 - 1 of 1
Loading...
- Name:
- Clinical Note Generation From Doctor-Patient Conversations Using Parameter-Efficient Fine-Tuning Large Language Models Comparative Study.pdf
- Size:
- 251.31 KB
- Format:
- Adobe Portable Document Format
License bundle
1 - 1 of 1
Loading...
- Name:
- license.txt
- Size:
- 1.71 KB
- Format:
- Item-specific license agreed upon to submission
- Description: