Automated chest X-Ray report generation using vision-language models: a Llama 3.2 based approch

bracu.type.groupResearch Publications
datacite.rightsMetadata Only
dc.contributor.authorArman, Mithila
dc.contributor.authorShovon, Reduanul Bari
dc.contributor.authorEmon, Sonet Barua
dc.contributor.authorSameha, Jannatul Asma
dc.contributor.authorFaisal, Fahad Siddique
dc.contributor.authorSayem, Moshiur Rahman
dc.contributor.departmentDepartment of Computer Science and Engineering
dc.date.accessioned2026-08-11T05:55:25Z
dc.date.available2026-08-11T05:55:25Z
dc.date.issued2026-01-01
dc.description.abstractAutomating radiology report generation from chest X-ray images offers the potential to ease the workload of radiologists while also improving diagnostic accuracy and efficiency. Recent advancements in vision-language models have shown strong promise in aligning visual data with natural language descriptions. In this study, we introduce a framework for automated chest X-ray report generation that leverages a LLaMA 3.2 vision-language architecture. Our approach combines image classification features with multimodal language modeling using a teacher-student learning paradigm. This integration allows the model to generate radiology reports that are both clinically accurate and contextually meaningful. By guiding the student model through teacher supervision, the framework enhances the coherence and relevance of the generated text in clinical settings. We evaluate the proposed method using a widely adopted chest X-ray dataset. The experimental results show that our model achieves strong performance, recording a BLEU score of 0.503 and a METEOR score of 0.657. These results surpass those of baseline methods in terms of both accuracy and fluency, demonstrating the effectiveness of the approach in producing high-quality medical reports. The findings suggest that large language models, when carefully adapted for multimodal medical data, can generate radiology reports that closely resemble those written by experts. This highlights the practical potential of our framework to support real-world clinical decision-making and improve healthcare delivery.
dc.description.versionPublished
dc.format.extent6 pages
dc.identifier.citationM. Arman, R. B. Shovon, S. B. Emon, J. A. Sameha, F. S. Faisal and M. R. Sayem, "Automated Chest X-Ray Report Generation Using Vision-Language Models: a Llama 3.2 Based Approch," 2026 IEEE 2nd International Conference on Quantum Photonics, Artificial Intelligence & Networking (QPAIN), Chittagong, Bangladesh, 2026, pp. 1-6, doi: 10.1109/QPAIN69676.2026.11545770.
dc.identifier.doi10.1109/QPAIN69676.2026.11545770
dc.identifier.issn9798331549909
dc.identifier.other2-s2.0-105042899571
dc.identifier.urihttps://hdl.handle.net/10361/28915
dc.language.isoen_US
dc.publisherInstitute of Electrical and Electronics Engineers Inc.
dc.relation.hasversion10.1109/QPAIN69676.2026.11545770
dc.relation.ispartof2026 IEEE 2nd International Conference on Quantum Photonics Artificial Intelligence and Networking Qpain 2026
dc.relation.ispartofseries2026 IEEE 2nd International Conference on Quantum Photonics Artificial Intelligence and Networking Qpain 2026
dc.relation.urihttps://ieeexplore.ieee.org/document/11545770
dc.rightsfalse
dc.subjectClassification
dc.subjectEfficientNetB7
dc.subjectLLaMa
dc.subjectLLMs
dc.subjectX-ray reports
dc.subject.lcshX-rays.
dc.subject.lcshComputational intelligence.
dc.subject.lcshComputer simulation.
dc.titleAutomated chest X-Ray report generation using vision-language models: a Llama 3.2 based approch
dc.typeConference Proceeding
person.affiliation.nameBRAC University
person.affiliation.nameUniversity of Scholars
person.affiliation.nameNoakhali Science and Technology University
person.affiliation.nameHazera Taju Degree College
person.affiliation.nameChittagong University of Engineering and Technology
person.affiliation.nameUniversity of Science and Technology Chittagong
person.identifier.scopus-author-id58144027900
person.identifier.scopus-author-id59390387100
person.identifier.scopus-author-id59464106800
person.identifier.scopus-author-id60709291700
person.identifier.scopus-author-id57211188719
person.identifier.scopus-author-id60709687700

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Demo.jpg
Size:
27.28 KB
Format:
Joint Photographic Experts Group/JPEG File Interchange Format (JFIF)

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: