Welcome to the upgraded BRAC University Institutional Repository. We are currently organizing collections after a recent system upgrade. Homepage category counters may temporarily show lower numbers while syncing, but over 27,000 repository items remain safe and accessible. Please use the search bar to find theses, scholarly outputs, and institutional documents.

Chained semantic retrieval for rare disease and gene identification using clinical phenotype

Loading...
Thumbnail Image

Publisher

BRAC University

Citation

Abstract

Accurate identification of rare diseases and associated genes from patient phenotypic data represents a critical challenge in precision medicine and genomic research. Current diagnostic approaches face substantial limitations in scalability, interpretability, and real-time processing when integrating heterogeneous phenotypic and genotypic databases. This study presents a novel artificial intelligence-driven system employing chained semantic retrieval and knowledge graph integration to identify rare diseases, genes, and Human Phenotype Ontology (HPO) terms from patient phenotype descriptions. The methodology leverages sentence transformers (based on Bidirectional and Auto-Regressive Transformers) for generating vector embeddings, Facebook AI Similarity Search (FAISS) for efficient similarity computation, and a fine-tuned Llama 3.2 model integrated with a Retrieval-Augmented Generation (RAG) pipeline. The system implements a tiered scoring mechanism that chains retrieval across Human Phenotype Ontology, disease, and gene databases to progressively refine predictions through contextual enhancement. Evaluation on 50 patient phenotypes with confirmed Duchenne Muscular Dystrophy diagnosis, consisting of 30 true positive and 20 true negative cases, demonstrated strong performance with 28 true positives, 17 true negatives, 2 false negatives, and 3 false positives in top-ten retrieval results. The system achieved 90 % overall accuracy, 93.3 % recall, 85 % specificity, 90.3 % precision, and an F1-score of 91.8 %, with an average computational efficiency of 3.2 seconds per response. The proposed framework effectively addresses critical gaps in cross-database integration while maintaining interpretability through tiered confidence scoring for clinical decision support applications.

Description

Cataloged from PDF version of thesis.
Includes bibliographical references (pages 57-58).
This thesis is submitted in partial fulfillment of the requirements for the degree of Master of Science in Computer Science, 2025.

Publisher Link

Type

Thesis