Knowledge distillation in split learning for heterogeneous clinical tabular data

bracu.degree.levelPostgraduate
bracu.type.groupStudent Works
datacite.rightsOpen Access
dc.contributor.advisorAlam, Md. Golam Rabiul
dc.contributor.authorFarheen, Tasneem Jahan
dc.contributor.departmentDepartment of Computer Science and Engineering
dc.date.accessioned2026-07-26T04:17:05Z
dc.date.available2026-07-26T04:17:05Z
dc.date.copyright2026
dc.date.issued2026-03
dc.descriptionThis thesis is submitted in partial fulfillment of the requirements for the degree of Master of Science in Computer Science and Engineering, 2026.
dc.descriptionCataloged from PDF version of thesis.
dc.descriptionIncludes bibliographical references (pages 63-66).
dc.description.abstractThe growing availability of clinical data across institutions creates new opportunities for improved disease risk prediction. However, privacy regulations, institutional policies, and limited labeled data restrict centralized model training, particularly in low-resource settings. To address these challenges, this study presents a privacypreserving framework that integrates split learning and knowledge distillation for diabetes risk prediction using heterogeneous tabular datasets. This work systematically evaluates two complementary transfer mechanisms. The first approach utilizes logit-level knowledge distillation to transfer predictive decision boundaries from high-capacity teacher models, collaboratively trained on large population datasets, to lightweight student models in low-resource target domains. The second approach implements encoder-level distillation to align latent feature representations within a split learning architecture. Both approaches maintain strict data privacy by keeping raw patient data local and exchanging only intermediate activations. Experimental results under few-shot supervision indicate that logit-level distillation substantially improves discriminative performance, particularly in extremely low-label scenarios while stabilizing training across heterogeneous domains. Encoder-level distillation further improves probabilistic calibration and representation alignment under crossdomain distribution shifts. These findings underscore the significance of structured knowledge transfer in privacy-preserving clinical machine learning and offer practical recommendations for deploying reliable models in heterogeneous healthcare environments.
dc.description.degreeMaster of Science in Computer Science and Engineering
dc.description.statementofresponsibilityTasneem Jahan Farheen
dc.format.extent81 pages
dc.identifier.otherID 24366032
dc.identifier.urihttps://hdl.handle.net/10361/28619
dc.language.isoen_US
dc.publisherBRAC University
dc.rightsAttribution-NonCommercial-NoDerivatives 4.0 International
dc.rightsBRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission.
dc.rights.urihttp://creativecommons.org/licenses/by-nc-nd/4.0/
dc.subjectSplit learning
dc.subjectKnowledge distillation
dc.subjectClinical prediction
dc.subjectTabular transformers
dc.subjectFew-shot learning
dc.subjectClinical data
dc.subject.lcshMachine learning.
dc.subject.lcshMedical records--Data processing.
dc.titleKnowledge distillation in split learning for heterogeneous clinical tabular data
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
24366032_CSE.pdf
Size:
977.61 KB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: