Knowledge distillation in split learning for heterogeneous clinical tabular data
| bracu.degree.level | Postgraduate | |
| bracu.type.group | Student Works | |
| datacite.rights | Open Access | |
| dc.contributor.advisor | Alam, Md. Golam Rabiul | |
| dc.contributor.author | Farheen, Tasneem Jahan | |
| dc.contributor.department | Department of Computer Science and Engineering | |
| dc.date.accessioned | 2026-07-26T04:17:05Z | |
| dc.date.available | 2026-07-26T04:17:05Z | |
| dc.date.copyright | 2026 | |
| dc.date.issued | 2026-03 | |
| dc.description | This thesis is submitted in partial fulfillment of the requirements for the degree of Master of Science in Computer Science and Engineering, 2026. | |
| dc.description | Cataloged from PDF version of thesis. | |
| dc.description | Includes bibliographical references (pages 63-66). | |
| dc.description.abstract | The growing availability of clinical data across institutions creates new opportunities for improved disease risk prediction. However, privacy regulations, institutional policies, and limited labeled data restrict centralized model training, particularly in low-resource settings. To address these challenges, this study presents a privacypreserving framework that integrates split learning and knowledge distillation for diabetes risk prediction using heterogeneous tabular datasets. This work systematically evaluates two complementary transfer mechanisms. The first approach utilizes logit-level knowledge distillation to transfer predictive decision boundaries from high-capacity teacher models, collaboratively trained on large population datasets, to lightweight student models in low-resource target domains. The second approach implements encoder-level distillation to align latent feature representations within a split learning architecture. Both approaches maintain strict data privacy by keeping raw patient data local and exchanging only intermediate activations. Experimental results under few-shot supervision indicate that logit-level distillation substantially improves discriminative performance, particularly in extremely low-label scenarios while stabilizing training across heterogeneous domains. Encoder-level distillation further improves probabilistic calibration and representation alignment under crossdomain distribution shifts. These findings underscore the significance of structured knowledge transfer in privacy-preserving clinical machine learning and offer practical recommendations for deploying reliable models in heterogeneous healthcare environments. | |
| dc.description.degree | Master of Science in Computer Science and Engineering | |
| dc.description.statementofresponsibility | Tasneem Jahan Farheen | |
| dc.format.extent | 81 pages | |
| dc.identifier.other | ID 24366032 | |
| dc.identifier.uri | https://hdl.handle.net/10361/28619 | |
| dc.language.iso | en_US | |
| dc.publisher | BRAC University | |
| dc.rights | Attribution-NonCommercial-NoDerivatives 4.0 International | |
| dc.rights | BRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. | |
| dc.rights.uri | http://creativecommons.org/licenses/by-nc-nd/4.0/ | |
| dc.subject | Split learning | |
| dc.subject | Knowledge distillation | |
| dc.subject | Clinical prediction | |
| dc.subject | Tabular transformers | |
| dc.subject | Few-shot learning | |
| dc.subject | Clinical data | |
| dc.subject.lcsh | Machine learning. | |
| dc.subject.lcsh | Medical records--Data processing. | |
| dc.title | Knowledge distillation in split learning for heterogeneous clinical tabular data | |
| dc.type | Thesis |