A self-adaptive data preprocessing pipeline for machine learning: an automated and dynamic approach
Loading...
Files
Date
Publisher
Institute of Electrical and Electronics Engineers Inc.
Citation
D. B. Niloy, M. E. Hossen Jony and H. K. Ruhani, "A Self-Adaptive Data Preprocessing Pipeline for Machine Learning: An Automated and Dynamic Approach," 2025 International Conference on Quantum Photonics, Artificial Intelligence, and Networking (QPAIN), Rangpur, Bangladesh, 2025, pp. 1-6, doi: 10.1109/QPAIN66474.2025.11171989.
Abstract
The optimization and generalization of performance of a machine learning model is profoundly influenced by efficient data preprocessing. A machine's learning model does not perform to its expectable capability because the traditional automation processes which are manual rule-triggered are not data diverse friendly and rely heavily on human input. In this paper, we propose the Self-Adaptive Data Preprocessing Pipeline which modifies its preprocessing phases according to specific dataset details. Automated feature selection, data imputation, normalization, and noise reduction are performed using a blend of machine learning heuristics, statistics, and reinforcement learning. Through iterative assessment of various preprocessing methods, SADPP enhances pipeline processes in real time - streamlining human input, enhancing adaptive functions, and increasing reliability. The results indicate that SADPP achieves comparable accuracy while minimizing human tuning time by 85 % and computational resources by 22 %. This adaptive framework offers a significant stride towards sophisticated, automated, scalable data preprocessing systems in machine learning optimized for the growth of AI technologies.
LC Subject Headings
Description
Publisher Link
Type
Conference Proceeding