Welcome to the upgraded BRAC University Institutional Repository. We are currently organizing collections after a recent system upgrade. Homepage category counters may temporarily show lower numbers while syncing, but over 27,000 repository items remain safe and accessible. Please use the search bar to find theses, scholarly outputs, and institutional documents.

Learning to win: A process reward model for competitive machine learning agents

bracu.degree.levelPostgraduate
bracu.type.groupStudent Works
datacite.rightsOpen Access
dc.contributor.advisorAlam, Md. Golam Rabiul
dc.contributor.authorAhmed, Sajjad
dc.contributor.departmentDepartment of Computer Science and Engineering
dc.date.accessioned2026-07-26T06:09:38Z
dc.date.available2026-07-26T06:09:38Z
dc.date.copyright2026
dc.date.issued2026-03
dc.descriptionThis thesis is submitted in partial fulfillment of the requirements for the degree of Master of Science in Computer Science and Engineering, 2026.
dc.descriptionCataloged from PDF version of thesis.
dc.descriptionIncludes bibliographical references (pages 44-46).
dc.description.abstractLarge Language Model (LLM)-based coding agents are increasingly deployed for multi-step data science tasks, yet no systematic study has examined how these agents behave across diverse competitive problems. This thesis presents a two-part investigation using 194 human-supervised agent human feedback JSONs across 103 Kaggle competitions (1,584 turns total). First, we conduct a comprehensive empirical study. We also propose and show that process-oriented quality metrics. Second, we develop a Process Reward Model (PRM) for data science workflows—the first application of PRMs beyond mathematical reasoning—that predicts whether an intermediate step will improve competition scores. Using 49 tabular features across 6 feature groups, we train XGBoost and LightGBM baselines and conduct a full feature ablation study. Code features (code length, imports, deltas) emerge as the strongest individual tabular predictors. We further fine-tune Qwen2.5-Coder-3B with QLoRA as a hybrid language-model-based PRM (v2), incorporating code snippets and tabular features into semantic prompts, which achieves AUROC 0.656—surpassing the best tabular baseline (0.530) by 12.6 points—by leveraging semantic content from agent plans, reflections, and code. To establish the necessity of domain-specific finetuning, we evaluate zero-shot frontier models (i.e., GPT-5, Gemini 3.0 Flash/Pro, Llama 3.3 70B) on the same task. GPT-4o achieves the highest zero-shot AUROC (0.709), outperforming our fine-tuned 3B model, while most other frontier models fall to majority class prediction (AUROC 0.500). This demonstrates that while powerful zero-shot models can reason about the task, domain-specific fine-tuning remains valuable for smaller models, achieving competitive performance with 200x fewer parameters. Ultimately, our research demonstrates that observable code signals are significantly more reliable than an agent’s articulated plans for predicting success, providing a robust foundation for the development of real-time guidance systems in competitive data science.
dc.description.degreeMaster of Science in Computer Science and Engineering
dc.description.statementofresponsibilitySajjad Ahmed
dc.format.extent57 pages
dc.identifier.otherID 20366018
dc.identifier.urihttps://hdl.handle.net/10361/28632
dc.language.isoen_US
dc.publisherBRAC University
dc.rightsAttribution-NonCommercial-NoDerivatives 4.0 International
dc.rightsBRAC University theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission.
dc.rights.urihttp://creativecommons.org/licenses/by-nc-nd/4.0/
dc.subjectLLM
dc.subjectLarge language models
dc.subjectData science
dc.subjectProcess reward model
dc.subjectReinforcement learning from human feedback
dc.subjectRLHF
dc.subject.lcshMachine learning.
dc.subject.lcshReinforcement learning.
dc.subject.lcshNatural language processing (Computer science).
dc.titleLearning to win: A process reward model for competitive machine learning agents
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
20366018_CSE.pdf
Size:
519.84 KB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: