Welcome to the upgraded BRAC University Institutional Repository. We are currently organizing collections after a recent system upgrade. Homepage category counters may temporarily show lower numbers while syncing, but over 27,000 repository items remain safe and accessible. Please use the search bar to find theses, scholarly outputs, and institutional documents.

Learning to win: A process reward model for competitive machine learning agents

Loading...
Thumbnail Image

Publisher

BRAC University

Citation

Abstract

Large Language Model (LLM)-based coding agents are increasingly deployed for multi-step data science tasks, yet no systematic study has examined how these agents behave across diverse competitive problems. This thesis presents a two-part investigation using 194 human-supervised agent human feedback JSONs across 103 Kaggle competitions (1,584 turns total). First, we conduct a comprehensive empirical study. We also propose and show that process-oriented quality metrics. Second, we develop a Process Reward Model (PRM) for data science workflows—the first application of PRMs beyond mathematical reasoning—that predicts whether an intermediate step will improve competition scores. Using 49 tabular features across 6 feature groups, we train XGBoost and LightGBM baselines and conduct a full feature ablation study. Code features (code length, imports, deltas) emerge as the strongest individual tabular predictors. We further fine-tune Qwen2.5-Coder-3B with QLoRA as a hybrid language-model-based PRM (v2), incorporating code snippets and tabular features into semantic prompts, which achieves AUROC 0.656—surpassing the best tabular baseline (0.530) by 12.6 points—by leveraging semantic content from agent plans, reflections, and code. To establish the necessity of domain-specific finetuning, we evaluate zero-shot frontier models (i.e., GPT-5, Gemini 3.0 Flash/Pro, Llama 3.3 70B) on the same task. GPT-4o achieves the highest zero-shot AUROC (0.709), outperforming our fine-tuned 3B model, while most other frontier models fall to majority class prediction (AUROC 0.500). This demonstrates that while powerful zero-shot models can reason about the task, domain-specific fine-tuning remains valuable for smaller models, achieving competitive performance with 200x fewer parameters. Ultimately, our research demonstrates that observable code signals are significantly more reliable than an agent’s articulated plans for predicting success, providing a robust foundation for the development of real-time guidance systems in competitive data science.

Description

This thesis is submitted in partial fulfillment of the requirements for the degree of Master of Science in Computer Science and Engineering, 2026.
Cataloged from PDF version of thesis.
Includes bibliographical references (pages 44-46).

Publisher Link

Type

Thesis

Creative Commons license

Attribution-NonCommercial-NoDerivatives 4.0 International

Except where otherwise noted, this item's license is described as

Attribution-NonCommercial-NoDerivatives 4.0 International