PESCO-BERT: An efficient prompt-based contrastive learning for Bangla news classification

Citation

M. Arman, A. Islam, M. M. Hoque and M. M. Rahman, "PESCO-BERT: An Efficient Prompt-Based Contrastive Learning for Bangla News Classification," 2025 28th International Conference on Computer and Information Technology (ICCIT), Cox's Bazar, Bangladesh, 2025, pp. 1463-1468, doi: 10.1109/ICCIT68739.2025.11491368.

Abstract

In this work, propose a scalable method for multiclass Bangla news categorization that combines a Bangla-specific data curation pipeline with contrastive, prompt-based fine-tuning of BanglaBERT. Through Unicode and label normalization, punctuation and digit harmonization, source and time-aware splits using shingled-n-gram MinHash, and light minority oversampling, the pipeline hops over label noise, orthographic variation, class imbalance, and data leakage, respectively. Modeling layer. PESCO (Prompt Ensemble Self-Contrastive) and therefore each article is represented as two semantically relevant but stylistically different prompts and trained using a combined loss of weighted cross-entropy and supervised contrastive loss per-class weighting ?=0.5. To make training practical on commodity hardware, employ QLoRA (4-bit NF4 with safe fallbacks), LoRA adapters on attention matrices, gradient checkpointing, mixed precision and conservative micro-batching. BanglaBERT-PESCO achieves 98.89% accuracy and time-aware splits of the Bangla Newspaper Dataset, outperforming other models including BanglaBERT (base and large), XLM-R (base and large), M-BERT, and a QLoRA-tuned LLaMA-3.

Description

Type

Conference Proceeding