PESCO-BERT: An efficient prompt-based contrastive learning for Bangla news classification
Loading...
Date
Publisher
Institute of Electrical and Electronics Engineers Inc.
Authors
Citation
M. Arman, A. Islam, M. M. Hoque and M. M. Rahman, "PESCO-BERT: An Efficient Prompt-Based Contrastive Learning for Bangla News Classification," 2025 28th International Conference on Computer and Information Technology (ICCIT), Cox's Bazar, Bangladesh, 2025, pp. 1463-1468, doi: 10.1109/ICCIT68739.2025.11491368.
Abstract
In this work, propose a scalable method for multiclass Bangla news categorization that combines a Bangla-specific data curation pipeline with contrastive, prompt-based fine-tuning of BanglaBERT. Through Unicode and label normalization, punctuation and digit harmonization, source and time-aware splits using shingled-n-gram MinHash, and light minority oversampling, the pipeline hops over label noise, orthographic variation, class imbalance, and data leakage, respectively. Modeling layer. PESCO (Prompt Ensemble Self-Contrastive) and therefore each article is represented as two semantically relevant but stylistically different prompts and trained using a combined loss of weighted cross-entropy and supervised contrastive loss per-class weighting ?=0.5. To make training practical on commodity hardware, employ QLoRA (4-bit NF4 with safe fallbacks), LoRA adapters on attention matrices, gradient checkpointing, mixed precision and conservative micro-batching. BanglaBERT-PESCO achieves 98.89% accuracy and time-aware splits of the Bangla Newspaper Dataset, outperforming other models including BanglaBERT (base and large), XLM-R (base and large), M-BERT, and a QLoRA-tuned LLaMA-3.
Description
Publisher Link
Type
Conference Proceeding