DarazMM: A multilingual multimodal e-commerce product review dataset and benchmark evaluation

Loading...
Thumbnail Image

Publisher

BRAC University

Citation

Abstract

Ordinal sentiment classification from customer reviews using star rating labels is challenging due to subtle distinctions between adjacent rating levels, particularly in low-resource and multilingual settings. Existing benchmarks are predominantly English-centric and text-only, limiting their ability to reflect real-world e-commerce scenarios. We introduce DarazMM, a large-scale Bangla-centric multilingual multimodal dataset for five-class star rating prediction, covering Bangla, English, codemixed Bangla-English and Romanized Bangla. The dataset is released in two settings: a text-only corpus of 257,817 reviews and a multimodal subset of 25,452 reviews that includes review text, customer-uploaded images, advertised product images and corresponding product descriptions. We benchmark zero-shot, few-shot and supervised models, and propose two neuro-symbolic frameworks that model rating ordinality and cross-modal interactions. Our best supervised framework achieves 0.87 accuracy and 0.83 macro-F1, while the best zero-shot and few-shot models reach 0.71 and 0.79 accuracy, respectively. The dataset will be made publicly available.

Description

This thesis is submitted in partial fulfillment of the requirements for the degree of Master of Science in Computer Science and Engineering, 2026.
Cataloged from PDF version of thesis.
Includes bibliographical references (pages 46-50).

Publisher Link

Type

Thesis

Creative Commons license

Attribution-NonCommercial-NoDerivatives 4.0 International

Except where otherwise noted, this item's license is described as

Attribution-NonCommercial-NoDerivatives 4.0 International