BanglaDocAtlas: A multi-class annotated dataset for complex Bangla document layout analysis

Citation

M. S. Hossain et al., "BanglaDocAtlas: A Multi-Class Annotated Dataset for Complex Bangla Document Layout Analysis," 2025 IEEE High Performance Extreme Computing Conference (HPEC), Wakefield, MA, USA, 2025, pp. 1-7, doi: 10.1109/HPEC67600.2025.11196300.

Abstract

Optical Character Recognition (OCR) technology is a vital tool for digitizing printed content, enabling efficient data extraction and enhancing document accessibility. Traditional OCR techniques rely on pre-stored templates for fonts or structured documents. Recent advancements in Machine Learning (ML), particularly Convolutional Neural Network (CNN) and transformer-based architectures, have enhanced OCR technologies with human-like intelligence. However, these models often fall short due to limitations in the diversity of document types, layouts, and content in the training datasets, particularly for complex Bangla documents. In this paper, we address the challenge of a limited, diverse dataset by introducing BanglaDocAtlas, a versatile and multi-class annotated dataset specifically designed to advance Bangla document layout analysis. The dataset includes eight distinct classes: paragraph, text, image, title, caption, table, advertisement, and page number, enabling comprehensive OCR applications. State-of-the-art segmentation models, i.e., You Only Look Once (YOLO), and a detection model, e.g., Real-Time DEtection TRansformer (RT-DETR), are trained and evaluated on the BanglaDocAtlas dataset. The results demonstrate that YOLOv9 achieves the highest precision, with values of 0.87 for bounding boxes and 0.79 for masks, while RT-DETR outperforms in recall, with a value of 0.86 for bounding boxes.

Description

Type

Conference Proceeding