Transformation of visual information into Bangla textual representation

bracu.type.groupResearch Publications
datacite.rightsMetadata Only
dc.contributor.authorNawer, Nafisa
dc.contributor.authorKhan, Md. Shakiful Islam
dc.contributor.authorAlam, Md. Mustakin
dc.contributor.authorMehedi, Md Humaion Kabir
dc.contributor.authorRasel, Annajiat Alim
dc.contributor.departmentDepartment of Computer Science and Engineering
dc.date.accessioned2026-07-26T11:13:13Z
dc.date.available2026-07-26T11:13:13Z
dc.date.issued2023-01-01
dc.description.abstractIn the past several years, the interest in research like generating humanoid descriptions of scenarios by detecting and analyzing their components has been increased tremendously. Even though a significant amount of research has been put into automating the process of converting visual information into written representation, some languages, such as Bangla, which have a limited amount of resources, continue to be quite unfocused due to a lack of standard datasets. In order to resolve this issue, we have introduced a new dataset named 'Biboron', in which we manually gathered information in Bangla of images extracted from the widely available Flickr30k dataset that were then post-processed and examined for quality assurance. 'Biboron' contains 1,58,915 distinct sentences describing 31,783 images which further specifies the versatile nature of the dataset. Furthermore, we have presented two models in order to enhance the automated extraction of visual information from images and represent in Bangla. The first model includes Local Attention, whilst the second model is based on Multi-Head Attention with Transformers. The image feature extractor of the models utilized VGG16, while bidirectional LSTM backed by CuDNN was used in the decoder network. The BLEU scores suggest that the second model appears to outperform the first one in terms of generating more relevant textual representations from images by achieving BLEU-1, BLEU -2, BLEU -3, BLEU -4 scores of 0.78, 0.53, 0.37, 0.21 respectively.
dc.description.versionPublished
dc.format.extent436-441
dc.identifier.citationN. Nawer, M. S. I. Khan, M. M. Alam, M. H. K. Mehedi and A. A. Rasel, "Transformation of Visual Information into Bangla Textual Representation," 2023 IEEE 13th Annual Computing and Communication Workshop and Conference (CCWC), Las Vegas, NV, USA, 2023, pp. 0436-0441, doi: 10.1109/CCWC57344.2023.10099345.
dc.identifier.doi10.1109/CCWC57344.2023.10099345
dc.identifier.issn9798350332865
dc.identifier.other2-s2.0-85156191088
dc.identifier.urihttps://hdl.handle.net/10361/28655
dc.language.isoen_US
dc.publisherInstitute of Electrical and Electronics Engineers Inc.
dc.relation.hasversion10.1109/CCWC57344.2023.10099345
dc.relation.ispartof2023 IEEE 13th Annual Computing and Communication Workshop and Conference Ccwc 2023
dc.relation.ispartofseries2023 IEEE 13th Annual Computing and Communication Workshop and Conference Ccwc 2023
dc.relation.urihttps://ieeexplore.ieee.org/document/10099345
dc.subjectPattern recognition systems
dc.subjectBLEU score
dc.subjectImage processing
dc.subjectLocal attention
dc.subjectMulti-head attention
dc.subjectComputational linguistics
dc.subject.lcshImage processing--Digital techniques--Data processing.
dc.subject.lcshNatural language processing (Computer science).
dc.subject.lcshDeep learning (Machine learning).
dc.subject.lcshBengali language--Data processing.
dc.titleTransformation of visual information into Bangla textual representation
dc.typeConference Proceeding
person.affiliation.nameBRAC University
person.affiliation.nameBRAC University
person.affiliation.nameBRAC University
person.affiliation.nameBRAC University
person.affiliation.nameBRAC University
person.identifier.scopus-author-id58222205600
person.identifier.scopus-author-id58144321800
person.identifier.scopus-author-id58184012600
person.identifier.scopus-author-id57422283000
person.identifier.scopus-author-id56495276900

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
IMG_8345.jpg
Size:
27.35 KB
Format:
Joint Photographic Experts Group/JPEG File Interchange Format (JFIF)

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: