Automated image caption generation using deep learning
| bracu.type.group | Research Publications | |
| datacite.rights | Metadata Only | |
| dc.contributor.author | Hasan, Mahmudul | |
| dc.contributor.author | Prithila, Sara Jerin | |
| dc.contributor.author | Al Ahasan, Tamim | |
| dc.contributor.author | Hassan, Mahmudul | |
| dc.contributor.author | Hossain, Abid | |
| dc.contributor.author | Bhuiyan, Sania Azhmee | |
| dc.contributor.author | Alim Rasel, Annajiat | |
| dc.contributor.department | Department of Computer Science and Engineering | |
| dc.date.accessioned | 2026-09-22T10:52:52Z | |
| dc.date.available | 2026-09-22T10:52:52Z | |
| dc.date.issued | 2023-01-01 | |
| dc.description.abstract | Image captioning means generating relevant texts from images that describe the image. With the growing advancement of deep learning, automatic image caption generation has become an exciting problem among researchers. Image captioning is used in various fields of computer science including computer vision and NLP. It is helpful in many cases such as visual impairment. Though many research works have already been published on this topic using different deep learning models, we can work with various features of the deep learning models to achieve better accuracy. In this paper, we are focusing on combining several deep-learning models to get the desired accuracy. VGG16 model is used which is a CNN architecture to extract essential features from the image. For the generation of relevant captions, LSTM is adopted rather than RNNs as LSTM generates better captions compared to RNNs and it consumes less time. Finally, the model is trained on Flickr 8k datasets. | |
| dc.description.version | Published | |
| dc.format.extent | 6 Pages | |
| dc.identifier.citation | M. Hasan et al., "Automated Image Caption Generation using Deep Learning," 2023 26th International Conference on Computer and Information Technology (ICCIT), Cox's Bazar, Bangladesh, 2023, pp. 1-6, doi: 10.1109/ICCIT60459.2023.10441058. | |
| dc.identifier.doi | 10.1109/ICCIT60459.2023.10441058 | |
| dc.identifier.issn | 9798350359015 | |
| dc.identifier.other | 2-s2.0-85187380165 | |
| dc.identifier.uri | https://hdl.handle.net/10361/30156 | |
| dc.language.iso | en_US | |
| dc.publisher | Institute of Electrical and Electronics Engineers Inc. | |
| dc.relation.hasversion | 10.1109/ICCIT60459.2023.10441058 | |
| dc.relation.ispartof | 2023 26th International Conference on Computer and Information Technology Iccit 2023 | |
| dc.relation.ispartofseries | 2023 26th International Conference on Computer and Information Technology Iccit 2023 | |
| dc.relation.uri | https://ieeexplore.ieee.org/document/10441058 | |
| dc.subject | Deep learning | |
| dc.subject | Computational modeling | |
| dc.subject | Visual impairment | |
| dc.subject | Feature extraction | |
| dc.subject | Multimedia communication | |
| dc.subject | Task analysis | |
| dc.subject | Image captioning | |
| dc.subject | Deep learning | |
| dc.subject.lcsh | Computer vision. | |
| dc.subject.lcsh | Image processing--Digital techniques. | |
| dc.title | Automated image caption generation using deep learning | |
| dc.type | Conference Proceeding | |
| person.affiliation.name | BRAC University | |
| person.affiliation.name | BRAC University | |
| person.affiliation.name | BRAC University | |
| person.affiliation.name | BRAC University | |
| person.affiliation.name | BRAC University | |
| person.affiliation.name | BRAC University | |
| person.affiliation.name | BRAC University | |
| person.identifier.scopus-author-id | 59283217400 | |
| person.identifier.scopus-author-id | 58908960600 | |
| person.identifier.scopus-author-id | 58931249400 | |
| person.identifier.scopus-author-id | 59282515000 | |
| person.identifier.scopus-author-id | 58931028500 | |
| person.identifier.scopus-author-id | 58930278800 | |
| person.identifier.scopus-author-id | 56495276900 |