Efficient action detection in video sequences: A hybrid CRNN-transformer approach with attention and HPC strategies

bracu.type.groupResearch Publications
datacite.rightsMetadata Only
dc.contributor.authorIshrak, Md Fatin
dc.contributor.authorHossain, Mohammad Ahad
dc.contributor.authorMitra, Anindya
dc.contributor.authorAraf, Sheikh Abdul
dc.contributor.authorRahman, Mushfiqur
dc.contributor.authorFaroque, Md Omar
dc.contributor.departmentDepartment of Computer Science and Engineering
dc.date.accessioned2026-09-30T04:38:17Z
dc.date.available2026-09-30T04:38:17Z
dc.date.issued2024-01-01
dc.description.abstractDue to the digital processing advancement and availability of big data the field of action recognition had a great growth in recent years. This work introduces a new action detection model that combines a hybrid CRNN-Transformer model with current computational methods. The proposed model architecture presents a shift from mainstream models, although it provides superior levels of accuracy and performance. One of these is the use of the attention mechanism for focusing selectively, the second stage focuses on object identification and proposing the regions of interest, and the third uses optimized High-Performance Computing methods in the process. The model adopts the highly effective Transformer-based architecture, applied to the spatiotemporal sequences of videos in the form of 5D tensors using ResNet-164 as a base network, thus enhancing the recognition capacities. Moreover, there are Long Short-Term Memory (LSTM) layer that helps capture dependency across time, so predictive chance will increase as well. The proposed hybrid model is better compared to the current approaches; the results of performance analysis on the UCF101 dataset yielded an mAP of 50.1% and a mean class-wise AP (mcAP) of 80.8% which outperforms TSN, LRCN, I3D, and all the other previous models. Other additional techniques are Mixed Precision Training, XLA Optimization, and Gradient Checkpointing to enhance the computational speed. Evaluation measures, including Average Precision (AP), confirm the efficiency of the proposed method in all aspects of action detection.
dc.description.versionPublished
dc.format.extent6 Pages
dc.identifier.citationM. F. Ishrak, M. A. Hossain, A. Mitra, S. A. Araf, M. Rahman and M. O. Faroque, "Efficient Action Detection in Video Sequences: A Hybrid CRNN-Transformer Approach with Attention and HPC Strategies," 2024 27th International Conference on Computer and Information Technology (ICCIT), Cox's Bazar, Bangladesh, 2024, pp. 1487-1492, doi: 10.1109/ICCIT64611.2024.11021885.
dc.identifier.doi10.1109/ICCIT64611.2024.11021885
dc.identifier.issn9798331519094
dc.identifier.other2-s2.0-105009031703
dc.identifier.urihttps://hdl.handle.net/10361/30300
dc.language.isoen_US
dc.publisherInstitute of Electrical and Electronics Engineers Inc.
dc.relation.hasversion10.1109/ICCIT64611.2024.11021885
dc.relation.ispartof2024 27th International Conference on Computer and Information Technology Iccit 2024 Proceedings
dc.relation.ispartofseries2024 27th International Conference on Computer and Information Technology Iccit 2024 Proceedings
dc.relation.urihttps://ieeexplore.ieee.org/document/11021885
dc.subjectTraining
dc.subjectAnalytical models
dc.subjectThree-dimensional displays
dc.subjectComputational modeling
dc.subjectHigh performance computing
dc.subjectVideo sequences
dc.subjectComputer architecture
dc.subjectTransformers
dc.subjectLong short term memory
dc.subjectAction detection
dc.subjectHigh-performance computing
dc.subject.lcshHuman activity recognition.
dc.titleEfficient action detection in video sequences: A hybrid CRNN-transformer approach with attention and HPC strategies
dc.typeConference Proceeding
person.affiliation.nameBRAC University
person.affiliation.nameBRAC University
person.affiliation.nameBRAC University
person.affiliation.nameBRAC University
person.affiliation.nameBRAC University
person.affiliation.nameBRAC University
person.identifier.scopus-author-id59484185300
person.identifier.scopus-author-id59963053500
person.identifier.scopus-author-id59964178600
person.identifier.scopus-author-id59963285100
person.identifier.scopus-author-id57706339100
person.identifier.scopus-author-id59962830000

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
IMG_8345.jpg
Size:
27.35 KB
Format:
Joint Photographic Experts Group/JPEG File Interchange Format (JFIF)

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: