Efficient action detection in video sequences: A hybrid CRNN-transformer approach with attention and HPC strategies
Loading...
Date
Publisher
Institute of Electrical and Electronics Engineers Inc.
Citation
M. F. Ishrak, M. A. Hossain, A. Mitra, S. A. Araf, M. Rahman and M. O. Faroque, "Efficient Action Detection in Video Sequences: A Hybrid CRNN-Transformer Approach with Attention and HPC Strategies," 2024 27th International Conference on Computer and Information Technology (ICCIT), Cox's Bazar, Bangladesh, 2024, pp. 1487-1492, doi: 10.1109/ICCIT64611.2024.11021885.
Abstract
Due to the digital processing advancement and availability of big data the field of action recognition had a great growth in recent years. This work introduces a new action detection model that combines a hybrid CRNN-Transformer model with current computational methods. The proposed model architecture presents a shift from mainstream models, although it provides superior levels of accuracy and performance. One of these is the use of the attention mechanism for focusing selectively, the second stage focuses on object identification and proposing the regions of interest, and the third uses optimized High-Performance Computing methods in the process. The model adopts the highly effective Transformer-based architecture, applied to the spatiotemporal sequences of videos in the form of 5D tensors using ResNet-164 as a base network, thus enhancing the recognition capacities. Moreover, there are Long Short-Term Memory (LSTM) layer that helps capture dependency across time, so predictive chance will increase as well. The proposed hybrid model is better compared to the current approaches; the results of performance analysis on the UCF101 dataset yielded an mAP of 50.1% and a mean class-wise AP (mcAP) of 80.8% which outperforms TSN, LRCN, I3D, and all the other previous models. Other additional techniques are Mixed Precision Training, XLA Optimization, and Gradient Checkpointing to enhance the computational speed. Evaluation measures, including Average Precision (AP), confirm the efficiency of the proposed method in all aspects of action detection.
LC Subject Headings
Description
Publisher Link
Type
Conference Proceeding