Chakravorty, Tirthendu ProsadAbeer, MobashraBaroi, Shaiane PremaRoy, SristyKarim, Dewan Ziaul2026-08-132026-08-132023-01-01T. P. Chakravorty, M. Abeer, S. P. Baroi, S. Roy and D. Z. Karim, "Interpretable Violence Detection Using Separable Convolution and Bidirectional LSTM," 2023 IEEE Asia-Pacific Conference on Computer Science and Data Engineering (CSDE), Nadi, Fiji, 2023, pp. 01-06, doi: 10.1109/CSDE59766.2023.10487654.97983503410722-s2.0-85190616644https://hdl.handle.net/10361/29062With the increasing demand for security concerns, security measures in all kinds of locations have become more dependent on the integration of surveillance cameras. Such devices are everywhere, which has significantly helped in the fight against violent crime. Continuous human monitoring becomes a laborious effort and frequently results in delayed reactions in bigger systems. Therefore, automated detection of aggressive behavior in surveillance systems can improve remote monitoring and boost reaction precision. The joint use of Convolutional Neural Networks (CNN) and Recurrent Neural Networks (RNN) have previously been used in studies to identify potentially violent actions accurately but many have struggled to minimize the computational resources required. The purpose of this study is to gain from decreased computing costs while preserving optimality for real-world applications. Hence, in this study, a rigorous model based on combinations of CNN and RNN architectures has been developed for spatiotemporal features from videos. To incorporate explainability for the AI's decision, the spatial feature extractor utilizes the LIME model, based on Explainable AI (XAI). The performance of the model then has been thoroughly analyzed using a compilation of several benchmark datasets. The suggested spatiotemporal feature-based model, in the final analysis, achieved a test accuracy of 98.75%.6 Pagesen-USAnalytical modelsRecurrent neural networksComputational modelingSurveillanceFeature extractionSpatiotemporal phenomenaConvolutional neural networksViolence detectionDeep learningRecurrent neural networksImage processingClosed-circuit television.Computational intelligence.Electronic surveillance.Deep learning (Machine learning).Interpretable violence detection using separable convolution and bidirectional LSTMConference Proceeding10.1109/CSDE59766.2023.10487654