Enhancing object detection interpretability for class-discriminative visualizations with Grad-CAM
Loading...
Date
Publisher
Institute of Electrical and Electronics Engineers Inc.
Citation
M. F. Ishrak, J. Nahar, F. A. Faiza and L. M. Siddiqua, "Enhancing Object Detection Interpretability for Class-Discriminative Visualizations with Grad-CAM," 2024 27th International Conference on Computer and Information Technology (ICCIT), Cox's Bazar, Bangladesh, 2024, pp. 1493-1498, doi: 10.1109/ICCIT64611.2024.11022089.
Abstract
Computer vision has benefited greatly from deep learning; it has saved the day by making it possible to achieve marvelous performances in visual recognition challenges. In the field of object detection, not only the detection of objects in video and image data but also the localization of those objects is essential. This work proposes a new way to perform object detection based on Grad-CAM with the main emphasis on the ILSVRC-17 dataset. A higher level of generality made it possible to adapt Grad-CAM originally designed for representing the image classification to highlight action regions in the frames of the video and images. Our visualizations offer several advantages: ig they are useful in the following ways: (a) They offer guidance on model failure modes and deepen our understanding of how seemingly irrational predictions are logical (b) They perform effectively and correctly on the ILSVRC-17 weakly supervised localization task compared to previous methodologies (c) They demonstrate resilience against adversarial disturbances (d) They are closer to the true model in their representations and (e) They help the models to generalize. The proposed work shows that the idea of Grad-CAM helps in increasing action localization effectiveness and, at the same time, provides meaningful information about the model's performance and decision-making mechanisms. These results demonstrate that Grad-CAM can be an efficacious tool for action recognition and spatial prediction, which lays the foundation for video and image processing, surveillance, and human-machine interaction. Lastly, we also talk about the possibilities of refining object detection approaches and increasing the model's interpretability, which we believe is important for the field of computer vision.
LC Subject Headings
Description
Publisher Link
Type
Conference Proceeding