Enhancing object detection interpretability for class-discriminative visualizations with Grad-CAM

bracu.type.groupResearch Publications
datacite.rightsMetadata Only
dc.contributor.authorIshrak, Md Fatin
dc.contributor.authorNahar, Jannatun
dc.contributor.authorFaiza, Fairooz Afnad
dc.contributor.authorSiddiqua, Lamia Mahzabin
dc.contributor.departmentDepartment of Computer Science and Engineering
dc.date.accessioned2026-09-30T06:38:35Z
dc.date.available2026-09-30T06:38:35Z
dc.date.issued2024-01-01
dc.description.abstractComputer vision has benefited greatly from deep learning; it has saved the day by making it possible to achieve marvelous performances in visual recognition challenges. In the field of object detection, not only the detection of objects in video and image data but also the localization of those objects is essential. This work proposes a new way to perform object detection based on Grad-CAM with the main emphasis on the ILSVRC-17 dataset. A higher level of generality made it possible to adapt Grad-CAM originally designed for representing the image classification to highlight action regions in the frames of the video and images. Our visualizations offer several advantages: ig they are useful in the following ways: (a) They offer guidance on model failure modes and deepen our understanding of how seemingly irrational predictions are logical (b) They perform effectively and correctly on the ILSVRC-17 weakly supervised localization task compared to previous methodologies (c) They demonstrate resilience against adversarial disturbances (d) They are closer to the true model in their representations and (e) They help the models to generalize. The proposed work shows that the idea of Grad-CAM helps in increasing action localization effectiveness and, at the same time, provides meaningful information about the model's performance and decision-making mechanisms. These results demonstrate that Grad-CAM can be an efficacious tool for action recognition and spatial prediction, which lays the foundation for video and image processing, surveillance, and human-machine interaction. Lastly, we also talk about the possibilities of refining object detection approaches and increasing the model's interpretability, which we believe is important for the field of computer vision.
dc.description.versionPublished
dc.format.extent1493-1498
dc.identifier.citationM. F. Ishrak, J. Nahar, F. A. Faiza and L. M. Siddiqua, "Enhancing Object Detection Interpretability for Class-Discriminative Visualizations with Grad-CAM," 2024 27th International Conference on Computer and Information Technology (ICCIT), Cox's Bazar, Bangladesh, 2024, pp. 1493-1498, doi: 10.1109/ICCIT64611.2024.11022089.
dc.identifier.doi10.1109/ICCIT64611.2024.11022089
dc.identifier.issn9798331519094
dc.identifier.other2-s2.0-105009154108
dc.identifier.urihttps://hdl.handle.net/10361/30306
dc.language.isoen_US
dc.publisherInstitute of Electrical and Electronics Engineers Inc.
dc.relation.hasversion10.1109/ICCIT64611.2024.11022089
dc.relation.ispartof2024 27th International Conference on Computer and Information Technology Iccit 2024 Proceedings
dc.relation.ispartofseries2024 27th International Conference on Computer and Information Technology Iccit 2024 Proceedings
dc.relation.urihttps://ieeexplore.ieee.org/document/11022089
dc.subjectLocation awareness
dc.subjectVisualization
dc.subjectImage segmentation
dc.subjectComputer vision
dc.subjectAccuracy
dc.subjectComputational modeling
dc.subjectSurveillance
dc.subjectObject detection
dc.subjectPredictive models
dc.subjectReliability
dc.subjectGrad-CAM
dc.subjectExplainable AI
dc.subject.lcshComputer vision.
dc.subject.lcshImage processing--Digital techniques.
dc.titleEnhancing object detection interpretability for class-discriminative visualizations with Grad-CAM
dc.typeConference Proceeding
person.affiliation.nameBRAC University
person.affiliation.nameBRAC University
person.affiliation.nameBRAC University
person.affiliation.nameBRAC University
person.identifier.scopus-author-id59484185300
person.identifier.scopus-author-id59962838600
person.identifier.scopus-author-id59964185900
person.identifier.scopus-author-id59963061000

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
IMG_8345.jpg
Size:
27.35 KB
Format:
Joint Photographic Experts Group/JPEG File Interchange Format (JFIF)

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: