MemeFusionNet: A cross-linguistic multimodal model for identifying troll memes

Citation

T. Sultan et al., "MemeFusionNet: A Cross-Linguistic Multimodal Model for Identifying Troll Memes," 2025 International Conference on Control, Automation and Diagnosis (ICCAD), Barcelona, Spain, 2025, pp. 1-6, doi: 10.1109/ICCAD64771.2025.11099160.

Abstract

The proliferation of troll memes, which exploit textual and visual elements to propagate misinformation and incite negativity, presents a critical challenge for online content moderation. Existing methods often struggle with cross-linguistic generalization, multimodal fusion, and contextual understanding, limiting their effectiveness in multilingual environments. To address these gaps, we propose MemeFusionNet, a transformer-driven multimodal fusion framework that effectively captures the intricate relationships between images and text. MemeFusionNet integrates a cross-modal attention mechanism based on ViLT to enhance contextual awareness and better detect implicit troll content, such as sarcasm and cultural nuances. Our model demonstrates superior performance on Bangla and English meme datasets, achieving 86% and 96% accuracy, respectively, outperforming all existing benchmarks. Its scalable architecture ensures robust cross-lingual adaptability, making it well-suited for large-scale, real-time content moderation.

Description

Type

Conference Proceeding