Sultan T.Sarker M.M.H.Bhuiyan, Md. Khairul BasharSaib M.Islam M.S.Hossain M.R.El-Shafai W.Azar A.T.Njima C.B.2026-09-102026-09-102025-01-01T. Sultan et al., "ViTDeBERTaNews: A Comparative Study of Single-modal, Multimodal, and LLM Techniques for Detecting Fake News," 2025 International Conference on Control, Automation and Diagnosis (ICCAD), Barcelona, Spain, 2025, pp. 1-6, doi: 10.1109/ICCAD64771.2025.11099172.97983315119132-s2.0-105014508621https://hdl.handle.net/10361/29832The spread of fake news online poses a significant challenge, particularly as social media increasingly combines images and text. To address this, we present the ViT+DeBERTaNews model, which effectively merges visual and textual information for fake news detection. This model utilizes ViT for detailed visual feature extraction and DeBERTa for deep textual understanding. Our experiments demonstrate its effectiveness, achieving 94.8% accuracy on the Weibo dataset and 92.6% on Twitter. The model's precision, recall, and F1 scores for fake news detection were 0.967, 0.945, and 0.956 on Weibo, and 0.929, 0.930, and 0.933 on Twitter, respectively. For real news, it scored 0.968, 0.944, and 0.956 on Weibo, and 0.925, 0.944, and 0.956 on Twitter. In contrast, text-based models like GPT-2 Epoch 3, while strong in precision and recall, are limited by their text-only approach. GPT-4 also faced challenges in recall on the Weibo dataset, indicating the need for task-specific optimizations. These findings underscore the necessity of advanced multimodal models for effective fake news detection.6 Pagesen-USVisualizationAccuracySocial networking (online)BlogsTransfer learningMetadataFeature extractionRobustnessFake newsOptimizationMultimodal deep learningLarge language modelsFake news.Disinformation.Natural language processing (Computer science).ViTDeBERTaNews: A comparative study of single-modal, multimodal, and LLM techniques for detecting fake newsConference Proceeding10.1109/ICCAD64771.2025.11099172