A multi-modal intelligent virtual assistant for extended reality environments based on computer vision, natural language processing, and large language models

Loading...
Thumbnail Image

Publisher

Institute of Electrical and Electronics Engineers Inc.

Citation

M. S. Islam, M. S. Shifa, M. S. Al Masud, M. N. W. Omi and R. Z. Eashan, "A Multi-Modal Intelligent Virtual Assistant for Extended Reality Environments Based on Computer Vision, Natural Language Processing, and Large Language Models," 2026 IEEE 2nd International Conference on Quantum Photonics, Artificial Intelligence & Networking (QPAIN), Chittagong, Bangladesh, 2026, pp. 1-6, doi: 10.1109/QPAIN69676.2026.11546475.

Abstract

Extended Reality (XR) environments increasingly demand natural and adaptive interaction mechanisms beyond traditional menu-driven interfaces. This paper presents a multi-modal intelligent virtual assistant for XR that jointly understands user intent and contextual information by integrating computer vision, natural language processing (NLP), and large language models (LLMs). The key novelty of the proposed approach lies in a context-grounded reasoning framework that fuses real-time visual perception with conversational state to generate reliable and explainable assistant actions. Visual information, such as objects, spatial cues, and user activities, is transformed into a structured scene representation, which is combined with linguistic intent signals to support accurate decision-making in immersive environments. Unlike conventional virtual assistants that process visual and textual inputs independently, the proposed system employs confidence-aware multi-modal fusion and retrieval-augmented LLM reasoning to maintain contextual consistency and reduce ambiguous responses. The framework is designed in a modular and engine-agnostic manner, enabling seamless integration with common XR platforms for applications including object-aware guidance, interactive task support, and contextual information retrieval. Experimental evaluation is conducted using task-completion scenarios and user experience metrics, demonstrating improved interaction naturalness, reduced user effort, and higher contextual accuracy compared to unimodal and loosely coupled multimodal baselines. These results highlight the effectiveness of multi-modal grounding for developing intelligent, trustworthy, and user-centric virtual assistants in extended reality environments.

Description

Type

Conference Proceeding