IntentVLM: Open-Vocabulary Intention Recognition through Forward-Inverse Modeling with Video-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Rahimi, Hamed, Grislain, Clemence, Cretides, Adrien Jacquet, Sigaud, Olivier, Chetouani, Mohamed |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
I-FailSense: Towards General Robotic Failure Detection with Vision-Language Models
by: Grislain, Clemence, et al.
Published: (2025)
by: Grislain, Clemence, et al.
Published: (2025)
Controlling Intent Expressiveness in Robot Motion with Diffusion Models
by: Shi, Wenli, et al.
Published: (2025)
by: Shi, Wenli, et al.
Published: (2025)
Iterative On-Policy Refinement of Hierarchical Diffusion Policies for Language-Conditioned Manipulation
by: Grislain, Clemence, et al.
Published: (2026)
by: Grislain, Clemence, et al.
Published: (2026)
Encoding Predictability and Legibility for Style-Conditioned Diffusion Policy
by: Crétides, Adrien Jacquet, et al.
Published: (2026)
by: Crétides, Adrien Jacquet, et al.
Published: (2026)
When Drawing Is Not Enough: Exploring Spontaneous Speech with Sketch for Intent Alignment in Multimodal LLMs
by: Shi, Weiyan, et al.
Published: (2026)
by: Shi, Weiyan, et al.
Published: (2026)
SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision Viewers
by: Ning, Zheng, et al.
Published: (2024)
by: Ning, Zheng, et al.
Published: (2024)
USER-VLM 360: Personalized Vision Language Models with User-aware Tuning for Social Human-Robot Interactions
by: Rahimi, Hamed, et al.
Published: (2025)
by: Rahimi, Hamed, et al.
Published: (2025)
Explainable Multimodal Emotion Recognition
by: Lian, Zheng, et al.
Published: (2023)
by: Lian, Zheng, et al.
Published: (2023)
VCEMO: Multi-Modal Emotion Recognition for Chinese Voiceprints
by: Tang, Jinghua, et al.
Published: (2024)
by: Tang, Jinghua, et al.
Published: (2024)
Physical-aware Cross-modal Adversarial Network for Wearable Sensor-based Human Action Recognition
by: Ni, Jianyuan, et al.
Published: (2023)
by: Ni, Jianyuan, et al.
Published: (2023)
Multimodal Infusion Tuning for Large Models
by: Sun, Hao, et al.
Published: (2024)
by: Sun, Hao, et al.
Published: (2024)
MV-Crafter: An Intelligent System for Music-guided Video Generation
by: Chen, Chuer, et al.
Published: (2025)
by: Chen, Chuer, et al.
Published: (2025)
MIMOSA: Human-AI Co-Creation of Computational Spatial Audio Effects on Videos
by: Ning, Zheng, et al.
Published: (2024)
by: Ning, Zheng, et al.
Published: (2024)
Anchorage: Visual Analysis of Satisfaction in Customer Service Videos via Anchor Events
by: Wong, Kam Kwai, et al.
Published: (2023)
by: Wong, Kam Kwai, et al.
Published: (2023)
Language-Guided Multimodal Texture Authoring via Generative Models
by: Qian, Wanli, et al.
Published: (2026)
by: Qian, Wanli, et al.
Published: (2026)
The Potential of Olfactory Stimuli in Stress Reduction through Virtual Reality
by: Valdivieso, Yasmin Elsaddik, et al.
Published: (2025)
by: Valdivieso, Yasmin Elsaddik, et al.
Published: (2025)
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations
by: Zhao, Baoquan, et al.
Published: (2025)
by: Zhao, Baoquan, et al.
Published: (2025)
Vocalize: Lead Acquisition and User Engagement through Gamified Voice Competitions
by: Teskeredzic, Edvin, et al.
Published: (2025)
by: Teskeredzic, Edvin, et al.
Published: (2025)
3D Modelling to Address Pandemic Challenges: A Project-Based Learning Methodology
by: Rocha, Tânia, et al.
Published: (2024)
by: Rocha, Tânia, et al.
Published: (2024)
Project Lx Conventos: Travelling through space and time in Lisbon's religious buildings
by: Gouveia, Joao, et al.
Published: (2024)
by: Gouveia, Joao, et al.
Published: (2024)
The Rhythm of Tai Chi: Revitalizing Cultural Heritage in Virtual Reality through Interactive Visuals
by: Wang, Xianghan
Published: (2025)
by: Wang, Xianghan
Published: (2025)
On Parallelism in Music and Language: A Perspective from Symbol Emergence Systems based on Probabilistic Generative Models
by: Taniguchi, Tadahiro
Published: (2025)
by: Taniguchi, Tadahiro
Published: (2025)
Human-Machine Collaboration-Guided Space Design: Combination of Machine Learning Models and Humanistic Design Concepts
by: Yang, Yuxuan
Published: (2025)
by: Yang, Yuxuan
Published: (2025)
DIGITWISE: Digital Twin-based Modeling of Adaptive Video Streaming Engagement
by: Artioli, Emanuele, et al.
Published: (2025)
by: Artioli, Emanuele, et al.
Published: (2025)
Laugh at Your Own Pace: Basic Performance Evaluation of Language Learning Assistance by Adjustment of Video Playback Speeds Based on Laughter Detection
by: Nishida, Naoto, et al.
Published: (2025)
by: Nishida, Naoto, et al.
Published: (2025)
VIPER: Visual Perception and Explainable Reasoning for Sequential Decision-Making
by: Aissi, Mohamed Salim, et al.
Published: (2025)
by: Aissi, Mohamed Salim, et al.
Published: (2025)
Facilitating Daily Practice in Intangible Cultural Heritage through Virtual Reality: A Case Study of Traditional Chinese Flower Arrangement
by: Wang, Yingna, et al.
Published: (2025)
by: Wang, Yingna, et al.
Published: (2025)
Photoshop Batch Rendering Using Actions for Stylistic Video Editing
by: De La Fuente, Tessa
Published: (2025)
by: De La Fuente, Tessa
Published: (2025)
Crafting Dynamic Virtual Activities with Advanced Multimodal Models
by: Li, Changyang, et al.
Published: (2024)
by: Li, Changyang, et al.
Published: (2024)
Human-Machine Ritual: Synergic Performance through Real-Time Motion Recognition
by: Cai, Zhuodi, et al.
Published: (2025)
by: Cai, Zhuodi, et al.
Published: (2025)
Vidmento: Creating Video Stories Through Context-Aware Expansion With Generative Video
by: Yeh, Catherine, et al.
Published: (2026)
by: Yeh, Catherine, et al.
Published: (2026)
A Multimedia Analytics Model for the Foundation Model Era
by: Worring, Marcel, et al.
Published: (2025)
by: Worring, Marcel, et al.
Published: (2025)
Plug-and-Play Clarifier: A Zero-Shot Multimodal Framework for Egocentric Intent Disambiguation
by: Yang, Sicheng, et al.
Published: (2025)
by: Yang, Sicheng, et al.
Published: (2025)
Development and Evaluation of Dental Image Exchange and Management System: A User-Centered Perspective
by: Rahimi, B, et al.
Published: (2022)
by: Rahimi, B, et al.
Published: (2022)
The perceptual gap between video see-through displays and natural human vision
by: Wang, Jialin, et al.
Published: (2026)
by: Wang, Jialin, et al.
Published: (2026)
Towards Open-Vocabulary Video Semantic Segmentation
by: Li, Xinhao, et al.
Published: (2024)
by: Li, Xinhao, et al.
Published: (2024)
MindCine: Multimodal EEG-to-Video Reconstruction with Large-Scale Pretrained Models
by: Zhou, Tian-Yi, et al.
Published: (2026)
by: Zhou, Tian-Yi, et al.
Published: (2026)
Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion Model
by: He, Xu, et al.
Published: (2024)
by: He, Xu, et al.
Published: (2024)
To Fuse or to Drop? Dual-Path Learning for Resolving Modality Conflicts in Multimodal Emotion Recognition
by: Yu, Yangchen, et al.
Published: (2026)
by: Yu, Yangchen, et al.
Published: (2026)
MotiBo: The Impact of Interactive Digital Storytelling Robots on Student Motivation through Self-Determination Theory
by: Fung, Ka Yan, et al.
Published: (2026)
by: Fung, Ka Yan, et al.
Published: (2026)
Similar Items
-
I-FailSense: Towards General Robotic Failure Detection with Vision-Language Models
by: Grislain, Clemence, et al.
Published: (2025) -
Controlling Intent Expressiveness in Robot Motion with Diffusion Models
by: Shi, Wenli, et al.
Published: (2025) -
Iterative On-Policy Refinement of Hierarchical Diffusion Policies for Language-Conditioned Manipulation
by: Grislain, Clemence, et al.
Published: (2026) -
Encoding Predictability and Legibility for Style-Conditioned Diffusion Policy
by: Crétides, Adrien Jacquet, et al.
Published: (2026) -
When Drawing Is Not Enough: Exploring Spontaneous Speech with Sketch for Intent Alignment in Multimodal LLMs
by: Shi, Weiyan, et al.
Published: (2026)