OLViT: Multi-Modal State Tracking via Attention-Based Embeddings for Video-Grounded Dialog
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Abdessaied, Adnen, von Hochmeister, Manuel, Bulling, Andreas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multi-Modal Video Dialog State Tracking in the Wild
von: Abdessaied, Adnen, et al.
Veröffentlicht: (2024)
von: Abdessaied, Adnen, et al.
Veröffentlicht: (2024)
V$^2$Dial: Unification of Video and Visual Dialog via Multimodal Experts
von: Abdessaied, Adnen, et al.
Veröffentlicht: (2025)
von: Abdessaied, Adnen, et al.
Veröffentlicht: (2025)
ActionDiffusion: An Action-aware Diffusion Model for Procedure Planning in Instructional Videos
von: Shi, Lei, et al.
Veröffentlicht: (2024)
von: Shi, Lei, et al.
Veröffentlicht: (2024)
CLAD: Constrained Latent Action Diffusion for Vision-Language Procedure Planning
von: Shi, Lei, et al.
Veröffentlicht: (2025)
von: Shi, Lei, et al.
Veröffentlicht: (2025)
Learning User Embeddings from Human Gaze for Personalised Saliency Prediction
von: Strohm, Florian, et al.
Veröffentlicht: (2024)
von: Strohm, Florian, et al.
Veröffentlicht: (2024)
Grounding is All You Need? Dual Temporal Grounding for Video Dialog
von: Qin, You, et al.
Veröffentlicht: (2024)
von: Qin, You, et al.
Veröffentlicht: (2024)
UP-FacE: User-predictable Fine-grained Face Shape Editing
von: Strohm, Florian, et al.
Veröffentlicht: (2024)
von: Strohm, Florian, et al.
Veröffentlicht: (2024)
HAIFAI: Human-AI Interaction for Mental Face Reconstruction
von: Strohm, Florian, et al.
Veröffentlicht: (2024)
von: Strohm, Florian, et al.
Veröffentlicht: (2024)
VQA-MHUG: A Gaze Dataset to Study Multimodal Neural Attention in Visual Question Answering
von: Sood, Ekta, et al.
Veröffentlicht: (2021)
von: Sood, Ekta, et al.
Veröffentlicht: (2021)
Multimodal Integration of Human-Like Attention in Visual Question Answering
von: Sood, Ekta, et al.
Veröffentlicht: (2021)
von: Sood, Ekta, et al.
Veröffentlicht: (2021)
UBATrack: Spatio-Temporal State Space Model for General Multi-Modal Tracking
von: Liang, Qihua, et al.
Veröffentlicht: (2026)
von: Liang, Qihua, et al.
Veröffentlicht: (2026)
DialogCC: An Automated Pipeline for Creating High-Quality Multi-Modal Dialogue Dataset
von: Lee, Young-Jun, et al.
Veröffentlicht: (2022)
von: Lee, Young-Jun, et al.
Veröffentlicht: (2022)
Single-Model and Any-Modality for Video Object Tracking
von: Wu, Zongwei, et al.
Veröffentlicht: (2023)
von: Wu, Zongwei, et al.
Veröffentlicht: (2023)
Multi-Modal Generative Embedding Model
von: Ma, Feipeng, et al.
Veröffentlicht: (2024)
von: Ma, Feipeng, et al.
Veröffentlicht: (2024)
DM$^3$T: Harmonizing Modalities via Diffusion for Multi-Object Tracking
von: Li, Weiran, et al.
Veröffentlicht: (2025)
von: Li, Weiran, et al.
Veröffentlicht: (2025)
SwiTrack: Tri-State Switch for Cross-Modal Object Tracking
von: Xu, Boyue, et al.
Veröffentlicht: (2025)
von: Xu, Boyue, et al.
Veröffentlicht: (2025)
SpikeMba: Multi-Modal Spiking Saliency Mamba for Temporal Video Grounding
von: Li, Wenrui, et al.
Veröffentlicht: (2024)
von: Li, Wenrui, et al.
Veröffentlicht: (2024)
Enhancing Visual Dialog State Tracking through Iterative Object-Entity Alignment in Multi-Round Conversations
von: Pang, Wei, et al.
Veröffentlicht: (2024)
von: Pang, Wei, et al.
Veröffentlicht: (2024)
Exploiting Modality-Specific Features For Multi-Modal Manipulation Detection And Grounding
von: Wang, Jiazhen, et al.
Veröffentlicht: (2023)
von: Wang, Jiazhen, et al.
Veröffentlicht: (2023)
GazeMotion: Gaze-guided Human Motion Forecasting
von: Hu, Zhiming, et al.
Veröffentlicht: (2024)
von: Hu, Zhiming, et al.
Veröffentlicht: (2024)
HOIGaze: Gaze Estimation During Hand-Object Interactions in Extended Reality Exploiting Eye-Hand-Head Coordination
von: Hu, Zhiming, et al.
Veröffentlicht: (2025)
von: Hu, Zhiming, et al.
Veröffentlicht: (2025)
GazeMoDiff: Gaze-guided Diffusion Model for Stochastic Human Motion Prediction
von: Yan, Haodong, et al.
Veröffentlicht: (2023)
von: Yan, Haodong, et al.
Veröffentlicht: (2023)
Pose2Gaze: Eye-body Coordination during Daily Activities for Gaze Prediction from Full-body Poses
von: Hu, Zhiming, et al.
Veröffentlicht: (2023)
von: Hu, Zhiming, et al.
Veröffentlicht: (2023)
Resolving Spatio-Temporal Entanglement in Video Prediction via Multi-Modal Attention
von: Gupta, Shreyam, et al.
Veröffentlicht: (2025)
von: Gupta, Shreyam, et al.
Veröffentlicht: (2025)
HAGI++: Head-Assisted Gaze Imputation and Generation
von: Jiao, Chuhan, et al.
Veröffentlicht: (2025)
von: Jiao, Chuhan, et al.
Veröffentlicht: (2025)
PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding?
von: Siam, Mennatullah
Veröffentlicht: (2025)
von: Siam, Mennatullah
Veröffentlicht: (2025)
Robust Message Embedding via Attention Flow-Based Steganography
von: Ye, Huayuan, et al.
Veröffentlicht: (2024)
von: Ye, Huayuan, et al.
Veröffentlicht: (2024)
VSA4VQA: Scaling a Vector Symbolic Architecture to Visual Question Answering on Natural Images
von: Penzkofer, Anna, et al.
Veröffentlicht: (2024)
von: Penzkofer, Anna, et al.
Veröffentlicht: (2024)
MLVTG: Mamba-Based Feature Alignment and LLM-Driven Purification for Multi-Modal Video Temporal Grounding
von: Zhu, Zhiyi, et al.
Veröffentlicht: (2025)
von: Zhu, Zhiyi, et al.
Veröffentlicht: (2025)
UniMoCo: Unified Modality Completion for Robust Multi-Modal Embeddings
von: Qin, Jiajun, et al.
Veröffentlicht: (2025)
von: Qin, Jiajun, et al.
Veröffentlicht: (2025)
MEVDT: Multi-Modal Event-Based Vehicle Detection and Tracking Dataset
von: Shair, Zaid A. El, et al.
Veröffentlicht: (2024)
von: Shair, Zaid A. El, et al.
Veröffentlicht: (2024)
Attention-Aware Multi-View Pedestrian Tracking
von: Alturki, Reef, et al.
Veröffentlicht: (2025)
von: Alturki, Reef, et al.
Veröffentlicht: (2025)
Multi-sentence Video Grounding for Long Video Generation
von: Feng, Wei, et al.
Veröffentlicht: (2024)
von: Feng, Wei, et al.
Veröffentlicht: (2024)
RegTrack: Simplicity Beneath Complexity in Robust Multi-Modal 3D Multi-Object Tracking
von: Gu, Lipeng, et al.
Veröffentlicht: (2024)
von: Gu, Lipeng, et al.
Veröffentlicht: (2024)
Tell Me Without Telling Me: Two-Way Prediction of Visualization Literacy and Visual Attention
von: Chang, Minsuk, et al.
Veröffentlicht: (2025)
von: Chang, Minsuk, et al.
Veröffentlicht: (2025)
Multi-State Tracker: Enhancing Efficient Object Tracking via Multi-State Specialization and Interaction
von: Wang, Shilei, et al.
Veröffentlicht: (2025)
von: Wang, Shilei, et al.
Veröffentlicht: (2025)
Exploring Modality-Aware Fusion and Decoupled Temporal Propagation for Multi-Modal Object Tracking
von: Wang, Shilei, et al.
Veröffentlicht: (2026)
von: Wang, Shilei, et al.
Veröffentlicht: (2026)
MultiCOIN: Multi-Modal COntrollable Video INbetweening
von: Tanveer, Maham, et al.
Veröffentlicht: (2025)
von: Tanveer, Maham, et al.
Veröffentlicht: (2025)
ObjectVisA-120: Object-based Visual Attention Prediction in Interactive Street-crossing Environments
von: Vozniak, Igor, et al.
Veröffentlicht: (2026)
von: Vozniak, Igor, et al.
Veröffentlicht: (2026)
Vision-Motion-Reference Alignment for Referring Multi-Object Tracking via Multi-Modal Large Language Models
von: Lv, Weiyi, et al.
Veröffentlicht: (2025)
von: Lv, Weiyi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Multi-Modal Video Dialog State Tracking in the Wild
von: Abdessaied, Adnen, et al.
Veröffentlicht: (2024) -
V$^2$Dial: Unification of Video and Visual Dialog via Multimodal Experts
von: Abdessaied, Adnen, et al.
Veröffentlicht: (2025) -
ActionDiffusion: An Action-aware Diffusion Model for Procedure Planning in Instructional Videos
von: Shi, Lei, et al.
Veröffentlicht: (2024) -
CLAD: Constrained Latent Action Diffusion for Vision-Language Procedure Planning
von: Shi, Lei, et al.
Veröffentlicht: (2025) -
Learning User Embeddings from Human Gaze for Personalised Saliency Prediction
von: Strohm, Florian, et al.
Veröffentlicht: (2024)