V$^2$Dial: Unification of Video and Visual Dialog via Multimodal Experts
Fuente:
arXiv
Salvato in:
| Autori principali: | Abdessaied, Adnen, Rohrbach, Anna, Rohrbach, Marcus, Bulling, Andreas |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Multi-Modal Video Dialog State Tracking in the Wild
di: Abdessaied, Adnen, et al.
Pubblicazione: (2024)
di: Abdessaied, Adnen, et al.
Pubblicazione: (2024)
OLViT: Multi-Modal State Tracking via Attention-Based Embeddings for Video-Grounded Dialog
di: Abdessaied, Adnen, et al.
Pubblicazione: (2024)
di: Abdessaied, Adnen, et al.
Pubblicazione: (2024)
DEFAME: Dynamic Evidence-based FAct-checking with Multimodal Experts
di: Braun, Tobias, et al.
Pubblicazione: (2024)
di: Braun, Tobias, et al.
Pubblicazione: (2024)
VeriTaS: The First Dynamic Benchmark for Multimodal Automated Fact-Checking
di: Rothermel, Mark, et al.
Pubblicazione: (2026)
di: Rothermel, Mark, et al.
Pubblicazione: (2026)
Chrono: A Simple Blueprint for Representing Time in MLLMs
di: Rodriguez, Hector, et al.
Pubblicazione: (2024)
di: Rodriguez, Hector, et al.
Pubblicazione: (2024)
SIEVES: Selective Prediction Generalizes through Visual Evidence Scoring
di: Rodriguez, Hector G., et al.
Pubblicazione: (2026)
di: Rodriguez, Hector G., et al.
Pubblicazione: (2026)
ReCap: Lightweight Referential Grounding for Coherent Story Visualization
di: Arora, Aditya, et al.
Pubblicazione: (2026)
di: Arora, Aditya, et al.
Pubblicazione: (2026)
HaloProbe: Bayesian Detection and Mitigation of Object Hallucinations in Vision-Language Models
di: Zohrabi, Reihaneh, et al.
Pubblicazione: (2026)
di: Zohrabi, Reihaneh, et al.
Pubblicazione: (2026)
Spurious-Aware Prototype Refinement for Reliable Out-of-Distribution Detection
di: Zohrabi, Reihaneh, et al.
Pubblicazione: (2025)
di: Zohrabi, Reihaneh, et al.
Pubblicazione: (2025)
Predicting Implicit Arguments in Procedural Video Instructions
di: Batra, Anil, et al.
Pubblicazione: (2025)
di: Batra, Anil, et al.
Pubblicazione: (2025)
Variational Visual Question Answering for Uncertainty-Aware Selective Prediction
di: Wieczorek, Tobias Jan, et al.
Pubblicazione: (2025)
di: Wieczorek, Tobias Jan, et al.
Pubblicazione: (2025)
Diffusion Classifiers Understand Compositionality, but Conditions Apply
di: Jeong, Yujin, et al.
Pubblicazione: (2025)
di: Jeong, Yujin, et al.
Pubblicazione: (2025)
Tuning Just Enough: Lightweight Backdoor Attacks on Multi-Encoder Diffusion Models
di: Chen, Ziyuan, et al.
Pubblicazione: (2026)
di: Chen, Ziyuan, et al.
Pubblicazione: (2026)
Evaluating the Impact of Post-Training Quantization on Reliable VQA with Multimodal LLMs
di: Kurz, Paul Jonas, et al.
Pubblicazione: (2026)
di: Kurz, Paul Jonas, et al.
Pubblicazione: (2026)
Efficient Pre-training for Localized Instruction Generation of Videos
di: Batra, Anil, et al.
Pubblicazione: (2023)
di: Batra, Anil, et al.
Pubblicazione: (2023)
DialNav: Multi-turn Dialog Navigation with a Remote Guide
di: Han, Leekyeung, et al.
Pubblicazione: (2025)
di: Han, Leekyeung, et al.
Pubblicazione: (2025)
When Do Diffusion Models learn to Generate Multiple Objects?
di: Jeong, Yujin, et al.
Pubblicazione: (2026)
di: Jeong, Yujin, et al.
Pubblicazione: (2026)
VSA4VQA: Scaling a Vector Symbolic Architecture to Visual Question Answering on Natural Images
di: Penzkofer, Anna, et al.
Pubblicazione: (2024)
di: Penzkofer, Anna, et al.
Pubblicazione: (2024)
ActionDiffusion: An Action-aware Diffusion Model for Procedure Planning in Instructional Videos
di: Shi, Lei, et al.
Pubblicazione: (2024)
di: Shi, Lei, et al.
Pubblicazione: (2024)
VQA-MHUG: A Gaze Dataset to Study Multimodal Neural Attention in Visual Question Answering
di: Sood, Ekta, et al.
Pubblicazione: (2021)
di: Sood, Ekta, et al.
Pubblicazione: (2021)
Multimodal Integration of Human-Like Attention in Visual Question Answering
di: Sood, Ekta, et al.
Pubblicazione: (2021)
di: Sood, Ekta, et al.
Pubblicazione: (2021)
Multi-axis Analysis of Image Manipulation Localization
di: Nichols, Keanu, et al.
Pubblicazione: (2026)
di: Nichols, Keanu, et al.
Pubblicazione: (2026)
CLAD: Constrained Latent Action Diffusion for Vision-Language Procedure Planning
di: Shi, Lei, et al.
Pubblicazione: (2025)
di: Shi, Lei, et al.
Pubblicazione: (2025)
Ontology-Guided Diffusion for Zero-Shot Visual Sim2Real Transfer
di: Youssef, Mohamed, et al.
Pubblicazione: (2026)
di: Youssef, Mohamed, et al.
Pubblicazione: (2026)
TokenDial: Continuous Attribute Control in Text-to-Video via Spatiotemporal Token Offsets
di: Liu, Zhixuan, et al.
Pubblicazione: (2026)
di: Liu, Zhixuan, et al.
Pubblicazione: (2026)
UP-FacE: User-predictable Fine-grained Face Shape Editing
di: Strohm, Florian, et al.
Pubblicazione: (2024)
di: Strohm, Florian, et al.
Pubblicazione: (2024)
HAIFAI: Human-AI Interaction for Mental Face Reconstruction
di: Strohm, Florian, et al.
Pubblicazione: (2024)
di: Strohm, Florian, et al.
Pubblicazione: (2024)
Pose2Gaze: Eye-body Coordination during Daily Activities for Gaze Prediction from Full-body Poses
di: Hu, Zhiming, et al.
Pubblicazione: (2023)
di: Hu, Zhiming, et al.
Pubblicazione: (2023)
Compression Tells Intelligence: Visual Coding, Visual Token Technology, and the Unification
di: Jin, Xin, et al.
Pubblicazione: (2026)
di: Jin, Xin, et al.
Pubblicazione: (2026)
HOIGaze: Gaze Estimation During Hand-Object Interactions in Extended Reality Exploiting Eye-Hand-Head Coordination
di: Hu, Zhiming, et al.
Pubblicazione: (2025)
di: Hu, Zhiming, et al.
Pubblicazione: (2025)
GazeMoDiff: Gaze-guided Diffusion Model for Stochastic Human Motion Prediction
di: Yan, Haodong, et al.
Pubblicazione: (2023)
di: Yan, Haodong, et al.
Pubblicazione: (2023)
GazeMotion: Gaze-guided Human Motion Forecasting
di: Hu, Zhiming, et al.
Pubblicazione: (2024)
di: Hu, Zhiming, et al.
Pubblicazione: (2024)
Unified Multimodal Visual Tracking with Dual Mixture-of-Experts
di: Hong, Lingyi, et al.
Pubblicazione: (2026)
di: Hong, Lingyi, et al.
Pubblicazione: (2026)
HAGI++: Head-Assisted Gaze Imputation and Generation
di: Jiao, Chuhan, et al.
Pubblicazione: (2025)
di: Jiao, Chuhan, et al.
Pubblicazione: (2025)
VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents
di: Meng, Rui, et al.
Pubblicazione: (2025)
di: Meng, Rui, et al.
Pubblicazione: (2025)
Tell Me Without Telling Me: Two-Way Prediction of Visualization Literacy and Visual Attention
di: Chang, Minsuk, et al.
Pubblicazione: (2025)
di: Chang, Minsuk, et al.
Pubblicazione: (2025)
Two Causal Principles for Improving Visual Dialog
di: Qi, Jiaxin, et al.
Pubblicazione: (2019)
di: Qi, Jiaxin, et al.
Pubblicazione: (2019)
Shape-Guided Diffusion with Inside-Outside Attention
di: Park, Dong Huk, et al.
Pubblicazione: (2022)
di: Park, Dong Huk, et al.
Pubblicazione: (2022)
Router-Suggest: Dynamic Routing for Multimodal Auto-Completion in Visually-Grounded Dialogs
di: Mishra, Sandeep, et al.
Pubblicazione: (2026)
di: Mishra, Sandeep, et al.
Pubblicazione: (2026)
Dynamic-I2V: Exploring Image-to-Video Generation Models via Multimodal LLM
di: Liu, Peng, et al.
Pubblicazione: (2025)
di: Liu, Peng, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Multi-Modal Video Dialog State Tracking in the Wild
di: Abdessaied, Adnen, et al.
Pubblicazione: (2024) -
OLViT: Multi-Modal State Tracking via Attention-Based Embeddings for Video-Grounded Dialog
di: Abdessaied, Adnen, et al.
Pubblicazione: (2024) -
DEFAME: Dynamic Evidence-based FAct-checking with Multimodal Experts
di: Braun, Tobias, et al.
Pubblicazione: (2024) -
VeriTaS: The First Dynamic Benchmark for Multimodal Automated Fact-Checking
di: Rothermel, Mark, et al.
Pubblicazione: (2026) -
Chrono: A Simple Blueprint for Representing Time in MLLMs
di: Rodriguez, Hector, et al.
Pubblicazione: (2024)