MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dang, Jisheng, Song, Huilin, Xiao, Junbin, Wang, Bimei, Peng, Han, Li, Haoxuan, Yang, Xun, Wang, Meng, Chua, Tat-Seng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scene-Text Grounding for Text-Based Video Question Answering
von: Zhou, Sheng, et al.
Veröffentlicht: (2024)
von: Zhou, Sheng, et al.
Veröffentlicht: (2024)
Can I Trust Your Answer? Visually Grounded Video Question Answering
von: Xiao, Junbin, et al.
Veröffentlicht: (2023)
von: Xiao, Junbin, et al.
Veröffentlicht: (2023)
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering
von: Zhou, Sheng, et al.
Veröffentlicht: (2025)
von: Zhou, Sheng, et al.
Veröffentlicht: (2025)
Reinforcing Video Reasoning with Focused Thinking
von: Dang, Jisheng, et al.
Veröffentlicht: (2025)
von: Dang, Jisheng, et al.
Veröffentlicht: (2025)
Ego-Grounding for Personalized Question-Answering in Egocentric Videos
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
On the Implicit Reward Overfitting and the Low-rank Dynamics in RLVR
von: Ye, Hao, et al.
Veröffentlicht: (2026)
von: Ye, Hao, et al.
Veröffentlicht: (2026)
Question-Answering Dense Video Events
von: Qin, Hangyu, et al.
Veröffentlicht: (2024)
von: Qin, Hangyu, et al.
Veröffentlicht: (2024)
Show Me What and Where has Changed? Question Answering and Grounding for Remote Sensing Change Detection
von: Li, Ke, et al.
Veröffentlicht: (2024)
von: Li, Ke, et al.
Veröffentlicht: (2024)
MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
Decoupled Seg Tokens Make Stronger Reasoning Video Segmenter and Grounder
von: Jisheng, Dang, et al.
Veröffentlicht: (2025)
von: Jisheng, Dang, et al.
Veröffentlicht: (2025)
Understanding Long Videos via LLM-Powered Entity Relation Graphs
von: Chu, Meng, et al.
Veröffentlicht: (2025)
von: Chu, Meng, et al.
Veröffentlicht: (2025)
Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model
von: Shen, Fei, et al.
Veröffentlicht: (2025)
von: Shen, Fei, et al.
Veröffentlicht: (2025)
Learning to Ask Critical Questions for Assisting Product Search
von: Li, Zixuan, et al.
Veröffentlicht: (2024)
von: Li, Zixuan, et al.
Veröffentlicht: (2024)
Agentic Keyframe Search for Video Question Answering
von: Fan, Sunqi, et al.
Veröffentlicht: (2025)
von: Fan, Sunqi, et al.
Veröffentlicht: (2025)
Dynamic Multimodal Fusion via Meta-Learning Towards Micro-Video Recommendation
von: Liu, Han, et al.
Veröffentlicht: (2025)
von: Liu, Han, et al.
Veröffentlicht: (2025)
Disentangling Masked Autoencoders for Unsupervised Domain Generalization
von: Zhang, An, et al.
Veröffentlicht: (2024)
von: Zhang, An, et al.
Veröffentlicht: (2024)
PathReasoner: Modeling Reasoning Path with Equivalent Extension for Logical Question Answering
von: Xu, Fangzhi, et al.
Veröffentlicht: (2024)
von: Xu, Fangzhi, et al.
Veröffentlicht: (2024)
Towards Natural Language-Guided Drones: GeoText-1652 Benchmark with Spatial Relation Matching
von: Chu, Meng, et al.
Veröffentlicht: (2023)
von: Chu, Meng, et al.
Veröffentlicht: (2023)
Causal-Entity Reflected Egocentric Traffic Accident Video Synthesis
von: Li, Lei-lei, et al.
Veröffentlicht: (2025)
von: Li, Lei-lei, et al.
Veröffentlicht: (2025)
A Survey on Neural Question Generation: Methods, Applications, and Prospects
von: Guo, Shasha, et al.
Veröffentlicht: (2024)
von: Guo, Shasha, et al.
Veröffentlicht: (2024)
Abductive Ego-View Accident Video Understanding for Safe Driving Perception
von: Fang, Jianwu, et al.
Veröffentlicht: (2024)
von: Fang, Jianwu, et al.
Veröffentlicht: (2024)
Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning
von: Liu, Huabin, et al.
Veröffentlicht: (2025)
von: Liu, Huabin, et al.
Veröffentlicht: (2025)
Compose Your Aesthetics: Empowering Text-to-Image Models with the Principles of Art
von: Jin, Zhe, et al.
Veröffentlicht: (2025)
von: Jin, Zhe, et al.
Veröffentlicht: (2025)
EPERM: An Evidence Path Enhanced Reasoning Model for Knowledge Graph Question and Answering
von: Long, Xiao, et al.
Veröffentlicht: (2025)
von: Long, Xiao, et al.
Veröffentlicht: (2025)
Towards Goal-oriented Intelligent Tutoring Systems in Online Education
von: Deng, Yang, et al.
Veröffentlicht: (2023)
von: Deng, Yang, et al.
Veröffentlicht: (2023)
Conformal Path Reasoning: Trustworthy Knowledge Graph Question Answering via Path-Level Calibration
von: Lin, Shuhang, et al.
Veröffentlicht: (2026)
von: Lin, Shuhang, et al.
Veröffentlicht: (2026)
VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering
von: Meng, Yiran, et al.
Veröffentlicht: (2025)
von: Meng, Yiran, et al.
Veröffentlicht: (2025)
Extending Visual Dynamics for Video-to-Music Generation
von: Liu, Xiaohao, et al.
Veröffentlicht: (2025)
von: Liu, Xiaohao, et al.
Veröffentlicht: (2025)
Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning
von: Liu, Chengwen, et al.
Veröffentlicht: (2026)
von: Liu, Chengwen, et al.
Veröffentlicht: (2026)
CARL: Criticality-Aware Agentic Reinforcement Learning
von: Shen, Leyang, et al.
Veröffentlicht: (2025)
von: Shen, Leyang, et al.
Veröffentlicht: (2025)
Generative Recommendation: Towards Next-generation Recommender Paradigm
von: Wang, Wenjie, et al.
Veröffentlicht: (2023)
von: Wang, Wenjie, et al.
Veröffentlicht: (2023)
Grounded Question-Answering in Long Egocentric Videos
von: Di, Shangzhe, et al.
Veröffentlicht: (2023)
von: Di, Shangzhe, et al.
Veröffentlicht: (2023)
DRAFT: Task Decoupled Latent Reasoning for Agent Safety
von: Wang, Lin, et al.
Veröffentlicht: (2026)
von: Wang, Lin, et al.
Veröffentlicht: (2026)
Reinforcing Chain-of-Thought Reasoning with Self-Evolving Rubrics
von: Sheng, Leheng, et al.
Veröffentlicht: (2026)
von: Sheng, Leheng, et al.
Veröffentlicht: (2026)
Object-centric Video Question Answering with Visual Grounding and Referring
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
GCoT-Decoding: Unlocking Deep Reasoning Paths for Universal Question Answering
von: Luo, Guanran, et al.
Veröffentlicht: (2026)
von: Luo, Guanran, et al.
Veröffentlicht: (2026)
Bridging Vision Language Models and Symbolic Grounding for Video Question Answering
von: Ma, Haodi, et al.
Veröffentlicht: (2025)
von: Ma, Haodi, et al.
Veröffentlicht: (2025)
LeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question Answering
von: Dong, Xinxin, et al.
Veröffentlicht: (2025)
von: Dong, Xinxin, et al.
Veröffentlicht: (2025)
TTOM: Test-Time Optimization and Memorization for Compositional Video Generation
von: Qu, Leigang, et al.
Veröffentlicht: (2025)
von: Qu, Leigang, et al.
Veröffentlicht: (2025)
Turing Patterns for Multimedia: Reaction-Diffusion Multi-Modal Fusion for Language-Guided Video Moment Retrieval
von: Fang, Xiang, et al.
Veröffentlicht: (2026)
von: Fang, Xiang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Scene-Text Grounding for Text-Based Video Question Answering
von: Zhou, Sheng, et al.
Veröffentlicht: (2024) -
Can I Trust Your Answer? Visually Grounded Video Question Answering
von: Xiao, Junbin, et al.
Veröffentlicht: (2023) -
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering
von: Zhou, Sheng, et al.
Veröffentlicht: (2025) -
Reinforcing Video Reasoning with Focused Thinking
von: Dang, Jisheng, et al.
Veröffentlicht: (2025) -
Ego-Grounding for Personalized Question-Answering in Egocentric Videos
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)