VDMA: Video Question Answering with Dynamically Generated Multi-Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kugo, Noriyuki, Ishibashi, Tatsuya, Ono, Kosuke, Sato, Yuji |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025
von: Kamoto, Umihiro, et al.
Veröffentlicht: (2025)
von: Kamoto, Umihiro, et al.
Veröffentlicht: (2025)
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering
von: Kugo, Noriyuki, et al.
Veröffentlicht: (2025)
von: Kugo, Noriyuki, et al.
Veröffentlicht: (2025)
Reflective Dialogue between Teacher and Solver Agents for Video Question Answering
von: Murakawa, Takuya, et al.
Veröffentlicht: (2026)
von: Murakawa, Takuya, et al.
Veröffentlicht: (2026)
Foundational Question Generation for Video Question Answering via an Embedding-Integrated Approach
von: Oh, Ju-Young
Veröffentlicht: (2025)
von: Oh, Ju-Young
Veröffentlicht: (2025)
Glance and Focus: Memory Prompting for Multi-Event Video Question Answering
von: Bai, Ziyi, et al.
Veröffentlicht: (2024)
von: Bai, Ziyi, et al.
Veröffentlicht: (2024)
FIQ: Fundamental Question Generation with the Integration of Question Embeddings for Video Question Answering
von: Oh, Ju-Young, et al.
Veröffentlicht: (2025)
von: Oh, Ju-Young, et al.
Veröffentlicht: (2025)
Static or Dynamic: Towards Query-Adaptive Token Selection for Video Question Answering
von: Shi, Yumeng, et al.
Veröffentlicht: (2025)
von: Shi, Yumeng, et al.
Veröffentlicht: (2025)
Grounded Question-Answering in Long Egocentric Videos
von: Di, Shangzhe, et al.
Veröffentlicht: (2023)
von: Di, Shangzhe, et al.
Veröffentlicht: (2023)
Agentic Keyframe Search for Video Question Answering
von: Fan, Sunqi, et al.
Veröffentlicht: (2025)
von: Fan, Sunqi, et al.
Veröffentlicht: (2025)
ChineseVideoBench: Benchmarking Multi-modal Large Models for Chinese Video Question Answering
von: Nie, Yuxiang, et al.
Veröffentlicht: (2025)
von: Nie, Yuxiang, et al.
Veröffentlicht: (2025)
Question-Instructed Visual Descriptions for Zero-Shot Video Question Answering
von: Romero, David, et al.
Veröffentlicht: (2024)
von: Romero, David, et al.
Veröffentlicht: (2024)
UDVideoQA: A Traffic Video Question Answering Dataset for Multi-Object Spatio-Temporal Reasoning in Urban Dynamics
von: Vishal, Joseph Raj, et al.
Veröffentlicht: (2026)
von: Vishal, Joseph Raj, et al.
Veröffentlicht: (2026)
Narrative Aligned Long Form Video Question Answering
von: Jain, Rahul, et al.
Veröffentlicht: (2026)
von: Jain, Rahul, et al.
Veröffentlicht: (2026)
VideoDistill: Language-aware Vision Distillation for Video Question Answering
von: Zou, Bo, et al.
Veröffentlicht: (2024)
von: Zou, Bo, et al.
Veröffentlicht: (2024)
MovieChat+: Question-aware Sparse Memory for Long Video Question Answering
von: Song, Enxin, et al.
Veröffentlicht: (2024)
von: Song, Enxin, et al.
Veröffentlicht: (2024)
ViLA: Efficient Video-Language Alignment for Video Question Answering
von: Wang, Xijun, et al.
Veröffentlicht: (2023)
von: Wang, Xijun, et al.
Veröffentlicht: (2023)
REVEAL: Relation-based Video Representation Learning for Video-Question-Answering
von: Chaybouti, Sofian, et al.
Veröffentlicht: (2025)
von: Chaybouti, Sofian, et al.
Veröffentlicht: (2025)
Streaming Video Question-Answering with In-context Video KV-Cache Retrieval
von: Di, Shangzhe, et al.
Veröffentlicht: (2025)
von: Di, Shangzhe, et al.
Veröffentlicht: (2025)
Towards Top-Down Reasoning: An Explainable Multi-Agent Approach for Visual Question Answering
von: Wang, Zeqing, et al.
Veröffentlicht: (2023)
von: Wang, Zeqing, et al.
Veröffentlicht: (2023)
Answering from Sure to Uncertain: Uncertainty-Aware Curriculum Learning for Video Question Answering
von: Li, Haopeng, et al.
Veröffentlicht: (2024)
von: Li, Haopeng, et al.
Veröffentlicht: (2024)
ETVA: Evaluation of Text-to-Video Alignment via Fine-grained Question Generation and Answering
von: Guan, Kaisi, et al.
Veröffentlicht: (2025)
von: Guan, Kaisi, et al.
Veröffentlicht: (2025)
CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering
von: Zhang, Mingfang, et al.
Veröffentlicht: (2026)
von: Zhang, Mingfang, et al.
Veröffentlicht: (2026)
Neuro Symbolic Knowledge Reasoning for Procedural Video Question Answering
von: Fernando, Basura, et al.
Veröffentlicht: (2025)
von: Fernando, Basura, et al.
Veröffentlicht: (2025)
Object-centric Video Question Answering with Visual Grounding and Referring
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
Admitting Ignorance Helps the Video Question Answering Models to Answer
von: Li, Haopeng, et al.
Veröffentlicht: (2025)
von: Li, Haopeng, et al.
Veröffentlicht: (2025)
Marten: Visual Question Answering with Mask Generation for Multi-modal Document Understanding
von: Wang, Zining, et al.
Veröffentlicht: (2025)
von: Wang, Zining, et al.
Veröffentlicht: (2025)
Multi-Sourced Compositional Generalization in Visual Question Answering
von: Li, Chuanhao, et al.
Veröffentlicht: (2025)
von: Li, Chuanhao, et al.
Veröffentlicht: (2025)
DEJIMA: A Novel Large-scale Japanese Dataset for Image Captioning and Visual Question Answering
von: Katsube, Toshiki, et al.
Veröffentlicht: (2025)
von: Katsube, Toshiki, et al.
Veröffentlicht: (2025)
LvBench: A Benchmark for Long-form Video Understanding with Versatile Multi-modal Question Answering
von: Zhang, Hongjie, et al.
Veröffentlicht: (2023)
von: Zhang, Hongjie, et al.
Veröffentlicht: (2023)
Align and Aggregate: Compositional Reasoning with Video Alignment and Answer Aggregation for Video Question-Answering
von: Liao, Zhaohe, et al.
Veröffentlicht: (2024)
von: Liao, Zhaohe, et al.
Veröffentlicht: (2024)
TraveLER: A Modular Multi-LMM Agent Framework for Video Question-Answering
von: Shang, Chuyi, et al.
Veröffentlicht: (2024)
von: Shang, Chuyi, et al.
Veröffentlicht: (2024)
Ranking Distillation for Open-Ended Video Question Answering with Insufficient Labels
von: Liang, Tianming, et al.
Veröffentlicht: (2024)
von: Liang, Tianming, et al.
Veröffentlicht: (2024)
STRIVE: Structured Spatiotemporal Exploration for Reinforcement Learning in Video Question Answering
von: Bahrami, Emad, et al.
Veröffentlicht: (2026)
von: Bahrami, Emad, et al.
Veröffentlicht: (2026)
Tool-Augmented Spatiotemporal Reasoning for Streamlining Video Question Answering Task
von: Fan, Sunqi, et al.
Veröffentlicht: (2025)
von: Fan, Sunqi, et al.
Veröffentlicht: (2025)
A Simple LLM Framework for Long-Range Video Question-Answering
von: Zhang, Ce, et al.
Veröffentlicht: (2023)
von: Zhang, Ce, et al.
Veröffentlicht: (2023)
SurgLQA: Scalable Long-Horizon Surgical Video Question Answering
von: Guo, Diandian, et al.
Veröffentlicht: (2026)
von: Guo, Diandian, et al.
Veröffentlicht: (2026)
BIMBA: Selective-Scan Compression for Long-Range Video Question Answering
von: Islam, Md Mohaiminul, et al.
Veröffentlicht: (2025)
von: Islam, Md Mohaiminul, et al.
Veröffentlicht: (2025)
Multi-object event graph representation learning for Video Question Answering
von: Wang, Yanan, et al.
Veröffentlicht: (2024)
von: Wang, Yanan, et al.
Veröffentlicht: (2024)
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering
von: Dang, Jisheng, et al.
Veröffentlicht: (2025)
von: Dang, Jisheng, et al.
Veröffentlicht: (2025)
Question-Answering Dense Video Events
von: Qin, Hangyu, et al.
Veröffentlicht: (2024)
von: Qin, Hangyu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025
von: Kamoto, Umihiro, et al.
Veröffentlicht: (2025) -
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering
von: Kugo, Noriyuki, et al.
Veröffentlicht: (2025) -
Reflective Dialogue between Teacher and Solver Agents for Video Question Answering
von: Murakawa, Takuya, et al.
Veröffentlicht: (2026) -
Foundational Question Generation for Video Question Answering via an Embedding-Integrated Approach
von: Oh, Ju-Young
Veröffentlicht: (2025) -
Glance and Focus: Memory Prompting for Multi-Event Video Question Answering
von: Bai, Ziyi, et al.
Veröffentlicht: (2024)