Reflective Dialogue between Teacher and Solver Agents for Video Question Answering
Fuente:
arXiv
Saved in:
| Main Authors: | Murakawa, Takuya, Tamaki, Toru |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
M3DDM+: An improved video outpainting by a modified masking strategy
by: Murakawa, Takuya, et al.
Published: (2026)
by: Murakawa, Takuya, et al.
Published: (2026)
MoExDA: Domain Adaptation for Edge-based Action Recognition
by: Sugimoto, Takuya, et al.
Published: (2025)
by: Sugimoto, Takuya, et al.
Published: (2025)
Shot2Tactic-Caption: Multi-Scale Captioning of Badminton Videos for Tactical Understanding
by: Ding, Ning, et al.
Published: (2025)
by: Ding, Ning, et al.
Published: (2025)
VDMA: Video Question Answering with Dynamically Generated Multi-Agents
by: Kugo, Noriyuki, et al.
Published: (2024)
by: Kugo, Noriyuki, et al.
Published: (2024)
Multi-model learning by sequential reading of untrimmed videos for action recognition
by: Kamiya, Kodai, et al.
Published: (2024)
by: Kamiya, Kodai, et al.
Published: (2024)
Shift and matching queries for video semantic segmentation
by: Mizuno, Tsubasa, et al.
Published: (2024)
by: Mizuno, Tsubasa, et al.
Published: (2024)
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering
by: Kugo, Noriyuki, et al.
Published: (2025)
by: Kugo, Noriyuki, et al.
Published: (2025)
Grounded Question-Answering in Long Egocentric Videos
by: Di, Shangzhe, et al.
Published: (2023)
by: Di, Shangzhe, et al.
Published: (2023)
Agentic Keyframe Search for Video Question Answering
by: Fan, Sunqi, et al.
Published: (2025)
by: Fan, Sunqi, et al.
Published: (2025)
Question-Instructed Visual Descriptions for Zero-Shot Video Question Answering
by: Romero, David, et al.
Published: (2024)
by: Romero, David, et al.
Published: (2024)
Narrative Aligned Long Form Video Question Answering
by: Jain, Rahul, et al.
Published: (2026)
by: Jain, Rahul, et al.
Published: (2026)
ViLA: Efficient Video-Language Alignment for Video Question Answering
by: Wang, Xijun, et al.
Published: (2023)
by: Wang, Xijun, et al.
Published: (2023)
REVEAL: Relation-based Video Representation Learning for Video-Question-Answering
by: Chaybouti, Sofian, et al.
Published: (2025)
by: Chaybouti, Sofian, et al.
Published: (2025)
VideoDistill: Language-aware Vision Distillation for Video Question Answering
by: Zou, Bo, et al.
Published: (2024)
by: Zou, Bo, et al.
Published: (2024)
Streaming Video Question-Answering with In-context Video KV-Cache Retrieval
by: Di, Shangzhe, et al.
Published: (2025)
by: Di, Shangzhe, et al.
Published: (2025)
Foundational Question Generation for Video Question Answering via an Embedding-Integrated Approach
by: Oh, Ju-Young
Published: (2025)
by: Oh, Ju-Young
Published: (2025)
MovieChat+: Question-aware Sparse Memory for Long Video Question Answering
by: Song, Enxin, et al.
Published: (2024)
by: Song, Enxin, et al.
Published: (2024)
BFMD: A Full-Match Badminton Dense Dataset for Dense Shot Captioning
by: Ding, Ning, et al.
Published: (2026)
by: Ding, Ning, et al.
Published: (2026)
Online pre-training with long-form videos
by: Kato, Itsuki, et al.
Published: (2024)
by: Kato, Itsuki, et al.
Published: (2024)
Query matching for spatio-temporal action detection with query-based object detector
by: Hori, Shimon, et al.
Published: (2024)
by: Hori, Shimon, et al.
Published: (2024)
Disentangling Static and Dynamic Information for Reducing Static Bias in Action Recognition
by: Kobayashi, Masato, et al.
Published: (2025)
by: Kobayashi, Masato, et al.
Published: (2025)
Action tube generation by person query matching for spatio-temporal action detection
by: Omi, Kazuki, et al.
Published: (2025)
by: Omi, Kazuki, et al.
Published: (2025)
Fine-grained length controllable video captioning with ordinal embeddings
by: Nitta, Tomoya, et al.
Published: (2024)
by: Nitta, Tomoya, et al.
Published: (2024)
Answering from Sure to Uncertain: Uncertainty-Aware Curriculum Learning for Video Question Answering
by: Li, Haopeng, et al.
Published: (2024)
by: Li, Haopeng, et al.
Published: (2024)
Neuro Symbolic Knowledge Reasoning for Procedural Video Question Answering
by: Fernando, Basura, et al.
Published: (2025)
by: Fernando, Basura, et al.
Published: (2025)
Object-centric Video Question Answering with Visual Grounding and Referring
by: Wang, Haochen, et al.
Published: (2025)
by: Wang, Haochen, et al.
Published: (2025)
Admitting Ignorance Helps the Video Question Answering Models to Answer
by: Li, Haopeng, et al.
Published: (2025)
by: Li, Haopeng, et al.
Published: (2025)
Align and Aggregate: Compositional Reasoning with Video Alignment and Answer Aggregation for Video Question-Answering
by: Liao, Zhaohe, et al.
Published: (2024)
by: Liao, Zhaohe, et al.
Published: (2024)
FIQ: Fundamental Question Generation with the Integration of Question Embeddings for Video Question Answering
by: Oh, Ju-Young, et al.
Published: (2025)
by: Oh, Ju-Young, et al.
Published: (2025)
STRIVE: Structured Spatiotemporal Exploration for Reinforcement Learning in Video Question Answering
by: Bahrami, Emad, et al.
Published: (2026)
by: Bahrami, Emad, et al.
Published: (2026)
SurgLQA: Scalable Long-Horizon Surgical Video Question Answering
by: Guo, Diandian, et al.
Published: (2026)
by: Guo, Diandian, et al.
Published: (2026)
Ranking Distillation for Open-Ended Video Question Answering with Insufficient Labels
by: Liang, Tianming, et al.
Published: (2024)
by: Liang, Tianming, et al.
Published: (2024)
Tool-Augmented Spatiotemporal Reasoning for Streamlining Video Question Answering Task
by: Fan, Sunqi, et al.
Published: (2025)
by: Fan, Sunqi, et al.
Published: (2025)
A Simple LLM Framework for Long-Range Video Question-Answering
by: Zhang, Ce, et al.
Published: (2023)
by: Zhang, Ce, et al.
Published: (2023)
BIMBA: Selective-Scan Compression for Long-Range Video Question Answering
by: Islam, Md Mohaiminul, et al.
Published: (2025)
by: Islam, Md Mohaiminul, et al.
Published: (2025)
Glance and Focus: Memory Prompting for Multi-Event Video Question Answering
by: Bai, Ziyi, et al.
Published: (2024)
by: Bai, Ziyi, et al.
Published: (2024)
Question-Answering Dense Video Events
by: Qin, Hangyu, et al.
Published: (2024)
by: Qin, Hangyu, et al.
Published: (2024)
ORCA: Orchestrated Reasoning with Collaborative Agents for Document Visual Question Answering
by: Lassoued, Aymen, et al.
Published: (2026)
by: Lassoued, Aymen, et al.
Published: (2026)
Actions and Objects Pathways for Domain Adaptation in Video Question Answering
by: Mohamud, Safaa Abdullahi Moallim, et al.
Published: (2024)
by: Mohamud, Safaa Abdullahi Moallim, et al.
Published: (2024)
ChineseVideoBench: Benchmarking Multi-modal Large Models for Chinese Video Question Answering
by: Nie, Yuxiang, et al.
Published: (2025)
by: Nie, Yuxiang, et al.
Published: (2025)
Similar Items
-
M3DDM+: An improved video outpainting by a modified masking strategy
by: Murakawa, Takuya, et al.
Published: (2026) -
MoExDA: Domain Adaptation for Edge-based Action Recognition
by: Sugimoto, Takuya, et al.
Published: (2025) -
Shot2Tactic-Caption: Multi-Scale Captioning of Badminton Videos for Tactical Understanding
by: Ding, Ning, et al.
Published: (2025) -
VDMA: Video Question Answering with Dynamically Generated Multi-Agents
by: Kugo, Noriyuki, et al.
Published: (2024) -
Multi-model learning by sequential reading of untrimmed videos for action recognition
by: Kamiya, Kodai, et al.
Published: (2024)