Video-in-the-Loop: Span-Grounded Long Video QA with Interleaved Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Chendong, Bai, Donglin, Yang, Yifan, Jin, Xiao, Zhang, Anlan, Wang, Rui, Jiang, Shiqi, Yang, Yuqing, Wu, Hao, Dai, Qi, Luo, Chong, Cao, Ting, Qiu, Lili, Banerjee, Suman |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VoLUT: Efficient Volumetric streaming enhanced by LUT-based super-resolution
von: Wang, Chendong, et al.
Veröffentlicht: (2025)
von: Wang, Chendong, et al.
Veröffentlicht: (2025)
AVA: Towards Agentic Video Analytics with Vision Language Models
von: Yan, Yuxuan, et al.
Veröffentlicht: (2025)
von: Yan, Yuxuan, et al.
Veröffentlicht: (2025)
Babel: A Scalable Pre-trained Model for Multi-Modal Sensing via Expandable Modality Alignment
von: Dai, Shenghong, et al.
Veröffentlicht: (2024)
von: Dai, Shenghong, et al.
Veröffentlicht: (2024)
AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation
von: Zhou, Ziwei, et al.
Veröffentlicht: (2026)
von: Zhou, Ziwei, et al.
Veröffentlicht: (2026)
StreamMind: Unlocking Full Frame Rate Streaming Video Dialogue through Event-Gated Cognition
von: Ding, Xin, et al.
Veröffentlicht: (2025)
von: Ding, Xin, et al.
Veröffentlicht: (2025)
ProRe: A Proactive Reward System for GUI Agents via Reasoner-Actor Collaboration
von: Dai, Gaole, et al.
Veröffentlicht: (2025)
von: Dai, Gaole, et al.
Veröffentlicht: (2025)
Zoomer: Adaptive Image Focus Optimization for Black-box MLLM
von: Qian, Jiaxu, et al.
Veröffentlicht: (2025)
von: Qian, Jiaxu, et al.
Veröffentlicht: (2025)
Making Every Frame Matter: Continuous Activity Recognition in Streaming Video via Adaptive Video Context Modeling
von: Wu, Hao, et al.
Veröffentlicht: (2024)
von: Wu, Hao, et al.
Veröffentlicht: (2024)
HiTVideo: Hierarchical Tokenizers for Enhancing Text-to-Video Generation with Autoregressive Large Language Models
von: Zhou, Ziqin, et al.
Veröffentlicht: (2025)
von: Zhou, Ziqin, et al.
Veröffentlicht: (2025)
Neural-Symbolic VideoQA: Learning Compositional Spatio-Temporal Reasoning for Real-world Video Question Answering
von: Liang, Lili, et al.
Veröffentlicht: (2024)
von: Liang, Lili, et al.
Veröffentlicht: (2024)
Hierarchical Memory for Long Video QA
von: Wang, Yiqin, et al.
Veröffentlicht: (2024)
von: Wang, Yiqin, et al.
Veröffentlicht: (2024)
Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning
von: Zeng, Xiangyu, et al.
Veröffentlicht: (2026)
von: Zeng, Xiangyu, et al.
Veröffentlicht: (2026)
ReasonGen-R1: CoT for Autoregressive Image generation models through SFT and RL
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos
von: Chen, Qirui, et al.
Veröffentlicht: (2024)
von: Chen, Qirui, et al.
Veröffentlicht: (2024)
Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding
von: Chen, Houlun, et al.
Veröffentlicht: (2026)
von: Chen, Houlun, et al.
Veröffentlicht: (2026)
NeuS-QA: Grounding Long-Form Video Understanding in Temporal Logic and Neuro-Symbolic Reasoning
von: Shah, Sahil, et al.
Veröffentlicht: (2025)
von: Shah, Sahil, et al.
Veröffentlicht: (2025)
Match Stereo Videos via Bidirectional Alignment
von: Jing, Junpeng, et al.
Veröffentlicht: (2024)
von: Jing, Junpeng, et al.
Veröffentlicht: (2024)
Advancing Mobile GUI Agents: A Verifier-Driven Approach to Practical Deployment
von: Dai, Gaole, et al.
Veröffentlicht: (2025)
von: Dai, Gaole, et al.
Veröffentlicht: (2025)
WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning
von: Zhang, Yuanhan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuanhan, et al.
Veröffentlicht: (2024)
CI-VID: A Coherent Interleaved Text-Video Dataset
von: Ju, Yiming, et al.
Veröffentlicht: (2025)
von: Ju, Yiming, et al.
Veröffentlicht: (2025)
InterAct-Video: Reasoning-Rich Video QA for Urban Traffic
von: Vishal, Joseph Raj, et al.
Veröffentlicht: (2025)
von: Vishal, Joseph Raj, et al.
Veröffentlicht: (2025)
VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos
von: Rasheed, Hanoona, et al.
Veröffentlicht: (2025)
von: Rasheed, Hanoona, et al.
Veröffentlicht: (2025)
LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
von: Wu, Haoning, et al.
Veröffentlicht: (2024)
von: Wu, Haoning, et al.
Veröffentlicht: (2024)
Beyond Isolated Facts: Synthesizing Narrative and Grounded Supervision for VideoQA
von: Liang, Jianxin, et al.
Veröffentlicht: (2025)
von: Liang, Jianxin, et al.
Veröffentlicht: (2025)
ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models
von: Wang, Yueqian, et al.
Veröffentlicht: (2025)
von: Wang, Yueqian, et al.
Veröffentlicht: (2025)
FrameMind: Frame-Interleaved Video Reasoning via Reinforcement Learning
von: Ge, Haonan, et al.
Veröffentlicht: (2025)
von: Ge, Haonan, et al.
Veröffentlicht: (2025)
Moment Sampling in Video LLMs for Long-Form Video QA
von: Chasmai, Mustafa, et al.
Veröffentlicht: (2025)
von: Chasmai, Mustafa, et al.
Veröffentlicht: (2025)
SpatialMem: Metric-Aligned Long-Horizon Video Memory for Language Grounding and QA
von: Zheng, Xinyi, et al.
Veröffentlicht: (2026)
von: Zheng, Xinyi, et al.
Veröffentlicht: (2026)
Weaver: End-to-End Agentic System Training for Video Interleaved Reasoning
von: Shi, Yudi, et al.
Veröffentlicht: (2026)
von: Shi, Yudi, et al.
Veröffentlicht: (2026)
YTCommentQA: Video Question Answerability in Instructional Videos
von: Yang, Saelyne, et al.
Veröffentlicht: (2024)
von: Yang, Saelyne, et al.
Veröffentlicht: (2024)
Multi-sentence Video Grounding for Long Video Generation
von: Feng, Wei, et al.
Veröffentlicht: (2024)
von: Feng, Wei, et al.
Veröffentlicht: (2024)
Motion-o: Trajectory-Grounded Video Reasoning
von: Galoaa, Bishoy, et al.
Veröffentlicht: (2026)
von: Galoaa, Bishoy, et al.
Veröffentlicht: (2026)
Bridging Modalities, Spanning Time: Structured Memory for Ultra-Long Agentic Video Reasoning
von: Li, Jiazheng, et al.
Veröffentlicht: (2026)
von: Li, Jiazheng, et al.
Veröffentlicht: (2026)
ReasVQA: Advancing VideoQA with Imperfect Reasoning Process
von: Liang, Jianxin, et al.
Veröffentlicht: (2025)
von: Liang, Jianxin, et al.
Veröffentlicht: (2025)
ENTER: Event Based Interpretable Reasoning for VideoQA
von: Ayyubi, Hammad, et al.
Veröffentlicht: (2025)
von: Ayyubi, Hammad, et al.
Veröffentlicht: (2025)
VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video Reasoning
von: Ding, Yang, et al.
Veröffentlicht: (2025)
von: Ding, Yang, et al.
Veröffentlicht: (2025)
REDUCIO! Generating 1K Video within 16 Seconds using Extremely Compressed Motion Latents
von: Tian, Rui, et al.
Veröffentlicht: (2024)
von: Tian, Rui, et al.
Veröffentlicht: (2024)
Adaptive Dense Evidence Refinement for Video Relational Reasoning for VRR-QA Challenge
von: Sun, Yuyang, et al.
Veröffentlicht: (2026)
von: Sun, Yuyang, et al.
Veröffentlicht: (2026)
StreamingCoT: A Dataset for Temporal Dynamics and Multimodal Chain-of-Thought Reasoning in Streaming VideoQA
von: Hu, Yuhang, et al.
Veröffentlicht: (2025)
von: Hu, Yuhang, et al.
Veröffentlicht: (2025)
SwinGS: Sliding Window Gaussian Splatting for Volumetric Video Streaming with Arbitrary Length
von: Liu, Bangya, et al.
Veröffentlicht: (2024)
von: Liu, Bangya, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
VoLUT: Efficient Volumetric streaming enhanced by LUT-based super-resolution
von: Wang, Chendong, et al.
Veröffentlicht: (2025) -
AVA: Towards Agentic Video Analytics with Vision Language Models
von: Yan, Yuxuan, et al.
Veröffentlicht: (2025) -
Babel: A Scalable Pre-trained Model for Multi-Modal Sensing via Expandable Modality Alignment
von: Dai, Shenghong, et al.
Veröffentlicht: (2024) -
AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation
von: Zhou, Ziwei, et al.
Veröffentlicht: (2026) -
StreamMind: Unlocking Full Frame Rate Streaming Video Dialogue through Event-Gated Cognition
von: Ding, Xin, et al.
Veröffentlicht: (2025)