Adaptive Dense Evidence Refinement for Video Relational Reasoning for VRR-QA Challenge
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Yuyang, Wu, Yongliang, Zhu, Xingyu, Chen, Yuxia, Jiang, Zhenxiang, Ji, Yangguang, Zhu, Wenbo, Shi, Yanxi, Wu, Jay, Wang, Shuo, Yang, Xu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Temporal Evidence Routing with Structured Visual Evidence for TimeLogicQA
by: Sun, Yuyang, et al.
Published: (2026)
by: Sun, Yuyang, et al.
Published: (2026)
VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit Cues
by: Swetha, Sirnam, et al.
Published: (2025)
by: Swetha, Sirnam, et al.
Published: (2025)
RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought
by: Lu, Yi, et al.
Published: (2025)
by: Lu, Yi, et al.
Published: (2025)
Reframe Anything: LLM Agent for Open World Video Reframing
by: Cao, Jiawang, et al.
Published: (2024)
by: Cao, Jiawang, et al.
Published: (2024)
VEU-Bench: Towards Comprehensive Understanding of Video Editing
by: Li, Bozheng, et al.
Published: (2025)
by: Li, Bozheng, et al.
Published: (2025)
Number it: Temporal Grounding Videos like Flipping Manga
by: Wu, Yongliang, et al.
Published: (2024)
by: Wu, Yongliang, et al.
Published: (2024)
OpusAnimation: Code-Based Dynamic Chart Generation
by: Li, Bozheng, et al.
Published: (2025)
by: Li, Bozheng, et al.
Published: (2025)
Hierarchical Budget Policy Optimization for Adaptive Reasoning
by: Lyu, Shangke, et al.
Published: (2025)
by: Lyu, Shangke, et al.
Published: (2025)
P‐9.14: A method of Improving Image Quality of VRR Flicker
by: Yizhuo Zhao, et al.
Published: (2024)
by: Yizhuo Zhao, et al.
Published: (2024)
Adapting Point Cloud Analysis via Multimodal Bayesian Distribution Learning
by: Zhu, Xingyu, et al.
Published: (2026)
by: Zhu, Xingyu, et al.
Published: (2026)
Zero-Shot Long-Form Video Understanding through Screenplay
by: Wu, Yongliang, et al.
Published: (2024)
by: Wu, Yongliang, et al.
Published: (2024)
Video Repurposing from User Generated Content: A Large-scale Dataset and Benchmark
by: Wu, Yongliang, et al.
Published: (2024)
by: Wu, Yongliang, et al.
Published: (2024)
LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization
by: Wu, Xingyu, et al.
Published: (2025)
by: Wu, Xingyu, et al.
Published: (2025)
Video-in-the-Loop: Span-Grounded Long Video QA with Interleaved Reasoning
by: Wang, Chendong, et al.
Published: (2025)
by: Wang, Chendong, et al.
Published: (2025)
DIBS: Enhancing Dense Video Captioning with Unlabeled Videos via Pseudo Boundary Enrichment and Online Refinement
by: Wu, Hao, et al.
Published: (2024)
by: Wu, Hao, et al.
Published: (2024)
Multifunctional agents based on 3‐dicycanovinylindan‐1‐one acceptor: Molecular design and phototheranostic application
by: Najia Zhu, et al.
Published: (2024)
by: Najia Zhu, et al.
Published: (2024)
ReAgent: Reversible Multi-Agent Reasoning for Knowledge-Enhanced Multi-Hop QA
by: Zhao, Xinjie, et al.
Published: (2025)
by: Zhao, Xinjie, et al.
Published: (2025)
RealFin: How Well Do LLMs Reason About Finance When Users Leave Things Unsaid?
by: Dai, Yuyang, et al.
Published: (2026)
by: Dai, Yuyang, et al.
Published: (2026)
Improving Visual Reasoning with Iterative Evidence Refinement
by: Shi, Zeru, et al.
Published: (2026)
by: Shi, Zeru, et al.
Published: (2026)
37‐2: Enhancing VRR Flicker Index Using Time‐Domain Analysis
by: Hyosun Kim, et al.
Published: (2025)
by: Hyosun Kim, et al.
Published: (2025)
First Place Solution to the Multiple-choice Video QA Track of The Second Perception Test Challenge
by: Peng, Yingzhe, et al.
Published: (2024)
by: Peng, Yingzhe, et al.
Published: (2024)
Learning Dual-Arm Push and Grasp Synergy in Dense Clutter
by: Wang, Yongliang, et al.
Published: (2024)
by: Wang, Yongliang, et al.
Published: (2024)
A Fano-Style Accuracy Upper Bound for LLM Single-Pass Reasoning in Multi-Hop QA
by: Wan, Kaiyang, et al.
Published: (2025)
by: Wan, Kaiyang, et al.
Published: (2025)
30‐1: Development of High‐integration HOP Panel with High‐frequency & VRR Driving
by: Hyeongseok Kim, et al.
Published: (2024)
by: Hyeongseok Kim, et al.
Published: (2024)
Reasoning Portability: Guiding Continual Learning for MLLMs in the RLVR Era
by: Hong, Qiuhe, et al.
Published: (2026)
by: Hong, Qiuhe, et al.
Published: (2026)
Question-Aware Evidence Ledgers for Video Relational Reasoning
by: Ou, Yilin, et al.
Published: (2026)
by: Ou, Yilin, et al.
Published: (2026)
ENTER: Event Based Interpretable Reasoning for VideoQA
by: Ayyubi, Hammad, et al.
Published: (2025)
by: Ayyubi, Hammad, et al.
Published: (2025)
BiDense: Binarization for Dense Prediction
by: Yin, Rui, et al.
Published: (2024)
by: Yin, Rui, et al.
Published: (2024)
InterAct-Video: Reasoning-Rich Video QA for Urban Traffic
by: Vishal, Joseph Raj, et al.
Published: (2025)
by: Vishal, Joseph Raj, et al.
Published: (2025)
VideoQA-SC: Adaptive Semantic Communication for Video Question Answering
by: Guo, Jiangyuan, et al.
Published: (2024)
by: Guo, Jiangyuan, et al.
Published: (2024)
Momentum Posterior Regularization for Multi-hop Dense Retrieval
by: Xia, Zehua, et al.
Published: (2024)
by: Xia, Zehua, et al.
Published: (2024)
rePIRL: Learn PRM with Inverse RL for LLM Reasoning
by: Wu, Xian, et al.
Published: (2026)
by: Wu, Xian, et al.
Published: (2026)
Nearly invariant subspaces and kernels of Toeplitz operators on the Hardy space over the bidisk
by: Zhu, Senhua, et al.
Published: (2024)
by: Zhu, Senhua, et al.
Published: (2024)
RAVEN: Robust Advertisement Video Violation Temporal Grounding via Reinforcement Reasoning
by: Ji, Deyi, et al.
Published: (2025)
by: Ji, Deyi, et al.
Published: (2025)
Look Carefully: Adaptive Visual Reinforcements in Multimodal Large Language Models for Hallucination Mitigation
by: Zhu, Xingyu, et al.
Published: (2026)
by: Zhu, Xingyu, et al.
Published: (2026)
Prediction-Augmented Mechanism Design for Weighted Facility Location
by: Shi, Yangguang, et al.
Published: (2025)
by: Shi, Yangguang, et al.
Published: (2025)
QA-MoE: Towards a Continuous Reliability Spectrum with Quality-Aware Mixture of Experts for Robust Multimodal Sentiment Analysis
by: Zhu, Yitong, et al.
Published: (2026)
by: Zhu, Yitong, et al.
Published: (2026)
Topical Application of Fluoxetine Improves DNCB‐Induced Atopic Dermatitis in Mice
by: Xue Jiang, et al.
Published: (2025)
by: Xue Jiang, et al.
Published: (2025)
HCR-Reasoner: Synergizing Large Language Models and Theory for Human-like Causal Reasoning
by: Zhang, Yanxi, et al.
Published: (2025)
by: Zhang, Yanxi, et al.
Published: (2025)
Adaptive Testing Environment Generation for Connected and Automated Vehicles with Dense Reinforcement Learning
by: Yang, Jingxuan, et al.
Published: (2024)
by: Yang, Jingxuan, et al.
Published: (2024)
Similar Items
-
Temporal Evidence Routing with Structured Visual Evidence for TimeLogicQA
by: Sun, Yuyang, et al.
Published: (2026) -
VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit Cues
by: Swetha, Sirnam, et al.
Published: (2025) -
RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought
by: Lu, Yi, et al.
Published: (2025) -
Reframe Anything: LLM Agent for Open World Video Reframing
by: Cao, Jiawang, et al.
Published: (2024) -
VEU-Bench: Towards Comprehensive Understanding of Video Editing
by: Li, Bozheng, et al.
Published: (2025)