Grounding Video Reasoning in Physical Signals
Fuente:
arXiv
Saved in:
| Main Authors: | Osmanli, Alibay, Cheng, Zixu, Gong, Shaogang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking
by: Cheng, Zixu, et al.
Published: (2026)
by: Cheng, Zixu, et al.
Published: (2026)
V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning
by: Cheng, Zixu, et al.
Published: (2025)
by: Cheng, Zixu, et al.
Published: (2025)
INT: Instance-Specific Negative Mining for Task-Generic Promptable Segmentation
by: Hu, Jian, et al.
Published: (2025)
by: Hu, Jian, et al.
Published: (2025)
SHINE: Saliency-aware HIerarchical NEgative Ranking for Compositional Temporal Grounding
by: Cheng, Zixu, et al.
Published: (2024)
by: Cheng, Zixu, et al.
Published: (2024)
CoS: Chain-of-Shot Prompting for Long Video Understanding
by: Hu, Jian, et al.
Published: (2025)
by: Hu, Jian, et al.
Published: (2025)
Uncertainty-quantified Rollout Policy Adaptation for Unlabelled Cross-domain Temporal Grounding
by: Hu, Jian, et al.
Published: (2025)
by: Hu, Jian, et al.
Published: (2025)
Neuro-Symbolic Spatial Reasoning in Segmentation
by: Lin, Jiayi, et al.
Published: (2025)
by: Lin, Jiayi, et al.
Published: (2025)
Hybrid-Learning Video Moment Retrieval across Multi-Domain Labels
by: Cai, Weitong, et al.
Published: (2024)
by: Cai, Weitong, et al.
Published: (2024)
Few-Shot Image Generation by Conditional Relaxing Diffusion Inversion
by: Cao, Yu, et al.
Published: (2024)
by: Cao, Yu, et al.
Published: (2024)
MLLM as Video Narrator: Mitigating Modality Imbalance in Video Moment Retrieval
by: Cai, Weitong, et al.
Published: (2024)
by: Cai, Weitong, et al.
Published: (2024)
Generative Video Diffusion for Unseen Novel Semantic Video Moment Retrieval
by: Luo, Dezhao, et al.
Published: (2024)
by: Luo, Dezhao, et al.
Published: (2024)
ViSMaP: Unsupervised Hour-long Video Summarisation by Meta-Prompting
by: Hu, Jian, et al.
Published: (2025)
by: Hu, Jian, et al.
Published: (2025)
Training-free Zero-shot Composed Image Retrieval with Local Concept Reranking
by: Sun, Shitong, et al.
Published: (2023)
by: Sun, Shitong, et al.
Published: (2023)
TCMA: Text-Conditioned Multi-granularity Alignment for Drone Cross-Modal Text-Video Retrieval
by: Zhao, Zixu, et al.
Published: (2025)
by: Zhao, Zixu, et al.
Published: (2025)
Temporal Score Analysis for Understanding and Correcting Diffusion Artifacts
by: Cao, Yu, et al.
Published: (2025)
by: Cao, Yu, et al.
Published: (2025)
Leveraging Hallucinations to Reduce Manual Prompt Dependency in Promptable Segmentation
by: Hu, Jian, et al.
Published: (2024)
by: Hu, Jian, et al.
Published: (2024)
InvSeg: Test-Time Prompt Inversion for Semantic Segmentation
by: Lin, Jiayi, et al.
Published: (2024)
by: Lin, Jiayi, et al.
Published: (2024)
Enhancing Zero-Shot Facial Expression Recognition by LLM Knowledge Transfer
by: Zhao, Zengqun, et al.
Published: (2024)
by: Zhao, Zengqun, et al.
Published: (2024)
R^3: Composed Video Retrieval via Reasoning-Guided Recalling and Re-ranking
by: Li, Zixu, et al.
Published: (2026)
by: Li, Zixu, et al.
Published: (2026)
Video-in-the-Loop: Span-Grounded Long Video QA with Interleaved Reasoning
by: Wang, Chendong, et al.
Published: (2025)
by: Wang, Chendong, et al.
Published: (2025)
Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding
by: Wu, Zhixuan, et al.
Published: (2026)
by: Wu, Zhixuan, et al.
Published: (2026)
LatSearch: Latent Reward-Guided Search for Faster Inference-Time Scaling in Video Diffusion
by: Zhao, Zengqun, et al.
Published: (2026)
by: Zhao, Zengqun, et al.
Published: (2026)
Clinically-Grounded Counterfactual Reasoning for Medical Video Diagnosis
by: Gao, Jianzhe, et al.
Published: (2026)
by: Gao, Jianzhe, et al.
Published: (2026)
Temporal Grounding as a Learning Signal for Referring Video Object Segmentation
by: Lee, Seunghun, et al.
Published: (2025)
by: Lee, Seunghun, et al.
Published: (2025)
Video Evidence to Reasoning Efficient Video Understanding via Explicit Evidence Grounding
by: Huang, Yanxiang, et al.
Published: (2026)
by: Huang, Yanxiang, et al.
Published: (2026)
XFMamba: Cross-Fusion Mamba for Multi-View Medical Image Classification
by: Zheng, Xiaoyu, et al.
Published: (2025)
by: Zheng, Xiaoyu, et al.
Published: (2025)
SYNCR: A Cross-Video Reasoning Benchmark with Synthetic Grounding
by: Ghazanfari, Sara, et al.
Published: (2026)
by: Ghazanfari, Sara, et al.
Published: (2026)
VTimeCoT: Thinking by Drawing for Video Temporal Grounding and Reasoning
by: Zhang, Jinglei, et al.
Published: (2025)
by: Zhang, Jinglei, et al.
Published: (2025)
Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding
by: Zheng, Zelin, et al.
Published: (2026)
by: Zheng, Zelin, et al.
Published: (2026)
Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding
by: Chen, Houlun, et al.
Published: (2026)
by: Chen, Houlun, et al.
Published: (2026)
Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning
by: Li, Yifei, et al.
Published: (2025)
by: Li, Yifei, et al.
Published: (2025)
Exploiting Auxiliary Caption for Video Grounding
by: Li, Hongxiang, et al.
Published: (2023)
by: Li, Hongxiang, et al.
Published: (2023)
NEWTON: Agentic Planning for Physically Grounded Video Generation
by: Feng, Yuxiang, et al.
Published: (2026)
by: Feng, Yuxiang, et al.
Published: (2026)
Adaptive Domain Shift in Diffusion Models for Cross-Modality Image Translation
by: Wang, Zihao, et al.
Published: (2026)
by: Wang, Zihao, et al.
Published: (2026)
PhysCtrl: Generative Physics for Controllable and Physics-Grounded Video Generation
by: Wang, Chen, et al.
Published: (2025)
by: Wang, Chen, et al.
Published: (2025)
Motion-o: Trajectory-Grounded Video Reasoning
by: Galoaa, Bishoy, et al.
Published: (2026)
by: Galoaa, Bishoy, et al.
Published: (2026)
Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
by: Maaz, Muhammad, et al.
Published: (2025)
by: Maaz, Muhammad, et al.
Published: (2025)
Know-Show: Benchmarking Video-Language Models on Spatio-Temporal Grounded Reasoning
by: Sugandhika, Chinthani, et al.
Published: (2025)
by: Sugandhika, Chinthani, et al.
Published: (2025)
From Priors to Perception: Grounding Video-LLMs in Physical Reality
by: Zhao, Zicheng, et al.
Published: (2026)
by: Zhao, Zicheng, et al.
Published: (2026)
VideoSAM: Open-World Video Segmentation
by: Guo, Pinxue, et al.
Published: (2024)
by: Guo, Pinxue, et al.
Published: (2024)
Similar Items
-
GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking
by: Cheng, Zixu, et al.
Published: (2026) -
V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning
by: Cheng, Zixu, et al.
Published: (2025) -
INT: Instance-Specific Negative Mining for Task-Generic Promptable Segmentation
by: Hu, Jian, et al.
Published: (2025) -
SHINE: Saliency-aware HIerarchical NEgative Ranking for Compositional Temporal Grounding
by: Cheng, Zixu, et al.
Published: (2024) -
CoS: Chain-of-Shot Prompting for Long Video Understanding
by: Hu, Jian, et al.
Published: (2025)