Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zheng, Zelin, Liu, Xinyan, Li, Ruixin, Chan, Antoni B., Li, Guorong, Huang, Qingming, Qing, Laiyun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Exploring the Temporal Consistency for Point-Level Weakly-Supervised Temporal Action Localization
von: Ma, Yunchuan, et al.
Veröffentlicht: (2026)
von: Ma, Yunchuan, et al.
Veröffentlicht: (2026)
Boosting Point-supervised Temporal Action Localization via Text Refinement and Alignment
von: Ma, Yunchuan, et al.
Veröffentlicht: (2026)
von: Ma, Yunchuan, et al.
Veröffentlicht: (2026)
RETTA: Retrieval-Enhanced Test-Time Adaptation for Zero-Shot Video Captioning
von: Ma, Yunchuan, et al.
Veröffentlicht: (2024)
von: Ma, Yunchuan, et al.
Veröffentlicht: (2024)
SDVPT: Semantic-Driven Visual Prompt Tuning for Open-World Object Counting
von: Zhao, Yiming, et al.
Veröffentlicht: (2025)
von: Zhao, Yiming, et al.
Veröffentlicht: (2025)
Keeping the Evidence Chain: Semantic Evidence Allocation for Training-Free Token Pruning in Video Temporal Grounding
von: Li, Jiaqi, et al.
Veröffentlicht: (2026)
von: Li, Jiaqi, et al.
Veröffentlicht: (2026)
Bias-Conflict Sample Synthesis and Adversarial Removal Debias Strategy for Temporal Sentence Grounding in Video
von: Qi, Zhaobo, et al.
Veröffentlicht: (2024)
von: Qi, Zhaobo, et al.
Veröffentlicht: (2024)
Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence
von: Meng, Jiahao, et al.
Veröffentlicht: (2025)
von: Meng, Jiahao, et al.
Veröffentlicht: (2025)
GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking
von: Cheng, Zixu, et al.
Veröffentlicht: (2026)
von: Cheng, Zixu, et al.
Veröffentlicht: (2026)
Dr.V: A Hierarchical Perception-Temporal-Cognition Framework to Diagnose Video Hallucination by Fine-grained Spatial-Temporal Grounding
von: Luo, Meng, et al.
Veröffentlicht: (2025)
von: Luo, Meng, et al.
Veröffentlicht: (2025)
Know-Show: Benchmarking Video-Language Models on Spatio-Temporal Grounded Reasoning
von: Sugandhika, Chinthani, et al.
Veröffentlicht: (2025)
von: Sugandhika, Chinthani, et al.
Veröffentlicht: (2025)
VTimeCoT: Thinking by Drawing for Video Temporal Grounding and Reasoning
von: Zhang, Jinglei, et al.
Veröffentlicht: (2025)
von: Zhang, Jinglei, et al.
Veröffentlicht: (2025)
Context-Guided Spatio-Temporal Video Grounding
von: Gu, Xin, et al.
Veröffentlicht: (2024)
von: Gu, Xin, et al.
Veröffentlicht: (2024)
EvoGround: Self-Evolving Video Agents for Video Temporal Grounding
von: Jung, Minjoon, et al.
Veröffentlicht: (2026)
von: Jung, Minjoon, et al.
Veröffentlicht: (2026)
VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
von: Wang, Shihao, et al.
Veröffentlicht: (2025)
von: Wang, Shihao, et al.
Veröffentlicht: (2025)
Moment Quantization for Video Temporal Grounding
von: Sun, Xiaolong, et al.
Veröffentlicht: (2025)
von: Sun, Xiaolong, et al.
Veröffentlicht: (2025)
VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video Reasoning
von: Liu, Ye, et al.
Veröffentlicht: (2025)
von: Liu, Ye, et al.
Veröffentlicht: (2025)
T2SGrid: Temporal-to-Spatial Gridification for Video Temporal Grounding
von: Guo, Chaohong, et al.
Veröffentlicht: (2026)
von: Guo, Chaohong, et al.
Veröffentlicht: (2026)
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding
von: Wasim, Syed Talal, et al.
Veröffentlicht: (2023)
von: Wasim, Syed Talal, et al.
Veröffentlicht: (2023)
SOVC: Subject-Oriented Video Captioning
von: Teng, Chang, et al.
Veröffentlicht: (2023)
von: Teng, Chang, et al.
Veröffentlicht: (2023)
Towards Long-Form Spatio-Temporal Video Grounding
von: Gu, Xin, et al.
Veröffentlicht: (2026)
von: Gu, Xin, et al.
Veröffentlicht: (2026)
Efficient Temporal Sentence Grounding in Videos with Multi-Teacher Knowledge Distillation
von: Liang, Renjie, et al.
Veröffentlicht: (2023)
von: Liang, Renjie, et al.
Veröffentlicht: (2023)
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding
von: Ahn, Geo, et al.
Veröffentlicht: (2026)
von: Ahn, Geo, et al.
Veröffentlicht: (2026)
Temporally Grounding Instructional Diagrams in Unconstrained Videos
von: Zhang, Jiahao, et al.
Veröffentlicht: (2024)
von: Zhang, Jiahao, et al.
Veröffentlicht: (2024)
Video Evidence to Reasoning Efficient Video Understanding via Explicit Evidence Grounding
von: Huang, Yanxiang, et al.
Veröffentlicht: (2026)
von: Huang, Yanxiang, et al.
Veröffentlicht: (2026)
Universal Video Temporal Grounding with Generative Multi-modal Large Language Models
von: Li, Zeqian, et al.
Veröffentlicht: (2025)
von: Li, Zeqian, et al.
Veröffentlicht: (2025)
Static and Dynamic Graph Alignment Network for Temporal Video Grounding
von: Hu, Zhanjie, et al.
Veröffentlicht: (2026)
von: Hu, Zhanjie, et al.
Veröffentlicht: (2026)
Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models
von: Wang, Haibo, et al.
Veröffentlicht: (2024)
von: Wang, Haibo, et al.
Veröffentlicht: (2024)
From Priors to Perception: Grounding Video-LLMs in Physical Reality
von: Zhao, Zicheng, et al.
Veröffentlicht: (2026)
von: Zhao, Zicheng, et al.
Veröffentlicht: (2026)
E.M.Ground: A Temporal Grounding Vid-LLM with Holistic Event Perception and Matching
von: Nie, Jiahao, et al.
Veröffentlicht: (2026)
von: Nie, Jiahao, et al.
Veröffentlicht: (2026)
Probing the Reliability of Driving VLMs: From Inconsistent Responses to Grounded Temporal Reasoning
von: Chang, Chun-Peng, et al.
Veröffentlicht: (2026)
von: Chang, Chun-Peng, et al.
Veröffentlicht: (2026)
Edit Temporal-Consistent Videos with Image Diffusion Model
von: Wang, Yuanzhi, et al.
Veröffentlicht: (2023)
von: Wang, Yuanzhi, et al.
Veröffentlicht: (2023)
TRACE: Temporal Grounding Video LLM via Causal Event Modeling
von: Guo, Yongxin, et al.
Veröffentlicht: (2024)
von: Guo, Yongxin, et al.
Veröffentlicht: (2024)
Multi-Scale Contrastive Learning for Video Temporal Grounding
von: Nguyen, Thong Thanh, et al.
Veröffentlicht: (2024)
von: Nguyen, Thong Thanh, et al.
Veröffentlicht: (2024)
Number it: Temporal Grounding Videos like Flipping Manga
von: Wu, Yongliang, et al.
Veröffentlicht: (2024)
von: Wu, Yongliang, et al.
Veröffentlicht: (2024)
VideoMolmo: Spatio-Temporal Grounding Meets Pointing
von: Ahmad, Ghazi Shazan, et al.
Veröffentlicht: (2025)
von: Ahmad, Ghazi Shazan, et al.
Veröffentlicht: (2025)
Unified Static and Dynamic Network: Efficient Temporal Filtering for Video Grounding
von: Hu, Jingjing, et al.
Veröffentlicht: (2024)
von: Hu, Jingjing, et al.
Veröffentlicht: (2024)
Grounding is All You Need? Dual Temporal Grounding for Video Dialog
von: Qin, You, et al.
Veröffentlicht: (2024)
von: Qin, You, et al.
Veröffentlicht: (2024)
Factorized Learning for Temporally Grounded Video-Language Models
von: Zeng, Wenzheng, et al.
Veröffentlicht: (2025)
von: Zeng, Wenzheng, et al.
Veröffentlicht: (2025)
CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering
von: Zhang, Mingfang, et al.
Veröffentlicht: (2026)
von: Zhang, Mingfang, et al.
Veröffentlicht: (2026)
Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding
von: Zheng, Minghang, et al.
Veröffentlicht: (2025)
von: Zheng, Minghang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Exploring the Temporal Consistency for Point-Level Weakly-Supervised Temporal Action Localization
von: Ma, Yunchuan, et al.
Veröffentlicht: (2026) -
Boosting Point-supervised Temporal Action Localization via Text Refinement and Alignment
von: Ma, Yunchuan, et al.
Veröffentlicht: (2026) -
RETTA: Retrieval-Enhanced Test-Time Adaptation for Zero-Shot Video Captioning
von: Ma, Yunchuan, et al.
Veröffentlicht: (2024) -
SDVPT: Semantic-Driven Visual Prompt Tuning for Open-World Object Counting
von: Zhao, Yiming, et al.
Veröffentlicht: (2025) -
Keeping the Evidence Chain: Semantic Evidence Allocation for Training-Free Token Pruning in Video Temporal Grounding
von: Li, Jiaqi, et al.
Veröffentlicht: (2026)