Tempo-R0: A Video-MLLM for Temporal Video Grounding through Efficient Temporal Sensing Reinforcement Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yue, Feng, Zhang, Zhaoxing, Jiao, Junming, Liang, Zhengyu, Cao, Shiwen, Zhang, Feifei, Shen, Rong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MASR: Self-Reflective Reasoning through Multimodal Hierarchical Attention Focusing for Agent-based Video Understanding
von: Cao, Shiwen, et al.
Veröffentlicht: (2025)
von: Cao, Shiwen, et al.
Veröffentlicht: (2025)
Datasets and Recipes for Video Temporal Grounding via Reinforcement Learning
von: Chen, Ruizhe, et al.
Veröffentlicht: (2025)
von: Chen, Ruizhe, et al.
Veröffentlicht: (2025)
$R^2$-Tuning: Efficient Image-to-Video Transfer Learning for Video Temporal Grounding
von: Liu, Ye, et al.
Veröffentlicht: (2024)
von: Liu, Ye, et al.
Veröffentlicht: (2024)
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding
von: Ahn, Geo, et al.
Veröffentlicht: (2026)
von: Ahn, Geo, et al.
Veröffentlicht: (2026)
Test-Time Temporal Sampling for Efficient MLLM Video Understanding
von: Wang, Kaibin, et al.
Veröffentlicht: (2025)
von: Wang, Kaibin, et al.
Veröffentlicht: (2025)
VideoTG-R1: Boosting Video Temporal Grounding via Curriculum Reinforcement Learning on Reflected Boundary Annotations
von: Dong, Lu, et al.
Veröffentlicht: (2025)
von: Dong, Lu, et al.
Veröffentlicht: (2025)
MASRA: MLLM-Assisted Semantic-Relational Consistent Alignment for Video Temporal Grounding
von: Ran, Ran, et al.
Veröffentlicht: (2026)
von: Ran, Ran, et al.
Veröffentlicht: (2026)
EvoGround: Self-Evolving Video Agents for Video Temporal Grounding
von: Jung, Minjoon, et al.
Veröffentlicht: (2026)
von: Jung, Minjoon, et al.
Veröffentlicht: (2026)
Temporally Grounding Instructional Diagrams in Unconstrained Videos
von: Zhang, Jiahao, et al.
Veröffentlicht: (2024)
von: Zhang, Jiahao, et al.
Veröffentlicht: (2024)
Efficient Temporal Sentence Grounding in Videos with Multi-Teacher Knowledge Distillation
von: Liang, Renjie, et al.
Veröffentlicht: (2023)
von: Liang, Renjie, et al.
Veröffentlicht: (2023)
ToG-Bench: Task-Oriented Spatio-Temporal Grounding in Egocentric Videos
von: Xu, Qi'ao, et al.
Veröffentlicht: (2025)
von: Xu, Qi'ao, et al.
Veröffentlicht: (2025)
VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video Reasoning
von: Ding, Yang, et al.
Veröffentlicht: (2025)
von: Ding, Yang, et al.
Veröffentlicht: (2025)
Context-Guided Spatio-Temporal Video Grounding
von: Gu, Xin, et al.
Veröffentlicht: (2024)
von: Gu, Xin, et al.
Veröffentlicht: (2024)
Learning Temporally Consistent Video Depth from Video Diffusion Priors
von: Shao, Jiahao, et al.
Veröffentlicht: (2024)
von: Shao, Jiahao, et al.
Veröffentlicht: (2024)
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding
von: Gao, Shida, et al.
Veröffentlicht: (2025)
von: Gao, Shida, et al.
Veröffentlicht: (2025)
End-to-End Streaming Video Temporal Action Segmentation with Reinforce Learning
von: Zhang, Jinrong, et al.
Veröffentlicht: (2023)
von: Zhang, Jinrong, et al.
Veröffentlicht: (2023)
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding
von: Luo, Fuwen, et al.
Veröffentlicht: (2025)
von: Luo, Fuwen, et al.
Veröffentlicht: (2025)
VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
von: Wang, Shihao, et al.
Veröffentlicht: (2025)
von: Wang, Shihao, et al.
Veröffentlicht: (2025)
Towards Long-Form Spatio-Temporal Video Grounding
von: Gu, Xin, et al.
Veröffentlicht: (2026)
von: Gu, Xin, et al.
Veröffentlicht: (2026)
STVG-R1: Incentivizing Instance-Level Reasoning and Grounding in Videos via Reinforcement Learning
von: Zhang, Xiaowen, et al.
Veröffentlicht: (2026)
von: Zhang, Xiaowen, et al.
Veröffentlicht: (2026)
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
von: Wang, Ye, et al.
Veröffentlicht: (2025)
von: Wang, Ye, et al.
Veröffentlicht: (2025)
Efficient Motion-Aware Video MLLM
von: Zhao, Zijia, et al.
Veröffentlicht: (2025)
von: Zhao, Zijia, et al.
Veröffentlicht: (2025)
Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence
von: Meng, Jiahao, et al.
Veröffentlicht: (2025)
von: Meng, Jiahao, et al.
Veröffentlicht: (2025)
Thinking With Bounding Boxes: Enhancing Spatio-Temporal Video Grounding via Reinforcement Fine-Tuning
von: Gu, Xin, et al.
Veröffentlicht: (2025)
von: Gu, Xin, et al.
Veröffentlicht: (2025)
TempoControl: Temporal Attention Guidance for Text-to-Video Models
von: Schiber, Shira, et al.
Veröffentlicht: (2025)
von: Schiber, Shira, et al.
Veröffentlicht: (2025)
Moment Quantization for Video Temporal Grounding
von: Sun, Xiaolong, et al.
Veröffentlicht: (2025)
von: Sun, Xiaolong, et al.
Veröffentlicht: (2025)
Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models
von: Wang, Haibo, et al.
Veröffentlicht: (2024)
von: Wang, Haibo, et al.
Veröffentlicht: (2024)
Multi-Scale Contrastive Learning for Video Temporal Grounding
von: Nguyen, Thong Thanh, et al.
Veröffentlicht: (2024)
von: Nguyen, Thong Thanh, et al.
Veröffentlicht: (2024)
Factorized Learning for Temporally Grounded Video-Language Models
von: Zeng, Wenzheng, et al.
Veröffentlicht: (2025)
von: Zeng, Wenzheng, et al.
Veröffentlicht: (2025)
VideoMolmo: Spatio-Temporal Grounding Meets Pointing
von: Ahmad, Ghazi Shazan, et al.
Veröffentlicht: (2025)
von: Ahmad, Ghazi Shazan, et al.
Veröffentlicht: (2025)
T2SGrid: Temporal-to-Spatial Gridification for Video Temporal Grounding
von: Guo, Chaohong, et al.
Veröffentlicht: (2026)
von: Guo, Chaohong, et al.
Veröffentlicht: (2026)
SmartSight: Mitigating Hallucination in Video-LLMs Without Compromising Video Understanding via Temporal Attention Collapse
von: Sun, Yiming, et al.
Veröffentlicht: (2025)
von: Sun, Yiming, et al.
Veröffentlicht: (2025)
TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval
von: Shen, Leqi, et al.
Veröffentlicht: (2024)
von: Shen, Leqi, et al.
Veröffentlicht: (2024)
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding
von: Wasim, Syed Talal, et al.
Veröffentlicht: (2023)
von: Wasim, Syed Talal, et al.
Veröffentlicht: (2023)
TimeScope: Towards Task-Oriented Temporal Grounding In Long Videos
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
Diversified Augmentation with Domain Adaptation for Debiased Video Temporal Grounding
von: Ren, Junlong, et al.
Veröffentlicht: (2025)
von: Ren, Junlong, et al.
Veröffentlicht: (2025)
TempSamp-R1: Effective Temporal Sampling with Reinforcement Fine-Tuning for Video LLMs
von: Li, Yunheng, et al.
Veröffentlicht: (2025)
von: Li, Yunheng, et al.
Veröffentlicht: (2025)
Unified Static and Dynamic Network: Efficient Temporal Filtering for Video Grounding
von: Hu, Jingjing, et al.
Veröffentlicht: (2024)
von: Hu, Jingjing, et al.
Veröffentlicht: (2024)
FlashVTG: Feature Layering and Adaptive Score Handling Network for Video Temporal Grounding
von: Cao, Zhuo, et al.
Veröffentlicht: (2024)
von: Cao, Zhuo, et al.
Veröffentlicht: (2024)
GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking
von: Cheng, Zixu, et al.
Veröffentlicht: (2026)
von: Cheng, Zixu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
MASR: Self-Reflective Reasoning through Multimodal Hierarchical Attention Focusing for Agent-based Video Understanding
von: Cao, Shiwen, et al.
Veröffentlicht: (2025) -
Datasets and Recipes for Video Temporal Grounding via Reinforcement Learning
von: Chen, Ruizhe, et al.
Veröffentlicht: (2025) -
$R^2$-Tuning: Efficient Image-to-Video Transfer Learning for Video Temporal Grounding
von: Liu, Ye, et al.
Veröffentlicht: (2024) -
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding
von: Ahn, Geo, et al.
Veröffentlicht: (2026) -
Test-Time Temporal Sampling for Efficient MLLM Video Understanding
von: Wang, Kaibin, et al.
Veröffentlicht: (2025)