Towards Long-Form Spatio-Temporal Video Grounding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gu, Xin, Fan, Bing, Yao, Jiali, Zhang, Zhipeng, Huang, Yan, Han, Cheng, Fan, Heng, Zhang, Libo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
OmniSTVG: Toward Spatio-Temporal Omni-Object Video Grounding
von: Yao, Jiali, et al.
Veröffentlicht: (2025)
von: Yao, Jiali, et al.
Veröffentlicht: (2025)
Context-Guided Spatio-Temporal Video Grounding
von: Gu, Xin, et al.
Veröffentlicht: (2024)
von: Gu, Xin, et al.
Veröffentlicht: (2024)
Thinking With Bounding Boxes: Enhancing Spatio-Temporal Video Grounding via Reinforcement Fine-Tuning
von: Gu, Xin, et al.
Veröffentlicht: (2025)
von: Gu, Xin, et al.
Veröffentlicht: (2025)
Knowing Your Target: Target-Aware Transformer Makes Better Spatio-Temporal Video Grounding
von: Gu, Xin, et al.
Veröffentlicht: (2025)
von: Gu, Xin, et al.
Veröffentlicht: (2025)
High-Fidelity Image Inpainting with Multimodal Guided GAN Inversion
von: Zhang, Libo, et al.
Veröffentlicht: (2025)
von: Zhang, Libo, et al.
Veröffentlicht: (2025)
Robust Ego-Exo Correspondence with Long-Term Memory
von: Hu, Yijun, et al.
Veröffentlicht: (2025)
von: Hu, Yijun, et al.
Veröffentlicht: (2025)
Flow-Guided Diffusion for Video Inpainting
von: Gu, Bohai, et al.
Veröffentlicht: (2023)
von: Gu, Bohai, et al.
Veröffentlicht: (2023)
DMTrack: Spatio-Temporal Multimodal Tracking via Dual-Adapter
von: Li, Weihong, et al.
Veröffentlicht: (2025)
von: Li, Weihong, et al.
Veröffentlicht: (2025)
Bridging Time and Space: Decoupled Spatio-Temporal Alignment for Video Grounding
von: Tu, Xuezhen, et al.
Veröffentlicht: (2026)
von: Tu, Xuezhen, et al.
Veröffentlicht: (2026)
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding
von: Wasim, Syed Talal, et al.
Veröffentlicht: (2023)
von: Wasim, Syed Talal, et al.
Veröffentlicht: (2023)
Towards Temporal Compositional Reasoning in Long-Form Sports Videos
von: Cao, Siyu, et al.
Veröffentlicht: (2026)
von: Cao, Siyu, et al.
Veröffentlicht: (2026)
Accurate and Fast Compressed Video Captioning
von: Shen, Yaojie, et al.
Veröffentlicht: (2023)
von: Shen, Yaojie, et al.
Veröffentlicht: (2023)
Video-Language Alignment via Spatio-Temporal Graph Transformer
von: Zhang, Shi-Xue, et al.
Veröffentlicht: (2024)
von: Zhang, Shi-Xue, et al.
Veröffentlicht: (2024)
VideoMolmo: Spatio-Temporal Grounding Meets Pointing
von: Ahmad, Ghazi Shazan, et al.
Veröffentlicht: (2025)
von: Ahmad, Ghazi Shazan, et al.
Veröffentlicht: (2025)
TimeScope: Towards Task-Oriented Temporal Grounding In Long Videos
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
von: Liu, Xiangrui, et al.
Veröffentlicht: (2025)
Beyond MOT: Semantic Multi-Object Tracking
von: Li, Yunhao, et al.
Veröffentlicht: (2024)
von: Li, Yunhao, et al.
Veröffentlicht: (2024)
SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability
von: Wang, Jiankang, et al.
Veröffentlicht: (2025)
von: Wang, Jiankang, et al.
Veröffentlicht: (2025)
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding
von: Gao, Shida, et al.
Veröffentlicht: (2025)
von: Gao, Shida, et al.
Veröffentlicht: (2025)
Edit3K: Universal Representation Learning for Video Editing Components
von: Gu, Xin, et al.
Veröffentlicht: (2024)
von: Gu, Xin, et al.
Veröffentlicht: (2024)
FiLA-Video: Spatio-Temporal Compression for Fine-Grained Long Video Understanding
von: Guo, Yanan, et al.
Veröffentlicht: (2025)
von: Guo, Yanan, et al.
Veröffentlicht: (2025)
NeuS-QA: Grounding Long-Form Video Understanding in Temporal Logic and Neuro-Symbolic Reasoning
von: Shah, Sahil, et al.
Veröffentlicht: (2025)
von: Shah, Sahil, et al.
Veröffentlicht: (2025)
CGTrack: Cascade Gating Network with Hierarchical Feature Aggregation for UAV Tracking
von: Li, Weihong, et al.
Veröffentlicht: (2025)
von: Li, Weihong, et al.
Veröffentlicht: (2025)
Agentic Spatio-Temporal Grounding via Collaborative Reasoning
von: Zhao, Heng, et al.
Veröffentlicht: (2026)
von: Zhao, Heng, et al.
Veröffentlicht: (2026)
CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering
von: Zhang, Mingfang, et al.
Veröffentlicht: (2026)
von: Zhang, Mingfang, et al.
Veröffentlicht: (2026)
Multi-sentence Video Grounding for Long Video Generation
von: Feng, Wei, et al.
Veröffentlicht: (2024)
von: Feng, Wei, et al.
Veröffentlicht: (2024)
Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos
von: Chen, Qirui, et al.
Veröffentlicht: (2024)
von: Chen, Qirui, et al.
Veröffentlicht: (2024)
Towards Visual Query Segmentation in the Wild
von: Fan, Bing, et al.
Veröffentlicht: (2026)
von: Fan, Bing, et al.
Veröffentlicht: (2026)
Enhancing Video-Language Representations with Structural Spatio-Temporal Alignment
von: Fei, Hao, et al.
Veröffentlicht: (2024)
von: Fei, Hao, et al.
Veröffentlicht: (2024)
SpatioTemporal Learning for Human Pose Estimation in Sparsely-Labeled Videos
von: Jiao, Yingying, et al.
Veröffentlicht: (2025)
von: Jiao, Yingying, et al.
Veröffentlicht: (2025)
Unleashing the Potential of Multimodal LLMs for Zero-Shot Spatio-Temporal Video Grounding
von: Yang, Zaiquan, et al.
Veröffentlicht: (2025)
von: Yang, Zaiquan, et al.
Veröffentlicht: (2025)
Contextual Self-paced Learning for Weakly Supervised Spatio-Temporal Video Grounding
von: Kumar, Akash, et al.
Veröffentlicht: (2025)
von: Kumar, Akash, et al.
Veröffentlicht: (2025)
Know-Show: Benchmarking Video-Language Models on Spatio-Temporal Grounded Reasoning
von: Sugandhika, Chinthani, et al.
Veröffentlicht: (2025)
von: Sugandhika, Chinthani, et al.
Veröffentlicht: (2025)
VastTrack: Vast Category Visual Object Tracking
von: Peng, Liang, et al.
Veröffentlicht: (2024)
von: Peng, Liang, et al.
Veröffentlicht: (2024)
TubeRMC: Tube-conditioned Reconstruction with Mutual Constraints for Weakly-supervised Spatio-Temporal Video Grounding
von: Li, Jinxuan, et al.
Veröffentlicht: (2025)
von: Li, Jinxuan, et al.
Veröffentlicht: (2025)
A Spatio-Temporal Attention-Based Method for Detecting Student Classroom Behaviors
von: Yang, Fan
Veröffentlicht: (2023)
von: Yang, Fan
Veröffentlicht: (2023)
Cyclic Refiner: Object-Aware Temporal Representation Learning for Multi-View 3D Detection and Tracking
von: Guo, Mingzhe, et al.
Veröffentlicht: (2024)
von: Guo, Mingzhe, et al.
Veröffentlicht: (2024)
Leveraging Consistent Spatio-Temporal Correspondence for Robust Visual Odometry
von: Zhang, Zhaoxing, et al.
Veröffentlicht: (2024)
von: Zhang, Zhaoxing, et al.
Veröffentlicht: (2024)
L-STEC: Learned Video Compression with Long-term Spatio-Temporal Enhanced Context
von: Zhang, Tiange, et al.
Veröffentlicht: (2025)
von: Zhang, Tiange, et al.
Veröffentlicht: (2025)
Spatio-Temporal Distortion Aware Omnidirectional Video Super-Resolution
von: An, Hongyu, et al.
Veröffentlicht: (2024)
von: An, Hongyu, et al.
Veröffentlicht: (2024)
STAF: 3D Human Mesh Recovery from Video with Spatio-Temporal Alignment Fusion
von: Yao, Wei, et al.
Veröffentlicht: (2024)
von: Yao, Wei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
OmniSTVG: Toward Spatio-Temporal Omni-Object Video Grounding
von: Yao, Jiali, et al.
Veröffentlicht: (2025) -
Context-Guided Spatio-Temporal Video Grounding
von: Gu, Xin, et al.
Veröffentlicht: (2024) -
Thinking With Bounding Boxes: Enhancing Spatio-Temporal Video Grounding via Reinforcement Fine-Tuning
von: Gu, Xin, et al.
Veröffentlicht: (2025) -
Knowing Your Target: Target-Aware Transformer Makes Better Spatio-Temporal Video Grounding
von: Gu, Xin, et al.
Veröffentlicht: (2025) -
High-Fidelity Image Inpainting with Multimodal Guided GAN Inversion
von: Zhang, Libo, et al.
Veröffentlicht: (2025)