Enregistré dans:
| Auteurs principaux: | Tu, Xuezhen, Wu, Jingyu, Kang, Fangyu, Nong, Qingpeng, Zhang, Kaijin, Niu, Chaoyue, Wu, Fan |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2604.08014 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Context-Guided Spatio-Temporal Video Grounding
par: Gu, Xin, et autres
Publié: (2024)
par: Gu, Xin, et autres
Publié: (2024)
Towards Long-Form Spatio-Temporal Video Grounding
par: Gu, Xin, et autres
Publié: (2026)
par: Gu, Xin, et autres
Publié: (2026)
Thinking With Bounding Boxes: Enhancing Spatio-Temporal Video Grounding via Reinforcement Fine-Tuning
par: Gu, Xin, et autres
Publié: (2025)
par: Gu, Xin, et autres
Publié: (2025)
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning
par: Qiu, Tianheng, et autres
Publié: (2025)
par: Qiu, Tianheng, et autres
Publié: (2025)
OmniSTVG: Toward Spatio-Temporal Omni-Object Video Grounding
par: Yao, Jiali, et autres
Publié: (2025)
par: Yao, Jiali, et autres
Publié: (2025)
Enhancing Video-Language Representations with Structural Spatio-Temporal Alignment
par: Fei, Hao, et autres
Publié: (2024)
par: Fei, Hao, et autres
Publié: (2024)
OmniGround: A Comprehensive Spatio-Temporal Grounding Benchmark for Real-World Complex Scenarios
par: Gao, Hong, et autres
Publié: (2025)
par: Gao, Hong, et autres
Publié: (2025)
SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability
par: Wang, Jiankang, et autres
Publié: (2025)
par: Wang, Jiankang, et autres
Publié: (2025)
UBATrack: Spatio-Temporal State Space Model for General Multi-Modal Tracking
par: Liang, Qihua, et autres
Publié: (2026)
par: Liang, Qihua, et autres
Publié: (2026)
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding
par: Wasim, Syed Talal, et autres
Publié: (2023)
par: Wasim, Syed Talal, et autres
Publié: (2023)
Video-Language Alignment via Spatio-Temporal Graph Transformer
par: Zhang, Shi-Xue, et autres
Publié: (2024)
par: Zhang, Shi-Xue, et autres
Publié: (2024)
VideoMolmo: Spatio-Temporal Grounding Meets Pointing
par: Ahmad, Ghazi Shazan, et autres
Publié: (2025)
par: Ahmad, Ghazi Shazan, et autres
Publié: (2025)
SpatioTemporal Learning for Human Pose Estimation in Sparsely-Labeled Videos
par: Jiao, Yingying, et autres
Publié: (2025)
par: Jiao, Yingying, et autres
Publié: (2025)
SegDebias: Test-Time Bias Mitigation for ViT-Based CLIP via Segmentation
par: Wu, Fangyu, et autres
Publié: (2025)
par: Wu, Fangyu, et autres
Publié: (2025)
STDR: Spatio-Temporal Decoupling for Real-Time Dynamic Scene Rendering
par: Li, Zehao, et autres
Publié: (2025)
par: Li, Zehao, et autres
Publié: (2025)
CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering
par: Zhang, Mingfang, et autres
Publié: (2026)
par: Zhang, Mingfang, et autres
Publié: (2026)
Rethinking the Spatio-Temporal Alignment of End-to-End 3D Perception
par: Li, Xiaoyu, et autres
Publié: (2025)
par: Li, Xiaoyu, et autres
Publié: (2025)
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding
par: Gao, Shida, et autres
Publié: (2025)
par: Gao, Shida, et autres
Publié: (2025)
Decoupling Spatio-Temporal Adapter for Fine-Grained Badminton Action Localization
par: Wang, Tianyu, et autres
Publié: (2026)
par: Wang, Tianyu, et autres
Publié: (2026)
VideoMamba: Spatio-Temporal Selective State Space Model
par: Park, Jinyoung, et autres
Publié: (2024)
par: Park, Jinyoung, et autres
Publié: (2024)
Static and Dynamic Graph Alignment Network for Temporal Video Grounding
par: Hu, Zhanjie, et autres
Publié: (2026)
par: Hu, Zhanjie, et autres
Publié: (2026)
Infinite-ID: Identity-preserved Personalization via ID-semantics Decoupling Paradigm
par: Wu, Yi, et autres
Publié: (2024)
par: Wu, Yi, et autres
Publié: (2024)
APVR: Hour-Level Long Video Understanding with Adaptive Pivot Visual Information Retrieval
par: Gao, Hong, et autres
Publié: (2025)
par: Gao, Hong, et autres
Publié: (2025)
Multimodal Spatio-temporal Graph Learning for Alignment-free RGBT Video Object Detection
par: Wang, Qishun, et autres
Publié: (2025)
par: Wang, Qishun, et autres
Publié: (2025)
Knowing Your Target: Target-Aware Transformer Makes Better Spatio-Temporal Video Grounding
par: Gu, Xin, et autres
Publié: (2025)
par: Gu, Xin, et autres
Publié: (2025)
DeRA: Decoupled Representation Alignment for Video Tokenization
par: Guo, Pengbo, et autres
Publié: (2025)
par: Guo, Pengbo, et autres
Publié: (2025)
NCSTR: Node-Centric Decoupled Spatio-Temporal Reasoning for Video-based Human Pose Estimation
par: Huynh, Quang Dang, et autres
Publié: (2026)
par: Huynh, Quang Dang, et autres
Publié: (2026)
Unleashing the Potential of Multimodal LLMs for Zero-Shot Spatio-Temporal Video Grounding
par: Yang, Zaiquan, et autres
Publié: (2025)
par: Yang, Zaiquan, et autres
Publié: (2025)
Contextual Self-paced Learning for Weakly Supervised Spatio-Temporal Video Grounding
par: Kumar, Akash, et autres
Publié: (2025)
par: Kumar, Akash, et autres
Publié: (2025)
Know-Show: Benchmarking Video-Language Models on Spatio-Temporal Grounded Reasoning
par: Sugandhika, Chinthani, et autres
Publié: (2025)
par: Sugandhika, Chinthani, et autres
Publié: (2025)
MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling
par: Zhang, Yue, et autres
Publié: (2024)
par: Zhang, Yue, et autres
Publié: (2024)
SpatioTemporal Difference Network for Video Depth Super-Resolution
par: Wang, Zhengxue, et autres
Publié: (2025)
par: Wang, Zhengxue, et autres
Publié: (2025)
Enhanced Textual Feature Extraction for Visual Question Answering: A Simple Convolutional Approach
par: Zhang, Zhilin, et autres
Publié: (2024)
par: Zhang, Zhilin, et autres
Publié: (2024)
STAF: 3D Human Mesh Recovery from Video with Spatio-Temporal Alignment Fusion
par: Yao, Wei, et autres
Publié: (2024)
par: Yao, Wei, et autres
Publié: (2024)
TRACE: Temporal Grounding Video LLM via Causal Event Modeling
par: Guo, Yongxin, et autres
Publié: (2024)
par: Guo, Yongxin, et autres
Publié: (2024)
Training-Free Spatio-temporal Decoupled Reasoning Video Segmentation with Adaptive Object Memory
par: Zhu, Zhengtong, et autres
Publié: (2026)
par: Zhu, Zhengtong, et autres
Publié: (2026)
Feature Alignment Determines Fusion Strategy: A Comparative Study of Cross-Attention and Concatenation in Multimodal Learning
par: Zhou, Zhiqiang, et autres
Publié: (2026)
par: Zhou, Zhiqiang, et autres
Publié: (2026)
Adaptive Routing of Text-to-Image Generation Requests Between Large Cloud Model and Light-Weight Edge Model
par: Xin, Zewei, et autres
Publié: (2024)
par: Xin, Zewei, et autres
Publié: (2024)
Decoupled Spatio-Temporal Consistency Learning for Self-Supervised Tracking
par: Zheng, Yaozong, et autres
Publié: (2025)
par: Zheng, Yaozong, et autres
Publié: (2025)
ToG-Bench: Task-Oriented Spatio-Temporal Grounding in Egocentric Videos
par: Xu, Qi'ao, et autres
Publié: (2025)
par: Xu, Qi'ao, et autres
Publié: (2025)
Documents similaires
-
Context-Guided Spatio-Temporal Video Grounding
par: Gu, Xin, et autres
Publié: (2024) -
Towards Long-Form Spatio-Temporal Video Grounding
par: Gu, Xin, et autres
Publié: (2026) -
Thinking With Bounding Boxes: Enhancing Spatio-Temporal Video Grounding via Reinforcement Fine-Tuning
par: Gu, Xin, et autres
Publié: (2025) -
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning
par: Qiu, Tianheng, et autres
Publié: (2025) -
OmniSTVG: Toward Spatio-Temporal Omni-Object Video Grounding
par: Yao, Jiali, et autres
Publié: (2025)