VTimeCoT: Thinking by Drawing for Video Temporal Grounding and Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Jinglei, Guo, Yuanfan, Potamias, Rolandos Alexandros, Deng, Jiankang, Xu, Hang, Ma, Chao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HaWoR: World-Space Hand Motion Reconstruction from Egocentric Videos
by: Zhang, Jinglei, et al.
Published: (2025)
by: Zhang, Jinglei, et al.
Published: (2025)
WiLoR: End-to-end 3D Hand Localization and Reconstruction in-the-wild
by: Potamias, Rolandos Alexandros, et al.
Published: (2024)
by: Potamias, Rolandos Alexandros, et al.
Published: (2024)
SAGS: Structure-Aware 3D Gaussian Splatting
by: Ververas, Evangelos, et al.
Published: (2024)
by: Ververas, Evangelos, et al.
Published: (2024)
ImHead: A Large-scale Implicit Morphable Model for Localized Head Modeling
by: Potamias, Rolandos Alexandros, et al.
Published: (2025)
by: Potamias, Rolandos Alexandros, et al.
Published: (2025)
Do You See What I Am Pointing At? Gesture-Based Egocentric Video Question Answering
by: Choi, Yura, et al.
Published: (2026)
by: Choi, Yura, et al.
Published: (2026)
Dex2HOI: Dexterous Bimanual Two-Object Interaction Generation
by: Pratikaki, Chrysa, et al.
Published: (2026)
by: Pratikaki, Chrysa, et al.
Published: (2026)
Signs as Tokens: A Retrieval-Enhanced Multilingual Sign Language Generator
by: Zuo, Ronglai, et al.
Published: (2024)
by: Zuo, Ronglai, et al.
Published: (2024)
CEDex: Cross-Embodiment Dexterous Grasp Generation at Scale from Human-like Contact Representations
by: Wu, Zhiyuan, et al.
Published: (2025)
by: Wu, Zhiyuan, et al.
Published: (2025)
Neural Sign Actors: A diffusion model for 3D sign language production from text
by: Baltatzis, Vasileios, et al.
Published: (2023)
by: Baltatzis, Vasileios, et al.
Published: (2023)
Interact2Ar: Full-Body Human-Human Interaction Generation via Autoregressive Diffusion Models
by: Ruiz-Ponce, Pablo, et al.
Published: (2025)
by: Ruiz-Ponce, Pablo, et al.
Published: (2025)
MaDiS: Taming Masked Diffusion Language Models for Sign Language Generation
by: Zuo, Ronglai, et al.
Published: (2026)
by: Zuo, Ronglai, et al.
Published: (2026)
HO-Flow: Generalizable Hand-Object Interaction Generation with Latent Flow Matching
by: Chen, Zerui, et al.
Published: (2026)
by: Chen, Zerui, et al.
Published: (2026)
Design2Cloth: 3D Cloth Generation from 2D Masks
by: Zheng, Jiali, et al.
Published: (2024)
by: Zheng, Jiali, et al.
Published: (2024)
STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits
by: Papantoniou, Foivos Paraperas, et al.
Published: (2025)
by: Papantoniou, Foivos Paraperas, et al.
Published: (2025)
ZeroGS: Training 3D Gaussian Splatting from Unposed Images
by: Chen, Yu, et al.
Published: (2024)
by: Chen, Yu, et al.
Published: (2024)
ShapeFusion: A 3D diffusion model for localized shape editing
by: Potamias, Rolandos Alexandros, et al.
Published: (2024)
by: Potamias, Rolandos Alexandros, et al.
Published: (2024)
HORT: Monocular Hand-held Objects Reconstruction with Transformers
by: Chen, Zerui, et al.
Published: (2025)
by: Chen, Zerui, et al.
Published: (2025)
Arc2Avatar: Generating Expressive 3D Avatars from a Single Image via ID Guidance
by: Gerogiannis, Dimitrios, et al.
Published: (2025)
by: Gerogiannis, Dimitrios, et al.
Published: (2025)
DreamCAD: Scaling Multi-modal CAD Generation using Differentiable Parametric Surfaces
by: Khan, Mohammad Sadil, et al.
Published: (2026)
by: Khan, Mohammad Sadil, et al.
Published: (2026)
StableHand: Quality-Aware Flow Matching for World-Space Dual-Hand Motion Estimation from Egocentric Video
by: Zeng, Huajian, et al.
Published: (2026)
by: Zeng, Huajian, et al.
Published: (2026)
Locally Adaptive Neural 3D Morphable Models
by: Tarasiou, Michail, et al.
Published: (2024)
by: Tarasiou, Michail, et al.
Published: (2024)
AnimateMe: 4D Facial Expressions via Diffusion Models
by: Gerogiannis, Dimitrios, et al.
Published: (2024)
by: Gerogiannis, Dimitrios, et al.
Published: (2024)
SlowFocus: Enhancing Fine-grained Temporal Understanding in Video LLM
by: Nie, Ming, et al.
Published: (2026)
by: Nie, Ming, et al.
Published: (2026)
AnyHand: A Large-Scale Synthetic Dataset for RGB(-D) Hand Pose Estimation
by: Si, Chen, et al.
Published: (2026)
by: Si, Chen, et al.
Published: (2026)
Brick-Diffusion: Generating Long Videos with Brick-to-Wall Denoising
by: Yuan, Yunlong, et al.
Published: (2025)
by: Yuan, Yunlong, et al.
Published: (2025)
EgoGrasp: World-Space Hand-Object Interaction Estimation from Egocentric Videos
by: Fu, Hongming, et al.
Published: (2026)
by: Fu, Hongming, et al.
Published: (2026)
Self-Adaptive Reality-Guided Diffusion for Artifact-Free Super-Resolution
by: Zheng, Qingping, et al.
Published: (2024)
by: Zheng, Qingping, et al.
Published: (2024)
VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos
by: Liu, Wenqi, et al.
Published: (2026)
by: Liu, Wenqi, et al.
Published: (2026)
GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking
by: Cheng, Zixu, et al.
Published: (2026)
by: Cheng, Zixu, et al.
Published: (2026)
VideoCoF: Unified Video Editing with Temporal Reasoner
by: Yang, Xiangpeng, et al.
Published: (2025)
by: Yang, Xiangpeng, et al.
Published: (2025)
T2SGrid: Temporal-to-Spatial Gridification for Video Temporal Grounding
by: Guo, Chaohong, et al.
Published: (2026)
by: Guo, Chaohong, et al.
Published: (2026)
TAR-TVG: Enhancing VLMs with Timestamp Anchor-Constrained Reasoning for Temporal Video Grounding
by: Guo, Chaohong, et al.
Published: (2025)
by: Guo, Chaohong, et al.
Published: (2025)
SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability
by: Wang, Jiankang, et al.
Published: (2025)
by: Wang, Jiankang, et al.
Published: (2025)
GLaVE-Cap: Global-Local Aligned Video Captioning with Vision Expert Integration
by: Xu, Wan, et al.
Published: (2025)
by: Xu, Wan, et al.
Published: (2025)
PanGu-Draw: Advancing Resource-Efficient Text-to-Image Synthesis with Time-Decoupled Training and Reusable Coop-Diffusion
by: Lu, Guansong, et al.
Published: (2023)
by: Lu, Guansong, et al.
Published: (2023)
Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding
by: Chen, Houlun, et al.
Published: (2026)
by: Chen, Houlun, et al.
Published: (2026)
Thinking With Bounding Boxes: Enhancing Spatio-Temporal Video Grounding via Reinforcement Fine-Tuning
by: Gu, Xin, et al.
Published: (2025)
by: Gu, Xin, et al.
Published: (2025)
Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding
by: Zheng, Zelin, et al.
Published: (2026)
by: Zheng, Zelin, et al.
Published: (2026)
Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding
by: Hu, Pengfei, et al.
Published: (2025)
by: Hu, Pengfei, et al.
Published: (2025)
EasyControl: Transfer ControlNet to Video Diffusion for Controllable Generation and Interpolation
by: Wang, Cong, et al.
Published: (2024)
by: Wang, Cong, et al.
Published: (2024)
Similar Items
-
HaWoR: World-Space Hand Motion Reconstruction from Egocentric Videos
by: Zhang, Jinglei, et al.
Published: (2025) -
WiLoR: End-to-end 3D Hand Localization and Reconstruction in-the-wild
by: Potamias, Rolandos Alexandros, et al.
Published: (2024) -
SAGS: Structure-Aware 3D Gaussian Splatting
by: Ververas, Evangelos, et al.
Published: (2024) -
ImHead: A Large-scale Implicit Morphable Model for Localized Head Modeling
by: Potamias, Rolandos Alexandros, et al.
Published: (2025) -
Do You See What I Am Pointing At? Gesture-Based Egocentric Video Question Answering
by: Choi, Yura, et al.
Published: (2026)