Invert4TVG: A Temporal Video Grounding Framework with Inversion Tasks Preserving Action Understanding Ability
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Zhaoyu, Lin, Hongnan, Nie, Yongwei, Ma, Fei, Xu, Xuemiao, Yu, Fei, Long, Chengjiang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TAR-TVG: Enhancing VLMs with Timestamp Anchor-Constrained Reasoning for Temporal Video Grounding
by: Guo, Chaohong, et al.
Published: (2025)
by: Guo, Chaohong, et al.
Published: (2025)
T2SGrid: Temporal-to-Spatial Gridification for Video Temporal Grounding
by: Guo, Chaohong, et al.
Published: (2026)
by: Guo, Chaohong, et al.
Published: (2026)
Incorporating Test-Time Optimization into Training with Dual Networks for Human Mesh Recovery
by: Nie, Yongwei, et al.
Published: (2024)
by: Nie, Yongwei, et al.
Published: (2024)
AutoTVG: A New Vision-language Pre-training Paradigm for Temporal Video Grounding
by: Zhang, Xing, et al.
Published: (2024)
by: Zhang, Xing, et al.
Published: (2024)
Interleaving One-Class and Weakly-Supervised Models with Adaptive Thresholding for Unsupervised Video Anomaly Detection
by: Nie, Yongwei, et al.
Published: (2024)
by: Nie, Yongwei, et al.
Published: (2024)
Enhancing Long Video Question Answering with Scene-Localized Frame Grouping
by: Yang, Xuyi, et al.
Published: (2025)
by: Yang, Xuyi, et al.
Published: (2025)
RecDreamer: Consistent Text-to-3D Generation via Uniform Score Distillation
by: Zheng, Chenxi, et al.
Published: (2025)
by: Zheng, Chenxi, et al.
Published: (2025)
Chinese expert consensus on an innovative patient‐centered approach to diagnosis and treatment of cancer
by: Hongnan Mo, et al.
Published: (2024)
by: Hongnan Mo, et al.
Published: (2024)
Multi-RoI Human Mesh Recovery with Camera Consistency and Contrastive Losses
by: Nie, Yongwei, et al.
Published: (2024)
by: Nie, Yongwei, et al.
Published: (2024)
Temperature fluctuations in mesoscopic systems
by: Fei, Zhaoyu, et al.
Published: (2023)
by: Fei, Zhaoyu, et al.
Published: (2023)
Berry-Phase-Induced Chirality in Thermodynamics
by: Fei, Zhaoyu, et al.
Published: (2026)
by: Fei, Zhaoyu, et al.
Published: (2026)
TGPO: Temporal Grounded Policy Optimization for Signal Temporal Logic Tasks
by: Meng, Yue, et al.
Published: (2025)
by: Meng, Yue, et al.
Published: (2025)
Quantum stochastic thermodynamics: A semiclassical theory in phase space
by: Fei, Zhaoyu
Published: (2023)
by: Fei, Zhaoyu
Published: (2023)
Registration is a Powerful Rotation-Invariance Learner for 3D Anomaly Detection
by: Yu, Yuyang, et al.
Published: (2025)
by: Yu, Yuyang, et al.
Published: (2025)
TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability
by: Chen, Shimin, et al.
Published: (2024)
by: Chen, Shimin, et al.
Published: (2024)
StakeBench: Evaluating Language Understanding Grounded in Market Commitment
by: Pei, Yunhua, et al.
Published: (2026)
by: Pei, Yunhua, et al.
Published: (2026)
FunduSAM: A Specialized Deep Learning Model for Enhanced Optic Disc and Cup Segmentation in Fundus Images
by: Yu, Jinchen, et al.
Published: (2025)
by: Yu, Jinchen, et al.
Published: (2025)
VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
by: Wang, Shihao, et al.
Published: (2025)
by: Wang, Shihao, et al.
Published: (2025)
Temporal2Seq: A Unified Framework for Temporal Video Understanding Tasks
by: Yang, Min, et al.
Published: (2024)
by: Yang, Min, et al.
Published: (2024)
Impact of HER2‐low expression on the efficacy of endocrine therapy with or without CDK4/6 inhibitor in HR‐positive/HER2‐negative metastatic breast cancer: A prospective study
by: Yun Wu, et al.
Published: (2024)
by: Yun Wu, et al.
Published: (2024)
Grounding is All You Need? Dual Temporal Grounding for Video Dialog
by: Qin, You, et al.
Published: (2024)
by: Qin, You, et al.
Published: (2024)
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding
by: Luo, Fuwen, et al.
Published: (2025)
by: Luo, Fuwen, et al.
Published: (2025)
What Happens Before Decoding? Prefill Determines GUI Grounding in VLMs
by: Lin, Jiaping, et al.
Published: (2026)
by: Lin, Jiaping, et al.
Published: (2026)
TVG: A Training-free Transition Video Generation Method with Diffusion Models
by: Zhang, Rui, et al.
Published: (2024)
by: Zhang, Rui, et al.
Published: (2024)
E4S: Fine-grained Face Swapping via Editing With Regional GAN Inversion
by: Li, Maomao, et al.
Published: (2023)
by: Li, Maomao, et al.
Published: (2023)
LeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question Answering
by: Dong, Xinxin, et al.
Published: (2025)
by: Dong, Xinxin, et al.
Published: (2025)
Metronomic Chemotherapy in Breast Cancer: Unleashing the Potential of Combination Regimens
by: Jiaxuan Liu, et al.
Published: (2025)
by: Jiaxuan Liu, et al.
Published: (2025)
Fewer Tokens and Fewer Videos: Extending Video Understanding Abilities in Large Vision-Language Models
by: Chen, Shimin, et al.
Published: (2024)
by: Chen, Shimin, et al.
Published: (2024)
T*: Re-thinking Temporal Search for Long-Form Video Understanding
by: Ye, Jinhui, et al.
Published: (2025)
by: Ye, Jinhui, et al.
Published: (2025)
Universal Visuo-Tactile Video Understanding for Embodied Interaction
by: Xie, Yifan, et al.
Published: (2025)
by: Xie, Yifan, et al.
Published: (2025)
Task-Specific Distance Correlation Matching for Few-Shot Action Recognition
by: Long, Fei, et al.
Published: (2025)
by: Long, Fei, et al.
Published: (2025)
VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos
by: Liu, Wenqi, et al.
Published: (2026)
by: Liu, Wenqi, et al.
Published: (2026)
A Survey on Video Temporal Grounding with Multimodal Large Language Model
by: Wu, Jianlong, et al.
Published: (2025)
by: Wu, Jianlong, et al.
Published: (2025)
A Video is Worth 256 Bases: Spatial-Temporal Expectation-Maximization Inversion for Zero-Shot Video Editing
by: Li, Maomao, et al.
Published: (2023)
by: Li, Maomao, et al.
Published: (2023)
(Table 2) Isotopic analyses from TVG 97
by: Greinert, Jens, et al.
Published: (2002)
by: Greinert, Jens, et al.
Published: (2002)
ToG-Bench: Task-Oriented Spatio-Temporal Grounding in Egocentric Videos
by: Xu, Qi'ao, et al.
Published: (2025)
by: Xu, Qi'ao, et al.
Published: (2025)
Repurposing 2D Diffusion Models with Gaussian Atlas for 3D Generation
by: Xiang, Tiange, et al.
Published: (2025)
by: Xiang, Tiange, et al.
Published: (2025)
Motion Keyframe Interpolation for Any Human Skeleton via Temporally Consistent Point Cloud Sampling and Reconstruction
by: Mo, Clinton, et al.
Published: (2024)
by: Mo, Clinton, et al.
Published: (2024)
Admissible Hermitian-Yang-Mills connections over normal varieties
by: Chen, Xuemiao
Published: (2022)
by: Chen, Xuemiao
Published: (2022)
On Vafa-Witten equations over Kaehler manifolds
by: Chen, Xuemiao
Published: (2023)
by: Chen, Xuemiao
Published: (2023)
Similar Items
-
TAR-TVG: Enhancing VLMs with Timestamp Anchor-Constrained Reasoning for Temporal Video Grounding
by: Guo, Chaohong, et al.
Published: (2025) -
T2SGrid: Temporal-to-Spatial Gridification for Video Temporal Grounding
by: Guo, Chaohong, et al.
Published: (2026) -
Incorporating Test-Time Optimization into Training with Dual Networks for Human Mesh Recovery
by: Nie, Yongwei, et al.
Published: (2024) -
AutoTVG: A New Vision-language Pre-training Paradigm for Temporal Video Grounding
by: Zhang, Xing, et al.
Published: (2024) -
Interleaving One-Class and Weakly-Supervised Models with Adaptive Thresholding for Unsupervised Video Anomaly Detection
by: Nie, Yongwei, et al.
Published: (2024)