Training-free Video Temporal Grounding using Large-scale Pre-trained Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zheng, Minghang, Cai, Xinhao, Chen, Qingchao, Peng, Yuxin, Liu, Yang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding
von: Zheng, Minghang, et al.
Veröffentlicht: (2026)
von: Zheng, Minghang, et al.
Veröffentlicht: (2026)
Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding
von: Zheng, Minghang, et al.
Veröffentlicht: (2025)
von: Zheng, Minghang, et al.
Veröffentlicht: (2025)
ResVG: Enhancing Relation and Semantic Understanding in Multiple Instances for Visual Grounding
von: Zheng, Minghang, et al.
Veröffentlicht: (2024)
von: Zheng, Minghang, et al.
Veröffentlicht: (2024)
Large-scale Pre-training for Grounded Video Caption Generation
von: Kazakos, Evangelos, et al.
Veröffentlicht: (2025)
von: Kazakos, Evangelos, et al.
Veröffentlicht: (2025)
InteractMove: Text-Controlled Human-Object Interaction Generation in 3D Scenes with Movable Objects
von: Cai, Xinhao, et al.
Veröffentlicht: (2025)
von: Cai, Xinhao, et al.
Veröffentlicht: (2025)
Advancing 3D Scene Understanding with MV-ScanQA Multi-View Reasoning Evaluation and TripAlign Pre-training Dataset
von: Mo, Wentao, et al.
Veröffentlicht: (2025)
von: Mo, Wentao, et al.
Veröffentlicht: (2025)
Towards 3D heart mesh generation using contactless radar imaging and physics-informed neural network
von: Li, Jinye, et al.
Veröffentlicht: (2026)
von: Li, Jinye, et al.
Veröffentlicht: (2026)
Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement
von: Gao, Jiayi, et al.
Veröffentlicht: (2025)
von: Gao, Jiayi, et al.
Veröffentlicht: (2025)
Incorporating Pre-training Data Matters in Unsupervised Domain Adaptation
von: Xu, Yinsong, et al.
Veröffentlicht: (2023)
von: Xu, Yinsong, et al.
Veröffentlicht: (2023)
Temporal-Consistent Video Restoration with Pre-trained Diffusion Models
von: Wang, Hengkang, et al.
Veröffentlicht: (2025)
von: Wang, Hengkang, et al.
Veröffentlicht: (2025)
OmnimatteZero: Fast Training-free Omnimatte with Pre-trained Video Diffusion Models
von: Samuel, Dvir, et al.
Veröffentlicht: (2025)
von: Samuel, Dvir, et al.
Veröffentlicht: (2025)
Fake It Right: Injecting Anatomical Logic into Synthetic Supervised Pre-training for Medical Segmentation
von: Tang, Jiaqi, et al.
Veröffentlicht: (2026)
von: Tang, Jiaqi, et al.
Veröffentlicht: (2026)
Foundation Model for Endoscopy Video Analysis via Large-scale Self-supervised Pre-train
von: Wang, Zhao, et al.
Veröffentlicht: (2023)
von: Wang, Zhao, et al.
Veröffentlicht: (2023)
Diff-BGM: A Diffusion Model for Video Background Music Generation
von: Li, Sizhe, et al.
Veröffentlicht: (2024)
von: Li, Sizhe, et al.
Veröffentlicht: (2024)
A Temporal Modeling Framework for Video Pre-Training on Video Instance Segmentation
von: Zhong, Qing, et al.
Veröffentlicht: (2025)
von: Zhong, Qing, et al.
Veröffentlicht: (2025)
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection
von: Zhuang, Weijun, et al.
Veröffentlicht: (2025)
von: Zhuang, Weijun, et al.
Veröffentlicht: (2025)
ZeroI2V: Zero-Cost Adaptation of Pre-trained Transformers from Image to Video
von: Li, Xinhao, et al.
Veröffentlicht: (2023)
von: Li, Xinhao, et al.
Veröffentlicht: (2023)
AutoTVG: A New Vision-language Pre-training Paradigm for Temporal Video Grounding
von: Zhang, Xing, et al.
Veröffentlicht: (2024)
von: Zhang, Xing, et al.
Veröffentlicht: (2024)
A Survey on Video Temporal Grounding with Multimodal Large Language Model
von: Wu, Jianlong, et al.
Veröffentlicht: (2025)
von: Wu, Jianlong, et al.
Veröffentlicht: (2025)
Open-Vocabulary HOI Detection with Interaction-aware Prompt and Concept Calibration
von: Lei, Ting, et al.
Veröffentlicht: (2025)
von: Lei, Ting, et al.
Veröffentlicht: (2025)
Investigating Domain Gaps for Indoor 3D Object Detection
von: Zhao, Zijing, et al.
Veröffentlicht: (2025)
von: Zhao, Zijing, et al.
Veröffentlicht: (2025)
TiFRe: Text-guided Video Frame Reduction for Efficient Video Multi-modal Large Language Models
von: Zheng, Xiangtian, et al.
Veröffentlicht: (2026)
von: Zheng, Xiangtian, et al.
Veröffentlicht: (2026)
Training-free Online Video Step Grounding
von: Zanella, Luca, et al.
Veröffentlicht: (2025)
von: Zanella, Luca, et al.
Veröffentlicht: (2025)
Med-GLIP: Advancing Medical Language-Image Pre-training with Large-scale Grounded Dataset
von: Deng, Ziye, et al.
Veröffentlicht: (2025)
von: Deng, Ziye, et al.
Veröffentlicht: (2025)
An Empirical Study of Parameter Efficient Fine-tuning on Vision-Language Pre-train Model
von: Tian, Yuxin, et al.
Veröffentlicht: (2024)
von: Tian, Yuxin, et al.
Veröffentlicht: (2024)
GroundVTS: Visual Token Sampling in Multimodal Large Language Models for Video Temporal Grounding
von: Fan, Rong, et al.
Veröffentlicht: (2026)
von: Fan, Rong, et al.
Veröffentlicht: (2026)
Repurposing Pre-trained Video Diffusion Models for Event-based Video Interpolation
von: Chen, Jingxi, et al.
Veröffentlicht: (2024)
von: Chen, Jingxi, et al.
Veröffentlicht: (2024)
TTSA3R: Training-Free Temporal-Spatial Adaptive Persistent State for Streaming 3D Reconstruction
von: Zheng, Zhijie, et al.
Veröffentlicht: (2026)
von: Zheng, Zhijie, et al.
Veröffentlicht: (2026)
An Evaluation of Large Pre-Trained Models for Gesture Recognition using Synthetic Videos
von: Reddy, Arun, et al.
Veröffentlicht: (2024)
von: Reddy, Arun, et al.
Veröffentlicht: (2024)
Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation
von: Li, Jiaze, et al.
Veröffentlicht: (2026)
von: Li, Jiaze, et al.
Veröffentlicht: (2026)
ChatVTG: Video Temporal Grounding via Chat with Video Dialogue Large Language Models
von: Qu, Mengxue, et al.
Veröffentlicht: (2024)
von: Qu, Mengxue, et al.
Veröffentlicht: (2024)
Training-Free Semantic Video Composition via Pre-trained Diffusion Model
von: Guo, Jiaqi, et al.
Veröffentlicht: (2024)
von: Guo, Jiaqi, et al.
Veröffentlicht: (2024)
TriDP-PTM: a three-stage distortion-perception tradeoff guides the pre-training model for radar cardiac sensing
von: Li, Jinye, et al.
Veröffentlicht: (2026)
von: Li, Jinye, et al.
Veröffentlicht: (2026)
DGHMesh: A Large-scale Dual-radar mmWave Dataset and Generalization-focused Benchmark for Human Mesh Reconstruction
von: Guo, Rongxiao, et al.
Veröffentlicht: (2026)
von: Guo, Rongxiao, et al.
Veröffentlicht: (2026)
Scanning Only Once: An End-to-end Framework for Fast Temporal Grounding in Long Videos
von: Pan, Yulin, et al.
Veröffentlicht: (2023)
von: Pan, Yulin, et al.
Veröffentlicht: (2023)
SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability
von: Wang, Jiankang, et al.
Veröffentlicht: (2025)
von: Wang, Jiankang, et al.
Veröffentlicht: (2025)
Efficient Temporal Extrapolation of Multimodal Large Language Models with Temporal Grounding Bridge
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
SPAST: Arbitrary Style Transfer with Style Priors via Pre-trained Large-scale Model
von: Zhang, Zhanjie, et al.
Veröffentlicht: (2025)
von: Zhang, Zhanjie, et al.
Veröffentlicht: (2025)
Training-free Temporal Object Tracking in Surgical Videos
von: Koley, Subhadeep, et al.
Veröffentlicht: (2026)
von: Koley, Subhadeep, et al.
Veröffentlicht: (2026)
Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding
von: Zheng, Zelin, et al.
Veröffentlicht: (2026)
von: Zheng, Zelin, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding
von: Zheng, Minghang, et al.
Veröffentlicht: (2026) -
Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding
von: Zheng, Minghang, et al.
Veröffentlicht: (2025) -
ResVG: Enhancing Relation and Semantic Understanding in Multiple Instances for Visual Grounding
von: Zheng, Minghang, et al.
Veröffentlicht: (2024) -
Large-scale Pre-training for Grounded Video Caption Generation
von: Kazakos, Evangelos, et al.
Veröffentlicht: (2025) -
InteractMove: Text-Controlled Human-Object Interaction Generation in 3D Scenes with Movable Objects
von: Cai, Xinhao, et al.
Veröffentlicht: (2025)