Universal Video Temporal Grounding with Generative Multi-modal Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Zeqian, Di, Shangzhe, Zhai, Zhonghua, Huang, Weilin, Wang, Yanfeng, Xie, Weidi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Revisiting Multi-Task Visual Representation Learning
by: Di, Shangzhe, et al.
Published: (2026)
by: Di, Shangzhe, et al.
Published: (2026)
Grounded Question-Answering in Long Egocentric Videos
by: Di, Shangzhe, et al.
Published: (2023)
by: Di, Shangzhe, et al.
Published: (2023)
Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos
by: Chen, Qirui, et al.
Published: (2024)
by: Chen, Qirui, et al.
Published: (2024)
Multi-Sentence Grounding for Long-term Instructional Video
by: Li, Zeqian, et al.
Published: (2023)
by: Li, Zeqian, et al.
Published: (2023)
Learning Streaming Video Representation via Multitask Training
by: Yan, Yibin, et al.
Published: (2025)
by: Yan, Yibin, et al.
Published: (2025)
Enhancing Video-LLM Reasoning via Agent-of-Thoughts Distillation
by: Shi, Yudi, et al.
Published: (2024)
by: Shi, Yudi, et al.
Published: (2024)
DenseStep2M: A Scalable, Training-Free Pipeline for Dense Instructional Video Annotation
by: Ge, Mingji, et al.
Published: (2026)
by: Ge, Mingji, et al.
Published: (2026)
Towards Universal Soccer Video Understanding
by: Rao, Jiayuan, et al.
Published: (2024)
by: Rao, Jiayuan, et al.
Published: (2024)
OmniStream: Mastering Perception, Reconstruction and Action in Continuous Streams
by: Yan, Yibin, et al.
Published: (2026)
by: Yan, Yibin, et al.
Published: (2026)
Weaver: End-to-End Agentic System Training for Video Interleaved Reasoning
by: Shi, Yudi, et al.
Published: (2026)
by: Shi, Yudi, et al.
Published: (2026)
Turbo: Informativity-Driven Acceleration Plug-In for Vision-Language Large Models
by: Ju, Chen, et al.
Published: (2024)
by: Ju, Chen, et al.
Published: (2024)
RadGenome-Chest CT: A Grounded Vision-Language Dataset for Chest CT Analysis
by: Zhang, Xiaoman, et al.
Published: (2024)
by: Zhang, Xiaoman, et al.
Published: (2024)
SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability
by: Wang, Jiankang, et al.
Published: (2025)
by: Wang, Jiankang, et al.
Published: (2025)
End-to-End Dense Video Grounding via Parallel Regression
by: Shi, Fengyuan, et al.
Published: (2021)
by: Shi, Fengyuan, et al.
Published: (2021)
Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models
by: Wang, Haibo, et al.
Published: (2024)
by: Wang, Haibo, et al.
Published: (2024)
MatchTime: Towards Automatic Soccer Game Commentary Generation
by: Rao, Jiayuan, et al.
Published: (2024)
by: Rao, Jiayuan, et al.
Published: (2024)
GroundingGPT:Language Enhanced Multi-modal Grounding Model
by: Li, Zhaowei, et al.
Published: (2024)
by: Li, Zhaowei, et al.
Published: (2024)
Multi-Agent System for Comprehensive Soccer Understanding
by: Rao, Jiayuan, et al.
Published: (2025)
by: Rao, Jiayuan, et al.
Published: (2025)
A Sanity Check on Composed Image Retrieval
by: Liu, Yikun, et al.
Published: (2026)
by: Liu, Yikun, et al.
Published: (2026)
VersaViT: Enhancing MLLM Vision Backbones via Task-Guided Optimization
by: Liu, Yikun, et al.
Published: (2026)
by: Liu, Yikun, et al.
Published: (2026)
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding
by: Gao, Shida, et al.
Published: (2025)
by: Gao, Shida, et al.
Published: (2025)
GroundVTS: Visual Token Sampling in Multimodal Large Language Models for Video Temporal Grounding
by: Fan, Rong, et al.
Published: (2026)
by: Fan, Rong, et al.
Published: (2026)
Q-Ground: Image Quality Grounding with Large Multi-modality Models
by: Chen, Chaofeng, et al.
Published: (2024)
by: Chen, Chaofeng, et al.
Published: (2024)
EchoSight: Advancing Visual-Language Models with Wiki Knowledge
by: Yan, Yibin, et al.
Published: (2024)
by: Yan, Yibin, et al.
Published: (2024)
VISA: Reasoning Video Object Segmentation via Large Language Models
by: Yan, Cilin, et al.
Published: (2024)
by: Yan, Cilin, et al.
Published: (2024)
A Survey on Video Temporal Grounding with Multimodal Large Language Model
by: Wu, Jianlong, et al.
Published: (2025)
by: Wu, Jianlong, et al.
Published: (2025)
Knowledge-enhanced Visual-Language Pretraining for Computational Pathology
by: Zhou, Xiao, et al.
Published: (2024)
by: Zhou, Xiao, et al.
Published: (2024)
ChatVTG: Video Temporal Grounding via Chat with Video Dialogue Large Language Models
by: Qu, Mengxue, et al.
Published: (2024)
by: Qu, Mengxue, et al.
Published: (2024)
CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models
by: Huang, Junming, et al.
Published: (2026)
by: Huang, Junming, et al.
Published: (2026)
EtC: Temporal Boundary Expand then Clarify for Weakly Supervised Video Grounding with Multimodal Large Language Model
by: Li, Guozhang, et al.
Published: (2023)
by: Li, Guozhang, et al.
Published: (2023)
Large Multi-modality Model Assisted AI-Generated Image Quality Assessment
by: Wang, Puyi, et al.
Published: (2024)
by: Wang, Puyi, et al.
Published: (2024)
LamRA: Large Multimodal Model as Your Advanced Retrieval Assistant
by: Liu, Yikun, et al.
Published: (2024)
by: Liu, Yikun, et al.
Published: (2024)
SeedEdit 3.0: Fast and High-Quality Generative Image Editing
by: Wang, Peng, et al.
Published: (2025)
by: Wang, Peng, et al.
Published: (2025)
A General Protocol to Probe Large Vision Models for 3D Physical Understanding
by: Zhan, Guanqi, et al.
Published: (2023)
by: Zhan, Guanqi, et al.
Published: (2023)
Zero-shot Composed Text-Image Retrieval
by: Liu, Yikun, et al.
Published: (2023)
by: Liu, Yikun, et al.
Published: (2023)
TiFRe: Text-guided Video Frame Reduction for Efficient Video Multi-modal Large Language Models
by: Zheng, Xiangtian, et al.
Published: (2026)
by: Zheng, Xiangtian, et al.
Published: (2026)
Intelligent Grimm -- Open-ended Visual Storytelling via Latent Diffusion Models
by: Liu, Chang, et al.
Published: (2023)
by: Liu, Chang, et al.
Published: (2023)
VideoFusion: A Spatio-Temporal Collaborative Network for Multi-modal Video Fusion
by: Tang, Linfeng, et al.
Published: (2025)
by: Tang, Linfeng, et al.
Published: (2025)
UniversalVTG: A Universal and Lightweight Foundation Model for Video Temporal Grounding
by: An, Joungbin, et al.
Published: (2026)
by: An, Joungbin, et al.
Published: (2026)
GMOS: Grounding Moving Object Segmentation in 3D Space and Time
by: Xie, Junyu, et al.
Published: (2026)
by: Xie, Junyu, et al.
Published: (2026)
Similar Items
-
Revisiting Multi-Task Visual Representation Learning
by: Di, Shangzhe, et al.
Published: (2026) -
Grounded Question-Answering in Long Egocentric Videos
by: Di, Shangzhe, et al.
Published: (2023) -
Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos
by: Chen, Qirui, et al.
Published: (2024) -
Multi-Sentence Grounding for Long-term Instructional Video
by: Li, Zeqian, et al.
Published: (2023) -
Learning Streaming Video Representation via Multitask Training
by: Yan, Yibin, et al.
Published: (2025)