SlowFocus: Enhancing Fine-grained Temporal Understanding in Video LLM
Fuente:
arXiv
Saved in:
| Main Authors: | Nie, Ming, Ding, Dan, Wang, Chunwei, Guo, Yuanfan, Han, Jianhua, Xu, Hang, Zhang, Li |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
KFFocus: Highlighting Keyframes for Enhanced Video Understanding
by: Nie, Ming, et al.
Published: (2025)
by: Nie, Ming, et al.
Published: (2025)
Brick-Diffusion: Generating Long Videos with Brick-to-Wall Denoising
by: Yuan, Yunlong, et al.
Published: (2025)
by: Yuan, Yunlong, et al.
Published: (2025)
Towards Unified Multimodal Interleaved Generation via Group Relative Policy Optimization
by: Nie, Ming, et al.
Published: (2026)
by: Nie, Ming, et al.
Published: (2026)
Reason2Drive: Towards Interpretable and Chain-based Reasoning for Autonomous Driving
by: Nie, Ming, et al.
Published: (2023)
by: Nie, Ming, et al.
Published: (2023)
VTimeCoT: Thinking by Drawing for Video Temporal Grounding and Reasoning
by: Zhang, Jinglei, et al.
Published: (2025)
by: Zhang, Jinglei, et al.
Published: (2025)
EasyControl: Transfer ControlNet to Video Diffusion for Controllable Generation and Interpolation
by: Wang, Cong, et al.
Published: (2024)
by: Wang, Cong, et al.
Published: (2024)
Focus Anywhere for Fine-grained Multi-page Document Understanding
by: Liu, Chenglong, et al.
Published: (2024)
by: Liu, Chenglong, et al.
Published: (2024)
TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos
by: Kong, Fanheng, et al.
Published: (2025)
by: Kong, Fanheng, et al.
Published: (2025)
GLaVE-Cap: Global-Local Aligned Video Captioning with Vision Expert Integration
by: Xu, Wan, et al.
Published: (2025)
by: Xu, Wan, et al.
Published: (2025)
ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance
by: Wang, Chunwei, et al.
Published: (2024)
by: Wang, Chunwei, et al.
Published: (2024)
PaMi-VDPO: Mitigating Video Hallucinations by Prompt-Aware Multi-Instance Video Preference Learning
by: Ding, Xinpeng, et al.
Published: (2025)
by: Ding, Xinpeng, et al.
Published: (2025)
Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models
by: Wang, Haibo, et al.
Published: (2024)
by: Wang, Haibo, et al.
Published: (2024)
WTS: A Pedestrian-Centric Traffic Video Dataset for Fine-grained Spatial-Temporal Understanding
by: Kong, Quan, et al.
Published: (2024)
by: Kong, Quan, et al.
Published: (2024)
HumanRefiner: Benchmarking Abnormal Human Generation and Refining with Coarse-to-fine Pose-Reversible Guidance
by: Fang, Guian, et al.
Published: (2024)
by: Fang, Guian, et al.
Published: (2024)
HiLM-D: Enhancing MLLMs with Multi-Scale High-Resolution Details for Autonomous Driving
by: Ding, Xinpeng, et al.
Published: (2023)
by: Ding, Xinpeng, et al.
Published: (2023)
MA-Bench: Towards Fine-grained Micro-Action Understanding
by: Li, Kun, et al.
Published: (2026)
by: Li, Kun, et al.
Published: (2026)
TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
by: Cai, Mu, et al.
Published: (2024)
by: Cai, Mu, et al.
Published: (2024)
FiLA-Video: Spatio-Temporal Compression for Fine-Grained Long Video Understanding
by: Guo, Yanan, et al.
Published: (2025)
by: Guo, Yanan, et al.
Published: (2025)
V-CORE: Temporally Consistent Video Understanding for Video-LLM
by: Kang, Zhengjian, et al.
Published: (2026)
by: Kang, Zhengjian, et al.
Published: (2026)
UNIT: Unifying Image and Text Recognition in One Vision Encoder
by: Zhu, Yi, et al.
Published: (2024)
by: Zhu, Yi, et al.
Published: (2024)
Mesh-RFT: Enhancing Mesh Generation via Fine-grained Reinforcement Fine-Tuning
by: Liu, Jian, et al.
Published: (2025)
by: Liu, Jian, et al.
Published: (2025)
SemHiTok: A Unified Image Tokenizer via Semantic-Guided Hierarchical Codebook for Multimodal Understanding and Generation
by: Chen, Zisheng, et al.
Published: (2025)
by: Chen, Zisheng, et al.
Published: (2025)
Large-scale and Fine-grained Vision-language Pre-training for Enhanced CT Image Understanding
by: Shui, Zhongyi, et al.
Published: (2025)
by: Shui, Zhongyi, et al.
Published: (2025)
FOCUS: Fine-grained Optimization with Semantic Guided Understanding for Pedestrian Attributes Recognition
by: An, Hongyan, et al.
Published: (2025)
by: An, Hongyan, et al.
Published: (2025)
Data-free Knowledge Distillation for Fine-grained Visual Categorization
by: Shao, Renrong, et al.
Published: (2024)
by: Shao, Renrong, et al.
Published: (2024)
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding
by: Han, Yudong, et al.
Published: (2024)
by: Han, Yudong, et al.
Published: (2024)
HiRes-LLaVA: Restoring Fragmentation Input in High-Resolution Large Vision-Language Models
by: Huang, Runhui, et al.
Published: (2024)
by: Huang, Runhui, et al.
Published: (2024)
Skim then Focus: Integrating Contextual and Fine-grained Views for Repetitive Action Counting
by: Zhao, Zhengqi, et al.
Published: (2024)
by: Zhao, Zhengqi, et al.
Published: (2024)
VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video Reasoning
by: Ding, Yang, et al.
Published: (2025)
by: Ding, Yang, et al.
Published: (2025)
VideoExpert: Augmented LLM for Temporal-Sensitive Video Understanding
by: Zhao, Henghao, et al.
Published: (2025)
by: Zhao, Henghao, et al.
Published: (2025)
FocusDiT: Masking Queries in Diffusion Transformers for Fine-grained Image Generation
by: Fang, Xueji, et al.
Published: (2026)
by: Fang, Xueji, et al.
Published: (2026)
From Summary to Action: Enhancing Large Language Models for Complex Tasks with Open World APIs
by: Liu, Yulong, et al.
Published: (2024)
by: Liu, Yulong, et al.
Published: (2024)
T2SGrid: Temporal-to-Spatial Gridification for Video Temporal Grounding
by: Guo, Chaohong, et al.
Published: (2026)
by: Guo, Chaohong, et al.
Published: (2026)
PanGu-Draw: Advancing Resource-Efficient Text-to-Image Synthesis with Time-Decoupled Training and Reusable Coop-Diffusion
by: Lu, Guansong, et al.
Published: (2023)
by: Lu, Guansong, et al.
Published: (2023)
FriendsQA: A New Large-Scale Deep Video Understanding Dataset with Fine-grained Topic Categorization for Story Videos
by: Wu, Zhengqian, et al.
Published: (2024)
by: Wu, Zhengqian, et al.
Published: (2024)
VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM
by: Yuan, Yuqian, et al.
Published: (2024)
by: Yuan, Yuqian, et al.
Published: (2024)
FC-VFI: Faithful and Consistent Video Frame Interpolation for High-FPS Slow Motion Video Generation
by: Ding, Ganggui, et al.
Published: (2026)
by: Ding, Ganggui, et al.
Published: (2026)
LION-FS: Fast & Slow Video-Language Thinker as Online Video Assistant
by: Li, Wei, et al.
Published: (2025)
by: Li, Wei, et al.
Published: (2025)
VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Grounding
by: Guo, Yongxin, et al.
Published: (2024)
by: Guo, Yongxin, et al.
Published: (2024)
Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting
by: Tang, Yunlong, et al.
Published: (2025)
by: Tang, Yunlong, et al.
Published: (2025)
Similar Items
-
KFFocus: Highlighting Keyframes for Enhanced Video Understanding
by: Nie, Ming, et al.
Published: (2025) -
Brick-Diffusion: Generating Long Videos with Brick-to-Wall Denoising
by: Yuan, Yunlong, et al.
Published: (2025) -
Towards Unified Multimodal Interleaved Generation via Group Relative Policy Optimization
by: Nie, Ming, et al.
Published: (2026) -
Reason2Drive: Towards Interpretable and Chain-based Reasoning for Autonomous Driving
by: Nie, Ming, et al.
Published: (2023) -
VTimeCoT: Thinking by Drawing for Video Temporal Grounding and Reasoning
by: Zhang, Jinglei, et al.
Published: (2025)