Enregistré dans:
| Auteurs principaux: | Yang, Min, Zhang, Zichen, Wang, Limin |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2409.18478 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning
par: Wu, Tao, et autres
Publié: (2025)
par: Wu, Tao, et autres
Publié: (2025)
VideoCompressa: Data-Efficient Video Understanding via Joint Temporal Compression and Spatial Reconstruction
par: Wang, Shaobo, et autres
Publié: (2025)
par: Wang, Shaobo, et autres
Publié: (2025)
STOP: Integrated Spatial-Temporal Dynamic Prompting for Video Understanding
par: Liu, Zichen, et autres
Publié: (2025)
par: Liu, Zichen, et autres
Publié: (2025)
Seq2Time: Sequential Knowledge Transfer for Video LLM Temporal Grounding
par: Deng, Andong, et autres
Publié: (2024)
par: Deng, Andong, et autres
Publié: (2024)
VideoCoF: Unified Video Editing with Temporal Reasoner
par: Yang, Xiangpeng, et autres
Publié: (2025)
par: Yang, Xiangpeng, et autres
Publié: (2025)
State Space Prompting via Gathering and Spreading Spatio-Temporal Information for Video Understanding
par: Zhou, Jiahuan, et autres
Publié: (2025)
par: Zhou, Jiahuan, et autres
Publié: (2025)
Video-QTR: Query-Driven Temporal Reasoning Framework for Lightweight Video Understanding
par: Zhao, Xinkui, et autres
Publié: (2025)
par: Zhao, Xinkui, et autres
Publié: (2025)
TUMTraffic-VideoQA: A Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes
par: Zhou, Xingcheng, et autres
Publié: (2025)
par: Zhou, Xingcheng, et autres
Publié: (2025)
Flash-Unified: A Training-Free and Task-Aware Acceleration Framework for Native Unified Models
par: Ke, Junlong, et autres
Publié: (2026)
par: Ke, Junlong, et autres
Publié: (2026)
SmartSight: Mitigating Hallucination in Video-LLMs Without Compromising Video Understanding via Temporal Attention Collapse
par: Sun, Yiming, et autres
Publié: (2025)
par: Sun, Yiming, et autres
Publié: (2025)
VTok: A Unified Video Tokenizer with Decoupled Spatial-Temporal Latents
par: Wang, Feng, et autres
Publié: (2026)
par: Wang, Feng, et autres
Publié: (2026)
TDViT: Temporal Dilated Video Transformer for Dense Video Tasks
par: Sun, Guanxiong, et autres
Publié: (2024)
par: Sun, Guanxiong, et autres
Publié: (2024)
V-CORE: Temporally Consistent Video Understanding for Video-LLM
par: Kang, Zhengjian, et autres
Publié: (2026)
par: Kang, Zhengjian, et autres
Publié: (2026)
Video Understanding: Through A Temporal Lens
par: Nguyen, Thong Thanh
Publié: (2026)
par: Nguyen, Thong Thanh
Publié: (2026)
FiLA-Video: Spatio-Temporal Compression for Fine-Grained Long Video Understanding
par: Guo, Yanan, et autres
Publié: (2025)
par: Guo, Yanan, et autres
Publié: (2025)
Zoom-Zero: Reinforced Coarse-to-Fine Video Understanding via Temporal Zoom-in
par: Shen, Xiaoqian, et autres
Publié: (2025)
par: Shen, Xiaoqian, et autres
Publié: (2025)
Unified Static and Dynamic Network: Efficient Temporal Filtering for Video Grounding
par: Hu, Jingjing, et autres
Publié: (2024)
par: Hu, Jingjing, et autres
Publié: (2024)
Incentivizing Temporal-Awareness in Egocentric Video Understanding Models
par: Xu, Zhiyang, et autres
Publié: (2026)
par: Xu, Zhiyang, et autres
Publié: (2026)
CurConMix+: A Unified Spatio-Temporal Framework for Hierarchical Surgical Workflow Understanding
par: Jeon, Yongjun, et autres
Publié: (2026)
par: Jeon, Yongjun, et autres
Publié: (2026)
Dual DETRs for Multi-Label Temporal Action Detection
par: Zhu, Yuhan, et autres
Publié: (2024)
par: Zhu, Yuhan, et autres
Publié: (2024)
Test-Time Temporal Sampling for Efficient MLLM Video Understanding
par: Wang, Kaibin, et autres
Publié: (2025)
par: Wang, Kaibin, et autres
Publié: (2025)
PiTe: Pixel-Temporal Alignment for Large Video-Language Model
par: Liu, Yang, et autres
Publié: (2024)
par: Liu, Yang, et autres
Publié: (2024)
VideoExpert: Augmented LLM for Temporal-Sensitive Video Understanding
par: Zhao, Henghao, et autres
Publié: (2025)
par: Zhao, Henghao, et autres
Publié: (2025)
TV-RAG: A Temporal-aware and Semantic Entropy-Weighted Framework for Long Video Retrieval and Understanding
par: Cao, Zongsheng, et autres
Publié: (2025)
par: Cao, Zongsheng, et autres
Publié: (2025)
Enhancing Temporal Understanding in Video-LLMs through Stacked Temporal Attention in Vision Encoders
par: Rasekh, Ali, et autres
Publié: (2025)
par: Rasekh, Ali, et autres
Publié: (2025)
SlowFocus: Enhancing Fine-grained Temporal Understanding in Video LLM
par: Nie, Ming, et autres
Publié: (2026)
par: Nie, Ming, et autres
Publié: (2026)
VideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding
par: Shi, Jiapeng, et autres
Publié: (2026)
par: Shi, Jiapeng, et autres
Publié: (2026)
TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
par: Zhang, Jun, et autres
Publié: (2025)
par: Zhang, Jun, et autres
Publié: (2025)
SVBench: A Benchmark with Temporal Multi-Turn Dialogues for Streaming Video Understanding
par: Yang, Zhenyu, et autres
Publié: (2025)
par: Yang, Zhenyu, et autres
Publié: (2025)
Temporal-Visual Semantic Alignment: A Unified Architecture for Transferring Spatial Priors from Vision Models to Zero-Shot Temporal Tasks
par: Ma, Xiangkai, et autres
Publié: (2025)
par: Ma, Xiangkai, et autres
Publié: (2025)
VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
par: Wang, Shihao, et autres
Publié: (2025)
par: Wang, Shihao, et autres
Publié: (2025)
VideoTG-R1: Boosting Video Temporal Grounding via Curriculum Reinforcement Learning on Reflected Boundary Annotations
par: Dong, Lu, et autres
Publié: (2025)
par: Dong, Lu, et autres
Publié: (2025)
Open-Vocabulary Spatio-Temporal Action Detection
par: Wu, Tao, et autres
Publié: (2024)
par: Wu, Tao, et autres
Publié: (2024)
Mobius: A High Efficient Spatial-Temporal Parallel Training Paradigm for Text-to-Video Generation Task
par: Yang, Yiran, et autres
Publié: (2024)
par: Yang, Yiran, et autres
Publié: (2024)
Harnessing Synthetic Preference Data for Enhancing Temporal Understanding of Video-LLMs
par: Vani, Sameep, et autres
Publié: (2025)
par: Vani, Sameep, et autres
Publié: (2025)
T*: Re-thinking Temporal Search for Long-Form Video Understanding
par: Ye, Jinhui, et autres
Publié: (2025)
par: Ye, Jinhui, et autres
Publié: (2025)
PixelRefer: A Unified Framework for Spatio-Temporal Object Referring with Arbitrary Granularity
par: Yuan, Yuqian, et autres
Publié: (2025)
par: Yuan, Yuqian, et autres
Publié: (2025)
Temporal Residual Guided Diffusion Framework for Event-Driven Video Reconstruction
par: Zhu, Lin, et autres
Publié: (2024)
par: Zhu, Lin, et autres
Publié: (2024)
RadarSeq: A Temporal Vision Framework for User Churn Prediction via Radar Chart Sequences
par: Najafi, Sina, et autres
Publié: (2025)
par: Najafi, Sina, et autres
Publié: (2025)
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
par: Cheng, Zesen, et autres
Publié: (2024)
par: Cheng, Zesen, et autres
Publié: (2024)
Documents similaires
-
TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning
par: Wu, Tao, et autres
Publié: (2025) -
VideoCompressa: Data-Efficient Video Understanding via Joint Temporal Compression and Spatial Reconstruction
par: Wang, Shaobo, et autres
Publié: (2025) -
STOP: Integrated Spatial-Temporal Dynamic Prompting for Video Understanding
par: Liu, Zichen, et autres
Publié: (2025) -
Seq2Time: Sequential Knowledge Transfer for Video LLM Temporal Grounding
par: Deng, Andong, et autres
Publié: (2024) -
VideoCoF: Unified Video Editing with Temporal Reasoner
par: Yang, Xiangpeng, et autres
Publié: (2025)