Gespeichert in:
| Hauptverfasser: | Wang, Kaibin, Lin, Mingbao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2511.17945 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Efficient Motion-Aware Video MLLM
von: Zhao, Zijia, et al.
Veröffentlicht: (2025)
von: Zhao, Zijia, et al.
Veröffentlicht: (2025)
RTV-Bench: Benchmarking MLLM Continuous Perception, Understanding and Reasoning through Real-Time Video
von: Xun, Shuhang, et al.
Veröffentlicht: (2025)
von: Xun, Shuhang, et al.
Veröffentlicht: (2025)
Query-Conditioned Evidential Keyframe Sampling for MLLM-Based Long-Form Video Understanding
von: Wang, Yiheng, et al.
Veröffentlicht: (2026)
von: Wang, Yiheng, et al.
Veröffentlicht: (2026)
Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution
von: Liu, Zuyan, et al.
Veröffentlicht: (2024)
von: Liu, Zuyan, et al.
Veröffentlicht: (2024)
MSJoE: Jointly Evolving MLLM and Sampler for Efficient Long-Form Video Understanding
von: Tan, Wenhui, et al.
Veröffentlicht: (2026)
von: Tan, Wenhui, et al.
Veröffentlicht: (2026)
Progressive Video Condensation with MLLM Agent for Long-form Video Understanding
von: Yin, Yufei, et al.
Veröffentlicht: (2026)
von: Yin, Yufei, et al.
Veröffentlicht: (2026)
LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM Agents
von: Chen, Boyu, et al.
Veröffentlicht: (2025)
von: Chen, Boyu, et al.
Veröffentlicht: (2025)
Tempo-R0: A Video-MLLM for Temporal Video Grounding through Efficient Temporal Sensing Reinforcement Learning
von: Yue, Feng, et al.
Veröffentlicht: (2025)
von: Yue, Feng, et al.
Veröffentlicht: (2025)
Exo2Ego: Exocentric Knowledge Guided MLLM for Egocentric Video Understanding
von: Zhang, Haoyu, et al.
Veröffentlicht: (2025)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2025)
Move and Act: Enhanced Object Manipulation and Background Integrity for Image Editing
von: Jiang, Pengfei, et al.
Veröffentlicht: (2024)
von: Jiang, Pengfei, et al.
Veröffentlicht: (2024)
ANTS: Adaptive Negative Textual Space Shaping for OOD Detection via Test-Time MLLM Understanding and Reasoning
von: Zhu, Wenjie, et al.
Veröffentlicht: (2025)
von: Zhu, Wenjie, et al.
Veröffentlicht: (2025)
MASRA: MLLM-Assisted Semantic-Relational Consistent Alignment for Video Temporal Grounding
von: Ran, Ran, et al.
Veröffentlicht: (2026)
von: Ran, Ran, et al.
Veröffentlicht: (2026)
In the Eye of MLLM: Benchmarking Egocentric Video Intent Understanding with Gaze-Guided Prompting
von: Peng, Taiying, et al.
Veröffentlicht: (2025)
von: Peng, Taiying, et al.
Veröffentlicht: (2025)
VideoJudge: Bootstrapping Enables Scalable Supervision of MLLM-as-a-Judge for Video Understanding
von: Waheed, Abdul, et al.
Veröffentlicht: (2025)
von: Waheed, Abdul, et al.
Veröffentlicht: (2025)
Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding
von: Hu, Pengfei, et al.
Veröffentlicht: (2025)
von: Hu, Pengfei, et al.
Veröffentlicht: (2025)
AccDiffusion: An Accurate Method for Higher-Resolution Image Generation
von: Lin, Zhihang, et al.
Veröffentlicht: (2024)
von: Lin, Zhihang, et al.
Veröffentlicht: (2024)
UniVST: A Unified Framework for Training-free Localized Video Style Transfer
von: Song, Quanjian, et al.
Veröffentlicht: (2024)
von: Song, Quanjian, et al.
Veröffentlicht: (2024)
Head-Aware KV Cache Compression for Efficient Visual Autoregressive Modeling
von: Qin, Ziran, et al.
Veröffentlicht: (2025)
von: Qin, Ziran, et al.
Veröffentlicht: (2025)
TSPO: Temporal Sampling Policy Optimization for Long-form Video Language Understanding
von: Tang, Canhui, et al.
Veröffentlicht: (2025)
von: Tang, Canhui, et al.
Veröffentlicht: (2025)
VideoCompressa: Data-Efficient Video Understanding via Joint Temporal Compression and Spatial Reconstruction
von: Wang, Shaobo, et al.
Veröffentlicht: (2025)
von: Wang, Shaobo, et al.
Veröffentlicht: (2025)
dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models
von: Xin, Yi, et al.
Veröffentlicht: (2025)
von: Xin, Yi, et al.
Veröffentlicht: (2025)
EasyInv: Toward Fast and Better DDIM Inversion
von: Zhang, Ziyue, et al.
Veröffentlicht: (2024)
von: Zhang, Ziyue, et al.
Veröffentlicht: (2024)
DiffusionTrend: A Minimalist Approach to Virtual Fashion Try-On
von: Zhan, Wengyi, et al.
Veröffentlicht: (2024)
von: Zhan, Wengyi, et al.
Veröffentlicht: (2024)
Parallel Vision Token Scheduling for Fast and Accurate Multimodal LMMs Inference
von: Zhan, Wengyi, et al.
Veröffentlicht: (2025)
von: Zhan, Wengyi, et al.
Veröffentlicht: (2025)
OCC-MLLM-Alpha:Empowering Multi-modal Large Language Model for the Understanding of Occluded Objects with Self-Supervised Test-Time Learning
von: Yang, Shuxin, et al.
Veröffentlicht: (2024)
von: Yang, Shuxin, et al.
Veröffentlicht: (2024)
Progressive Supernet Training for Efficient Visual Autoregressive Modeling
von: Chen, Xiaoyue, et al.
Veröffentlicht: (2025)
von: Chen, Xiaoyue, et al.
Veröffentlicht: (2025)
AdaptToken: Entropy-based Adaptive Token Selection for MLLM Long Video Understanding
von: Qi, Haozhe, et al.
Veröffentlicht: (2026)
von: Qi, Haozhe, et al.
Veröffentlicht: (2026)
Elysium: Exploring Object-level Perception in Videos via MLLM
von: Wang, Han, et al.
Veröffentlicht: (2024)
von: Wang, Han, et al.
Veröffentlicht: (2024)
Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid Inference
von: Lin, Zhihang, et al.
Veröffentlicht: (2024)
von: Lin, Zhihang, et al.
Veröffentlicht: (2024)
Towards Efficient and Effective Text-to-Video Retrieval with Coarse-to-Fine Visual Representation Learning
von: Tian, Kaibin, et al.
Veröffentlicht: (2024)
von: Tian, Kaibin, et al.
Veröffentlicht: (2024)
Improving MLLM Historical Record Extraction with Test-Time Image
von: Archibald, Taylor, et al.
Veröffentlicht: (2025)
von: Archibald, Taylor, et al.
Veröffentlicht: (2025)
Omni-Video 2: Scaling MLLM-Conditioned Diffusion for Unified Video Generation and Editing
von: Yang, Hao, et al.
Veröffentlicht: (2026)
von: Yang, Hao, et al.
Veröffentlicht: (2026)
Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?
von: Cheng, Junhao, et al.
Veröffentlicht: (2025)
von: Cheng, Junhao, et al.
Veröffentlicht: (2025)
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams
von: Zhang, Haoji, et al.
Veröffentlicht: (2025)
von: Zhang, Haoji, et al.
Veröffentlicht: (2025)
Temporal Cluster Assignment for Efficient Real-Time Video Segmentation
von: Yung, Ka-Wai, et al.
Veröffentlicht: (2025)
von: Yung, Ka-Wai, et al.
Veröffentlicht: (2025)
AccDiffusion v2: Towards More Accurate Higher-Resolution Diffusion Extrapolation
von: Lin, Zhihang, et al.
Veröffentlicht: (2024)
von: Lin, Zhihang, et al.
Veröffentlicht: (2024)
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding
von: Ahn, Geo, et al.
Veröffentlicht: (2026)
von: Ahn, Geo, et al.
Veröffentlicht: (2026)
Temporal2Seq: A Unified Framework for Temporal Video Understanding Tasks
von: Yang, Min, et al.
Veröffentlicht: (2024)
von: Yang, Min, et al.
Veröffentlicht: (2024)
Efficient Video Sampling: Pruning Temporally Redundant Tokens for Faster VLM Inference
von: Bagrov, Natan, et al.
Veröffentlicht: (2025)
von: Bagrov, Natan, et al.
Veröffentlicht: (2025)
ERTACache: Error Rectification and Timesteps Adjustment for Efficient Diffusion
von: Peng, Xurui, et al.
Veröffentlicht: (2025)
von: Peng, Xurui, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Efficient Motion-Aware Video MLLM
von: Zhao, Zijia, et al.
Veröffentlicht: (2025) -
RTV-Bench: Benchmarking MLLM Continuous Perception, Understanding and Reasoning through Real-Time Video
von: Xun, Shuhang, et al.
Veröffentlicht: (2025) -
Query-Conditioned Evidential Keyframe Sampling for MLLM-Based Long-Form Video Understanding
von: Wang, Yiheng, et al.
Veröffentlicht: (2026) -
Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution
von: Liu, Zuyan, et al.
Veröffentlicht: (2024) -
MSJoE: Jointly Evolving MLLM and Sampler for Efficient Long-Form Video Understanding
von: Tan, Wenhui, et al.
Veröffentlicht: (2026)