MSJoE: Jointly Evolving MLLM and Sampler for Efficient Long-Form Video Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tan, Wenhui, Yu, Xiaoyi, Li, Jiaze, Chen, Yijing, Ju, Jianzhong, Luo, Zhenbo, Song, Ruihua, Luan, Jian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Think-Clip-Sample: Slow-Fast Frame Selection for Video Understanding
von: Tan, Wenhui, et al.
Veröffentlicht: (2026)
von: Tan, Wenhui, et al.
Veröffentlicht: (2026)
Think Silently, Think Fast: Dynamic Latent Compression of LLM Reasoning Chains
von: Tan, Wenhui, et al.
Veröffentlicht: (2025)
von: Tan, Wenhui, et al.
Veröffentlicht: (2025)
REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding
von: Li, Jiaze, et al.
Veröffentlicht: (2025)
von: Li, Jiaze, et al.
Veröffentlicht: (2025)
TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding
von: Xu, Boshen, et al.
Veröffentlicht: (2025)
von: Xu, Boshen, et al.
Veröffentlicht: (2025)
Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation
von: Li, Jiaze, et al.
Veröffentlicht: (2026)
von: Li, Jiaze, et al.
Veröffentlicht: (2026)
Federated Joint Learning for Domain and Class Generalization
von: Xu, Haoran, et al.
Veröffentlicht: (2026)
von: Xu, Haoran, et al.
Veröffentlicht: (2026)
Restoring Exploration after Post-Training: Latent Exploration Decoding for Large Reasoning Models
von: Tan, Wenhui, et al.
Veröffentlicht: (2026)
von: Tan, Wenhui, et al.
Veröffentlicht: (2026)
Visual Para-Thinker: Divide-and-Conquer Reasoning for Visual Comprehension
von: Xu, Haoran, et al.
Veröffentlicht: (2026)
von: Xu, Haoran, et al.
Veröffentlicht: (2026)
Xiaomi MiMo-VL-Miloco Technical Report
von: Li, Jiaze, et al.
Veröffentlicht: (2025)
von: Li, Jiaze, et al.
Veröffentlicht: (2025)
Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
von: Guan, Yiran, et al.
Veröffentlicht: (2026)
von: Guan, Yiran, et al.
Veröffentlicht: (2026)
Listening with Time: Precise Temporal Awareness for Long-Form Audio Understanding
von: Shao, Mingchen, et al.
Veröffentlicht: (2026)
von: Shao, Mingchen, et al.
Veröffentlicht: (2026)
JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation
von: Chao, Jianghan, et al.
Veröffentlicht: (2025)
von: Chao, Jianghan, et al.
Veröffentlicht: (2025)
Shuffle-R1: Efficient RL framework for Multimodal Large Language Models via Data-centric Dynamic Shuffle
von: Zhu, Linghao, et al.
Veröffentlicht: (2025)
von: Zhu, Linghao, et al.
Veröffentlicht: (2025)
Generative Frame Sampler for Long Video Understanding
von: Yao, Linli, et al.
Veröffentlicht: (2025)
von: Yao, Linli, et al.
Veröffentlicht: (2025)
Query-Conditioned Evidential Keyframe Sampling for MLLM-Based Long-Form Video Understanding
von: Wang, Yiheng, et al.
Veröffentlicht: (2026)
von: Wang, Yiheng, et al.
Veröffentlicht: (2026)
Federated Balanced Learning
von: Li, Jiaze, et al.
Veröffentlicht: (2026)
von: Li, Jiaze, et al.
Veröffentlicht: (2026)
StreamPro: From Reactive Perception to Proactive Decision-Making in Streaming Video
von: Li, Ao, et al.
Veröffentlicht: (2026)
von: Li, Ao, et al.
Veröffentlicht: (2026)
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs
von: Zhang, Shaojie, et al.
Veröffentlicht: (2025)
von: Zhang, Shaojie, et al.
Veröffentlicht: (2025)
Progressive Video Condensation with MLLM Agent for Long-form Video Understanding
von: Yin, Yufei, et al.
Veröffentlicht: (2026)
von: Yin, Yufei, et al.
Veröffentlicht: (2026)
ThinkOmni: Lifting Textual Reasoning to Omni-modal Scenarios via Guidance Decoding
von: Guan, Yiran, et al.
Veröffentlicht: (2026)
von: Guan, Yiran, et al.
Veröffentlicht: (2026)
Direction-Aware Diagonal Autoregressive Image Generation
von: Xu, Yijia, et al.
Veröffentlicht: (2025)
von: Xu, Yijia, et al.
Veröffentlicht: (2025)
LLaVA-SG: Leveraging Scene Graphs as Visual Semantic Expression in Vision-Language Models
von: Wang, Jingyi, et al.
Veröffentlicht: (2024)
von: Wang, Jingyi, et al.
Veröffentlicht: (2024)
GeoFocus: Blending Efficient Global-to-Local Perception for Multimodal Geometry Problem-Solving
von: Deng, Linger, et al.
Veröffentlicht: (2026)
von: Deng, Linger, et al.
Veröffentlicht: (2026)
VSSFlow: Unifying Video-conditioned Sound and Speech Generation via Joint Learning
von: Cheng, Xin, et al.
Veröffentlicht: (2025)
von: Cheng, Xin, et al.
Veröffentlicht: (2025)
Test-Time Temporal Sampling for Efficient MLLM Video Understanding
von: Wang, Kaibin, et al.
Veröffentlicht: (2025)
von: Wang, Kaibin, et al.
Veröffentlicht: (2025)
Cook and Clean Together: Teaching Embodied Agents for Parallel Task Execution
von: Liang, Dingkang, et al.
Veröffentlicht: (2025)
von: Liang, Dingkang, et al.
Veröffentlicht: (2025)
LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM Agents
von: Chen, Boyu, et al.
Veröffentlicht: (2025)
von: Chen, Boyu, et al.
Veröffentlicht: (2025)
BFS-PO: Best-First Search for Large Reasoning Models
von: Parascandolo, Fiorenzo, et al.
Veröffentlicht: (2026)
von: Parascandolo, Fiorenzo, et al.
Veröffentlicht: (2026)
RoLD: Robot Latent Diffusion for Multi-task Policy Modeling
von: Tan, Wenhui, et al.
Veröffentlicht: (2024)
von: Tan, Wenhui, et al.
Veröffentlicht: (2024)
CacheFlow: Compressive Streaming Memory for Efficient Long-Form Video Understanding
von: Patel, Shrenik, et al.
Veröffentlicht: (2025)
von: Patel, Shrenik, et al.
Veröffentlicht: (2025)
AdaptToken: Entropy-based Adaptive Token Selection for MLLM Long Video Understanding
von: Qi, Haozhe, et al.
Veröffentlicht: (2026)
von: Qi, Haozhe, et al.
Veröffentlicht: (2026)
Thinking in cocktail party: Chain-of-Thought and reinforcement learning for target speaker automatic speech recognition
von: Zhang, Yiru, et al.
Veröffentlicht: (2025)
von: Zhang, Yiru, et al.
Veröffentlicht: (2025)
LoVA: Long-form Video-to-Audio Generation
von: Cheng, Xin, et al.
Veröffentlicht: (2024)
von: Cheng, Xin, et al.
Veröffentlicht: (2024)
Efficient Motion-Aware Video MLLM
von: Zhao, Zijia, et al.
Veröffentlicht: (2025)
von: Zhao, Zijia, et al.
Veröffentlicht: (2025)
Syn-GRPO: Self-Evolving Data Synthesis for MLLM Perception Reasoning
von: Huang, Qihan, et al.
Veröffentlicht: (2025)
von: Huang, Qihan, et al.
Veröffentlicht: (2025)
Temporal Preference Optimization for Long-Form Video Understanding
von: Li, Rui, et al.
Veröffentlicht: (2025)
von: Li, Rui, et al.
Veröffentlicht: (2025)
Text-Conditioned Resampler For Long Form Video Understanding
von: Korbar, Bruno, et al.
Veröffentlicht: (2023)
von: Korbar, Bruno, et al.
Veröffentlicht: (2023)
LongViTU: Instruction Tuning for Long-Form Video Understanding
von: Wu, Rujie, et al.
Veröffentlicht: (2025)
von: Wu, Rujie, et al.
Veröffentlicht: (2025)
Pair-In, Pair-Out: Latent Multi-Token Prediction for Efficient LLMs
von: Tan, Wenhui, et al.
Veröffentlicht: (2026)
von: Tan, Wenhui, et al.
Veröffentlicht: (2026)
Mosaic: Cross-Modal Clustering for Efficient Video Understanding
von: Wang, Tuowei, et al.
Veröffentlicht: (2026)
von: Wang, Tuowei, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Think-Clip-Sample: Slow-Fast Frame Selection for Video Understanding
von: Tan, Wenhui, et al.
Veröffentlicht: (2026) -
Think Silently, Think Fast: Dynamic Latent Compression of LLM Reasoning Chains
von: Tan, Wenhui, et al.
Veröffentlicht: (2025) -
REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding
von: Li, Jiaze, et al.
Veröffentlicht: (2025) -
TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding
von: Xu, Boshen, et al.
Veröffentlicht: (2025) -
Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation
von: Li, Jiaze, et al.
Veröffentlicht: (2026)