VideoLLM-MoD: Efficient Video-Language Streaming with Mixture-of-Depths Vision Computation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wu, Shiwei, Chen, Joya, Lin, Kevin Qinghong, Wang, Qimeng, Gao, Yan, Xu, Qianli, Xu, Tong, Hu, Yao, Chen, Enhong, Shou, Mike Zheng |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
VideoLLM-online: Online Video Large Language Model for Streaming Video
par: Chen, Joya, et autres
Publié: (2024)
par: Chen, Joya, et autres
Publié: (2024)
Bootstrapping SparseFormers from Vision Foundation Models
par: Gao, Ziteng, et autres
Publié: (2023)
par: Gao, Ziteng, et autres
Publié: (2023)
Code2Video: A Code-centric Paradigm for Educational Video Generation
par: Chen, Yanzhe, et autres
Publié: (2025)
par: Chen, Yanzhe, et autres
Publié: (2025)
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale
par: Chen, Joya, et autres
Publié: (2025)
par: Chen, Joya, et autres
Publié: (2025)
VLog: Video-Language Models by Generative Retrieval of Narration Vocabulary
par: Lin, Kevin Qinghong, et autres
Publié: (2025)
par: Lin, Kevin Qinghong, et autres
Publié: (2025)
Paper2Video: Automatic Video Generation from Scientific Papers
par: Zhu, Zeyu, et autres
Publié: (2025)
par: Zhu, Zeyu, et autres
Publié: (2025)
Learning Video Context as Interleaved Multimodal Sequences
par: Lin, Kevin Qinghong, et autres
Publié: (2024)
par: Lin, Kevin Qinghong, et autres
Publié: (2024)
Benchmarking Large Language Models for Conversational Question Answering in Multi-instructional Documents
par: Wu, Shiwei, et autres
Publié: (2024)
par: Wu, Shiwei, et autres
Publié: (2024)
VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video Reasoning
par: Liu, Ye, et autres
Publié: (2025)
par: Liu, Ye, et autres
Publié: (2025)
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths
par: Mao, Weijia, et autres
Publié: (2025)
par: Mao, Weijia, et autres
Publié: (2025)
p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay
par: Zhang, Jun, et autres
Publié: (2024)
par: Zhang, Jun, et autres
Publié: (2024)
$γ-$MoD: Exploring Mixture-of-Depth Adaptation for Multimodal Large Language Models
par: Luo, Yaxin, et autres
Publié: (2024)
par: Luo, Yaxin, et autres
Publié: (2024)
Sharp Eyes and Memory for VideoLLMs: Information-Aware Visual Token Pruning for Efficient and Reliable VideoLLM Reasoning
par: Qin, Jialong, et autres
Publié: (2025)
par: Qin, Jialong, et autres
Publié: (2025)
Memory-efficient Streaming VideoLLMs for Real-time Procedural Video Understanding
par: Chatterjee, Dibyadip, et autres
Publié: (2025)
par: Chatterjee, Dibyadip, et autres
Publié: (2025)
Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
par: Guan, Yiran, et autres
Publié: (2026)
par: Guan, Yiran, et autres
Publié: (2026)
MoDification: Mixture of Depths Made Easy
par: Zhang, Chen, et autres
Publié: (2024)
par: Zhang, Chen, et autres
Publié: (2024)
From a Social Cognitive Perspective: Context-aware Visual Social Relationship Recognition
par: Wu, Shiwei, et autres
Publié: (2024)
par: Wu, Shiwei, et autres
Publié: (2024)
VideoLLM Benchmarks and Evaluation: A Survey
par: Kumar, Yogesh
Publié: (2025)
par: Kumar, Yogesh
Publié: (2025)
VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning
par: Li, Chenglin, et autres
Publié: (2026)
par: Li, Chenglin, et autres
Publié: (2026)
Language-Guided Temporal Token Pruning for Efficient VideoLLM Processing
par: Kumar, Yogesh
Publié: (2025)
par: Kumar, Yogesh
Publié: (2025)
When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding
par: Fang, Pengcheng, et autres
Publié: (2025)
par: Fang, Pengcheng, et autres
Publié: (2025)
Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models
par: Wang, Haibo, et autres
Publié: (2024)
par: Wang, Haibo, et autres
Publié: (2024)
VideoGUI: A Benchmark for GUI Automation from Instructional Videos
par: Lin, Kevin Qinghong, et autres
Publié: (2024)
par: Lin, Kevin Qinghong, et autres
Publié: (2024)
Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration
par: Song, Yiren, et autres
Publié: (2026)
par: Song, Yiren, et autres
Publié: (2026)
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting
par: Ilaslan, Muhammet Furkan, et autres
Publié: (2024)
par: Ilaslan, Muhammet Furkan, et autres
Publié: (2024)
WeaveTime: Stream from Earlier Frames into Emergent Memory in VideoLLMs
par: Zhang, Yulin, et autres
Publié: (2026)
par: Zhang, Yulin, et autres
Publié: (2026)
IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes
par: Liang, Yujia, et autres
Publié: (2025)
par: Liang, Yujia, et autres
Publié: (2025)
One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos
par: Bai, Zechen, et autres
Publié: (2024)
par: Bai, Zechen, et autres
Publié: (2024)
Think or Not? Selective Reasoning via Reinforcement Learning for Vision-Language Models
par: Wang, Jiaqi, et autres
Publié: (2025)
par: Wang, Jiaqi, et autres
Publié: (2025)
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation
par: Yin, Shukang, et autres
Publié: (2024)
par: Yin, Shukang, et autres
Publié: (2024)
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM
par: Wang, Han, et autres
Publié: (2024)
par: Wang, Han, et autres
Publié: (2024)
NoteLLM: A Retrievable Large Language Model for Note Recommendation
par: Zhang, Chao, et autres
Publié: (2024)
par: Zhang, Chao, et autres
Publié: (2024)
EchoPrune: Interpreting Redundancy as Temporal Echoes for Efficient VideoLLMs
par: Li, Jiameng, et autres
Publié: (2026)
par: Li, Jiameng, et autres
Publié: (2026)
PANDA: Towards Generalist Video Anomaly Detection via Agentic AI Engineer
par: Yang, Zhiwei, et autres
Publié: (2025)
par: Yang, Zhiwei, et autres
Publié: (2025)
Reinforcement Learning for Large Model: A Survey
par: Wu, Weijia, et autres
Publié: (2025)
par: Wu, Weijia, et autres
Publié: (2025)
VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interaction Format
par: Wang, Yueqian, et autres
Publié: (2024)
par: Wang, Yueqian, et autres
Publié: (2024)
Proact-VL: A Proactive VideoLLM for Real-Time AI Companions
par: Yan, Weicai, et autres
Publié: (2026)
par: Yan, Weicai, et autres
Publié: (2026)
ShowUI-$π$: Flow-based Generative Models as GUI Dexterous Hands
par: Hu, Siyuan, et autres
Publié: (2025)
par: Hu, Siyuan, et autres
Publié: (2025)
Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs
par: Chung, Hyungjin, et autres
Publié: (2025)
par: Chung, Hyungjin, et autres
Publié: (2025)
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
par: Lin, Kevin Qinghong, et autres
Publié: (2024)
par: Lin, Kevin Qinghong, et autres
Publié: (2024)
Documents similaires
-
VideoLLM-online: Online Video Large Language Model for Streaming Video
par: Chen, Joya, et autres
Publié: (2024) -
Bootstrapping SparseFormers from Vision Foundation Models
par: Gao, Ziteng, et autres
Publié: (2023) -
Code2Video: A Code-centric Paradigm for Educational Video Generation
par: Chen, Yanzhe, et autres
Publié: (2025) -
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale
par: Chen, Joya, et autres
Publié: (2025) -
VLog: Video-Language Models by Generative Retrieval of Narration Vocabulary
par: Lin, Kevin Qinghong, et autres
Publié: (2025)