Efficient Multi-modal Large Language Models via Progressive Consistency Distillation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wen, Zichen, Wang, Shaobo, Zhou, Yufa, Zhang, Junyuan, Zhang, Qintong, Gao, Yifeng, Chen, Zhaorun, Wang, Bin, Li, Weijia, He, Conghui, Zhang, Linfeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More
von: Wen, Zichen, et al.
Veröffentlicht: (2025)
von: Wen, Zichen, et al.
Veröffentlicht: (2025)
Token Pruning in Multimodal Large Language Models: Are We Solving the Right Problem?
von: Wen, Zichen, et al.
Veröffentlicht: (2025)
von: Wen, Zichen, et al.
Veröffentlicht: (2025)
Respecting Self-Uncertainty in On-Policy Self-Distillation for Efficient LLM Reasoning
von: Ke, Junlong, et al.
Veröffentlicht: (2026)
von: Ke, Junlong, et al.
Veröffentlicht: (2026)
Grounding and Enhancing Informativeness and Utility in Dataset Distillation
von: Wang, Shaobo, et al.
Veröffentlicht: (2026)
von: Wang, Shaobo, et al.
Veröffentlicht: (2026)
OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation
von: Zhang, Junyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Junyuan, et al.
Veröffentlicht: (2024)
Prune2Drive: A Plug-and-Play Framework for Accelerating Vision-Language Models in Autonomous Driving
von: Xiong, Minhao, et al.
Veröffentlicht: (2025)
von: Xiong, Minhao, et al.
Veröffentlicht: (2025)
VideoCompressa: Data-Efficient Video Understanding via Joint Temporal Compression and Spatial Reconstruction
von: Wang, Shaobo, et al.
Veröffentlicht: (2025)
von: Wang, Shaobo, et al.
Veröffentlicht: (2025)
DOCR-Inspector: Fine-Grained and Automated Evaluation of Document Parsing with VLM
von: Zhang, Qintong, et al.
Veröffentlicht: (2025)
von: Zhang, Qintong, et al.
Veröffentlicht: (2025)
Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction
von: Zhang, Qintong, et al.
Veröffentlicht: (2024)
von: Zhang, Qintong, et al.
Veröffentlicht: (2024)
Dataset Distillation with Neural Characteristic Function: A Minmax Perspective
von: Wang, Shaobo, et al.
Veröffentlicht: (2025)
von: Wang, Shaobo, et al.
Veröffentlicht: (2025)
Data Whisperer: Efficient Data Selection for Task-Specific LLM Fine-Tuning via Few-Shot In-Context Learning
von: Wang, Shaobo, et al.
Veröffentlicht: (2025)
von: Wang, Shaobo, et al.
Veröffentlicht: (2025)
Flash-Unified: A Training-Free and Task-Aware Acceleration Framework for Native Unified Models
von: Ke, Junlong, et al.
Veröffentlicht: (2026)
von: Ke, Junlong, et al.
Veröffentlicht: (2026)
LEGION: Learning to Ground and Explain for Synthetic Image Detection
von: Kang, Hengrui, et al.
Veröffentlicht: (2025)
von: Kang, Hengrui, et al.
Veröffentlicht: (2025)
PM4Bench: Benchmarking Large Vision-Language Models with Parallel Multilingual Multi-Modal Multi-task Corpus
von: Gao, Junyuan, et al.
Veröffentlicht: (2025)
von: Gao, Junyuan, et al.
Veröffentlicht: (2025)
Stability Implies Redundancy: Delta Attention Selective Halting for Efficient Long-Context Prefilling
von: Chen, Yujie, et al.
Veröffentlicht: (2026)
von: Chen, Yujie, et al.
Veröffentlicht: (2026)
TRivia: Self-supervised Fine-tuning of Vision-Language Models for Table Recognition
von: Zhang, Junyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Junyuan, et al.
Veröffentlicht: (2025)
OmniDocLayout: Towards Diverse Document Layout Generation via Coarse-to-Fine LLM Learning
von: Kang, Hengrui, et al.
Veröffentlicht: (2025)
von: Kang, Hengrui, et al.
Veröffentlicht: (2025)
IPCV: Information-Preserving Compression for MLLM Visual Encoders
von: Chen, Yuan, et al.
Veröffentlicht: (2025)
von: Chen, Yuan, et al.
Veröffentlicht: (2025)
AudioKV: KV Cache Eviction in Efficient Large Audio Language Models
von: Wang, Yuxuan, et al.
Veröffentlicht: (2026)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2026)
From Consistency to Complementarity: Aligned and Disentangled Multi-modal Learning for Time Series Understanding and Reasoning
von: Ni, Hang, et al.
Veröffentlicht: (2026)
von: Ni, Hang, et al.
Veröffentlicht: (2026)
Diffusion LLM with Native Variable Generation Lengths: Let [EOS] Lead the Way
von: Yang, Yicun, et al.
Veröffentlicht: (2025)
von: Yang, Yicun, et al.
Veröffentlicht: (2025)
Thinking Inside the Mask: In-Place Prompting in Diffusion LLMs
von: Jin, Xiangqi, et al.
Veröffentlicht: (2025)
von: Jin, Xiangqi, et al.
Veröffentlicht: (2025)
Towards Principled Dataset Distillation: A Spectral Distribution Perspective
von: Wu, Ruixi, et al.
Veröffentlicht: (2026)
von: Wu, Ruixi, et al.
Veröffentlicht: (2026)
Self Speculative Decoding for Diffusion Large Language Models
von: Gao, Yifeng, et al.
Veröffentlicht: (2025)
von: Gao, Yifeng, et al.
Veröffentlicht: (2025)
EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action Models
von: Yang, Yantai, et al.
Veröffentlicht: (2025)
von: Yang, Yantai, et al.
Veröffentlicht: (2025)
One-shot Federated Learning via Synthetic Distiller-Distillate Communication
von: Zhang, Junyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Junyuan, et al.
Veröffentlicht: (2024)
ImagebindDC: Compressing Multi-modal Data with Imagebind-based Condensation
von: Min, Yue, et al.
Veröffentlicht: (2025)
von: Min, Yue, et al.
Veröffentlicht: (2025)
Not All Samples Should Be Utilized Equally: Towards Understanding and Improving Dataset Distillation
von: Wang, Shaobo, et al.
Veröffentlicht: (2024)
von: Wang, Shaobo, et al.
Veröffentlicht: (2024)
Cross-View Consistency Regularisation for Knowledge Distillation
von: Zhang, Weijia, et al.
Veröffentlicht: (2024)
von: Zhang, Weijia, et al.
Veröffentlicht: (2024)
AgentSlimming: Towards Efficient and Cost-Aware Multi-Agent Systems
von: Chen, Yulang, et al.
Veröffentlicht: (2026)
von: Chen, Yulang, et al.
Veröffentlicht: (2026)
The Devil behind the mask: An emergent safety vulnerability of Diffusion LLMs
von: Wen, Zichen, et al.
Veröffentlicht: (2025)
von: Wen, Zichen, et al.
Veröffentlicht: (2025)
TemCoCo: Temporally Consistent Multi-modal Video Fusion with Visual-Semantic Collaboration
von: Gong, Meiqi, et al.
Veröffentlicht: (2025)
von: Gong, Meiqi, et al.
Veröffentlicht: (2025)
DocGenome: An Open Large-scale Scientific Document Benchmark for Training and Testing Multi-modal Large Language Models
von: Xia, Renqiu, et al.
Veröffentlicht: (2024)
von: Xia, Renqiu, et al.
Veröffentlicht: (2024)
Shifting AI Efficiency From Model-Centric to Data-Centric Compression
von: Liu, Xuyang, et al.
Veröffentlicht: (2025)
von: Liu, Xuyang, et al.
Veröffentlicht: (2025)
OPERA: Alleviating Hallucination in Multi-Modal Large Language Models via Over-Trust Penalty and Retrospection-Allocation
von: Huang, Qidong, et al.
Veröffentlicht: (2023)
von: Huang, Qidong, et al.
Veröffentlicht: (2023)
Reasoning Like an Economist: Post-Training on Economic Problems Induces Strategic Generalization in LLMs
von: Zhou, Yufa, et al.
Veröffentlicht: (2025)
von: Zhou, Yufa, et al.
Veröffentlicht: (2025)
Agentic Proposing: Enhancing Large Language Model Reasoning via Compositional Skill Synthesis
von: Jiao, Zhengbo, et al.
Veröffentlicht: (2026)
von: Jiao, Zhengbo, et al.
Veröffentlicht: (2026)
Winning the Pruning Gamble: A Unified Approach to Joint Sample and Token Pruning for Efficient Supervised Fine-Tuning
von: Wang, Shaobo, et al.
Veröffentlicht: (2025)
von: Wang, Shaobo, et al.
Veröffentlicht: (2025)
Learnable Cross-modal Knowledge Distillation for Multi-modal Learning with Missing Modality
von: Wang, Hu, et al.
Veröffentlicht: (2023)
von: Wang, Hu, et al.
Veröffentlicht: (2023)
dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive Caching
von: Liu, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Liu, Zhiyuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More
von: Wen, Zichen, et al.
Veröffentlicht: (2025) -
Token Pruning in Multimodal Large Language Models: Are We Solving the Right Problem?
von: Wen, Zichen, et al.
Veröffentlicht: (2025) -
Respecting Self-Uncertainty in On-Policy Self-Distillation for Efficient LLM Reasoning
von: Ke, Junlong, et al.
Veröffentlicht: (2026) -
Grounding and Enhancing Informativeness and Utility in Dataset Distillation
von: Wang, Shaobo, et al.
Veröffentlicht: (2026) -
OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation
von: Zhang, Junyuan, et al.
Veröffentlicht: (2024)