p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Jun, Meng, Desen, Zhang, Zhengming, Huang, Zhenpeng, Wu, Tao, Wang, Limin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking
di: Meng, Desen, et al.
Pubblicazione: (2025)
di: Meng, Desen, et al.
Pubblicazione: (2025)
$γ-$MoD: Exploring Mixture-of-Depth Adaptation for Multimodal Large Language Models
di: Luo, Yaxin, et al.
Pubblicazione: (2024)
di: Luo, Yaxin, et al.
Pubblicazione: (2024)
VideoLLM-MoD: Efficient Video-Language Streaming with Mixture-of-Depths Vision Computation
di: Wu, Shiwei, et al.
Pubblicazione: (2024)
di: Wu, Shiwei, et al.
Pubblicazione: (2024)
MoD-DPO: Towards Mitigating Cross-modal Hallucinations in Omni LLMs using Modality Decoupled Preference Optimization
di: Chaubey, Ashutosh, et al.
Pubblicazione: (2026)
di: Chaubey, Ashutosh, et al.
Pubblicazione: (2026)
LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation
di: Shu, Fangxun, et al.
Pubblicazione: (2024)
di: Shu, Fangxun, et al.
Pubblicazione: (2024)
MoDES: Accelerating Mixture-of-Experts Multimodal Large Language Models via Dynamic Expert Skipping
di: Huang, Yushi, et al.
Pubblicazione: (2025)
di: Huang, Yushi, et al.
Pubblicazione: (2025)
Octavius: Mitigating Task Interference in MLLMs via LoRA-MoE
di: Chen, Zeren, et al.
Pubblicazione: (2023)
di: Chen, Zeren, et al.
Pubblicazione: (2023)
MoD-SLAM: Monocular Dense Mapping for Unbounded 3D Scene Reconstruction
di: Zhou, Heng, et al.
Pubblicazione: (2024)
di: Zhou, Heng, et al.
Pubblicazione: (2024)
MedRCube: A Multidimensional Framework for Fine-Grained and In-Depth Evaluation of MLLMs in Medical Imaging
di: Bao, Zhijie, et al.
Pubblicazione: (2026)
di: Bao, Zhijie, et al.
Pubblicazione: (2026)
Mitigating Object Hallucinations in MLLMs via Multi-Frequency Perturbations
di: Li, Shuo, et al.
Pubblicazione: (2025)
di: Li, Shuo, et al.
Pubblicazione: (2025)
CameraBench: Benchmarking Visual Reasoning in MLLMs via Photography
di: Fang, I-Sheng, et al.
Pubblicazione: (2025)
di: Fang, I-Sheng, et al.
Pubblicazione: (2025)
Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs
di: Zhang, Huanyu, et al.
Pubblicazione: (2025)
di: Zhang, Huanyu, et al.
Pubblicazione: (2025)
Seeing is Believing: Rich-Context Hallucination Detection for MLLMs via Backward Visual Grounding
di: Guo, Pinxue, et al.
Pubblicazione: (2025)
di: Guo, Pinxue, et al.
Pubblicazione: (2025)
Interpretable and Reliable Detection of AI-Generated Images via Grounded Reasoning in MLLMs
di: Ji, Yikun, et al.
Pubblicazione: (2025)
di: Ji, Yikun, et al.
Pubblicazione: (2025)
Unhackable Temporal Rewarding for Scalable Video MLLMs
di: Yu, En, et al.
Pubblicazione: (2025)
di: Yu, En, et al.
Pubblicazione: (2025)
MoPD: Mixture-of-Prompts Distillation for Vision-Language Models
di: Chen, Yang, et al.
Pubblicazione: (2024)
di: Chen, Yang, et al.
Pubblicazione: (2024)
MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs
di: Shi, Baorong, et al.
Pubblicazione: (2026)
di: Shi, Baorong, et al.
Pubblicazione: (2026)
MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI
di: Yao, Huanjin, et al.
Pubblicazione: (2025)
di: Yao, Huanjin, et al.
Pubblicazione: (2025)
LongVPO: From Anchored Cues to Self-Reasoning for Long-Form Video Preference Optimization
di: Huang, Zhenpeng, et al.
Pubblicazione: (2026)
di: Huang, Zhenpeng, et al.
Pubblicazione: (2026)
Linking Perception, Confidence and Accuracy in MLLMs
di: Du, Yuetian, et al.
Pubblicazione: (2026)
di: Du, Yuetian, et al.
Pubblicazione: (2026)
RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback
di: Yu, Tianyu, et al.
Pubblicazione: (2023)
di: Yu, Tianyu, et al.
Pubblicazione: (2023)
Med-MoE: Mixture of Domain-Specific Experts for Lightweight Medical Vision-Language Models
di: Jiang, Songtao, et al.
Pubblicazione: (2024)
di: Jiang, Songtao, et al.
Pubblicazione: (2024)
MoDification: Mixture of Depths Made Easy
di: Zhang, Chen, et al.
Pubblicazione: (2024)
di: Zhang, Chen, et al.
Pubblicazione: (2024)
The Side Effects of Being Smart: Safety Risks in MLLMs' Multi-Image Reasoning
di: Chen, Renmiao, et al.
Pubblicazione: (2026)
di: Chen, Renmiao, et al.
Pubblicazione: (2026)
3D-MoE: A Mixture-of-Experts Multi-modal LLM for 3D Vision and Pose Diffusion via Rectified Flow
di: Ma, Yueen, et al.
Pubblicazione: (2025)
di: Ma, Yueen, et al.
Pubblicazione: (2025)
Can MLLMs Read the Room? A Multimodal Benchmark for Assessing Deception in Multi-Party Social Interactions
di: Kang, Caixin, et al.
Pubblicazione: (2025)
di: Kang, Caixin, et al.
Pubblicazione: (2025)
Breaking Down and Building Up: Mixture of Skill-Based Vision-and-Language Navigation Agents
di: Ma, Tianyi, et al.
Pubblicazione: (2025)
di: Ma, Tianyi, et al.
Pubblicazione: (2025)
AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation
di: Wang, Junyang, et al.
Pubblicazione: (2023)
di: Wang, Junyang, et al.
Pubblicazione: (2023)
Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs
di: Huang, Jen-Tse, et al.
Pubblicazione: (2025)
di: Huang, Jen-Tse, et al.
Pubblicazione: (2025)
MaskMoE: Boosting Token-Level Learning via Routing Mask in Mixture-of-Experts
di: Su, Zhenpeng, et al.
Pubblicazione: (2024)
di: Su, Zhenpeng, et al.
Pubblicazione: (2024)
D$^{3}$ToM: Decider-Guided Dynamic Token Merging for Accelerating Diffusion MLLMs
di: Chang, Shuochen, et al.
Pubblicazione: (2025)
di: Chang, Shuochen, et al.
Pubblicazione: (2025)
MoEController: Instruction-based Arbitrary Image Manipulation with Mixture-of-Expert Controllers
di: Li, Sijia, et al.
Pubblicazione: (2023)
di: Li, Sijia, et al.
Pubblicazione: (2023)
Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
di: Li, Yunxin, et al.
Pubblicazione: (2024)
di: Li, Yunxin, et al.
Pubblicazione: (2024)
The Instinctive Bias: Spurious Images lead to Illusion in MLLMs
di: Han, Tianyang, et al.
Pubblicazione: (2024)
di: Han, Tianyang, et al.
Pubblicazione: (2024)
HiDrop: Hierarchical Vision Token Reduction in MLLMs via Late Injection, Concave Pyramid Pruning, and Early Exit
di: Wu, Hao, et al.
Pubblicazione: (2026)
di: Wu, Hao, et al.
Pubblicazione: (2026)
SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMs
di: Wang, Siting, et al.
Pubblicazione: (2025)
di: Wang, Siting, et al.
Pubblicazione: (2025)
NeMo: Needle in a Montage for Video-Language Understanding
di: Hu, Zi-Yuan, et al.
Pubblicazione: (2025)
di: Hu, Zi-Yuan, et al.
Pubblicazione: (2025)
Exploring the Design Space of Visual Context Representation in Video MLLMs
di: Du, Yifan, et al.
Pubblicazione: (2024)
di: Du, Yifan, et al.
Pubblicazione: (2024)
MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning
di: Liang, Yiqing, et al.
Pubblicazione: (2025)
di: Liang, Yiqing, et al.
Pubblicazione: (2025)
Can MLLMs Understand the Deep Implication Behind Chinese Images?
di: Zhang, Chenhao, et al.
Pubblicazione: (2024)
di: Zhang, Chenhao, et al.
Pubblicazione: (2024)
Documenti analoghi
-
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking
di: Meng, Desen, et al.
Pubblicazione: (2025) -
$γ-$MoD: Exploring Mixture-of-Depth Adaptation for Multimodal Large Language Models
di: Luo, Yaxin, et al.
Pubblicazione: (2024) -
VideoLLM-MoD: Efficient Video-Language Streaming with Mixture-of-Depths Vision Computation
di: Wu, Shiwei, et al.
Pubblicazione: (2024) -
MoD-DPO: Towards Mitigating Cross-modal Hallucinations in Omni LLMs using Modality Decoupled Preference Optimization
di: Chaubey, Ashutosh, et al.
Pubblicazione: (2026) -
LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation
di: Shu, Fangxun, et al.
Pubblicazione: (2024)