Mixpert: Mitigating Multimodal Learning Conflicts with Efficient Mixture-of-Vision-Experts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | He, Xin, Han, Xumeng, Wei, Longhui, Xie, Lingxi, Tian, Qi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Incorporating Visual Experts to Resolve the Information Loss in Multimodal Large Language Models
von: He, Xin, et al.
Veröffentlicht: (2024)
von: He, Xin, et al.
Veröffentlicht: (2024)
ViMoE: An Empirical Study of Designing Vision Mixture-of-Experts
von: Han, Xumeng, et al.
Veröffentlicht: (2024)
von: Han, Xumeng, et al.
Veröffentlicht: (2024)
EMMA: Efficient Multimodal Understanding, Generation, and Editing with a Unified Architecture
von: He, Xin, et al.
Veröffentlicht: (2025)
von: He, Xin, et al.
Veröffentlicht: (2025)
DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
von: Wu, Zhiyu, et al.
Veröffentlicht: (2024)
von: Wu, Zhiyu, et al.
Veröffentlicht: (2024)
When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models
von: Yakun, Cui, et al.
Veröffentlicht: (2026)
von: Yakun, Cui, et al.
Veröffentlicht: (2026)
HalluciDoctor: Mitigating Hallucinatory Toxicity in Visual Instruction Data
von: Yu, Qifan, et al.
Veröffentlicht: (2023)
von: Yu, Qifan, et al.
Veröffentlicht: (2023)
Tighnari v2: Mitigating Label Noise and Distribution Shift in Multimodal Plant Distribution Prediction via Mixture of Experts and Weakly Supervised Learning
von: Liu, Haixu, et al.
Veröffentlicht: (2026)
von: Liu, Haixu, et al.
Veröffentlicht: (2026)
Quant Experts: Token-aware Adaptive Error Reconstruction with Mixture of Experts for Large Vision-Language Models Quantization
von: Jia, Chenwei, et al.
Veröffentlicht: (2026)
von: Jia, Chenwei, et al.
Veröffentlicht: (2026)
DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image Generation
von: Ma, Zehong, et al.
Veröffentlicht: (2025)
von: Ma, Zehong, et al.
Veröffentlicht: (2025)
I2MoE: Interpretable Multimodal Interaction-aware Mixture-of-Experts
von: Xin, Jiayi, et al.
Veröffentlicht: (2025)
von: Xin, Jiayi, et al.
Veröffentlicht: (2025)
Metis-HOME: Hybrid Optimized Mixture-of-Experts for Multimodal Reasoning
von: Lan, Xiaohan, et al.
Veröffentlicht: (2025)
von: Lan, Xiaohan, et al.
Veröffentlicht: (2025)
Fast-iTPN: Integrally Pre-Trained Transformer Pyramid Network with Token Migration
von: Tian, Yunjie, et al.
Veröffentlicht: (2022)
von: Tian, Yunjie, et al.
Veröffentlicht: (2022)
Efficient Mixture-of-Expert for Video-based Driver State and Physiological Multi-task Estimation in Conditional Autonomous Driving
von: Wang, Jiyao, et al.
Veröffentlicht: (2024)
von: Wang, Jiyao, et al.
Veröffentlicht: (2024)
Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-Experts
von: Xu, Haolei, et al.
Veröffentlicht: (2026)
von: Xu, Haolei, et al.
Veröffentlicht: (2026)
Dynamic Multimodal Activation Steering for Hallucination Mitigation in Large Vision-Language Models
von: Yin, Jianghao, et al.
Veröffentlicht: (2026)
von: Yin, Jianghao, et al.
Veröffentlicht: (2026)
Boosting Segment Anything Model Towards Open-Vocabulary Learning
von: Han, Xumeng, et al.
Veröffentlicht: (2023)
von: Han, Xumeng, et al.
Veröffentlicht: (2023)
Exploiting Mixture-of-Experts Redundancy Unlocks Multimodal Generative Abilities
von: Dutt, Raman, et al.
Veröffentlicht: (2025)
von: Dutt, Raman, et al.
Veröffentlicht: (2025)
Routers in Vision Mixture of Experts: An Empirical Study
von: Liu, Tianlin, et al.
Veröffentlicht: (2024)
von: Liu, Tianlin, et al.
Veröffentlicht: (2024)
AD^2-Bench: A Hierarchical CoT Benchmark for MLLM in Autonomous Driving under Adverse Conditions
von: Wei, Zhaoyang, et al.
Veröffentlicht: (2025)
von: Wei, Zhaoyang, et al.
Veröffentlicht: (2025)
A Parameter-Efficient Mixture-of-Experts Framework for Cross-Modal Geo-Localization
von: Li, LinFeng, et al.
Veröffentlicht: (2025)
von: Li, LinFeng, et al.
Veröffentlicht: (2025)
Intelligent Communication Mixture-of-Experts Boosted-Medical Image Segmentation Foundation Model
von: Zhang, Xinwei, et al.
Veröffentlicht: (2025)
von: Zhang, Xinwei, et al.
Veröffentlicht: (2025)
MoPE: Mixture of Prompt Experts for Parameter-Efficient and Scalable Multimodal Fusion
von: Jiang, Ruixiang, et al.
Veröffentlicht: (2024)
von: Jiang, Ruixiang, et al.
Veröffentlicht: (2024)
Prefill-Time Intervention for Mitigating Hallucination in Large Vision-Language Models
von: Zhang, Chengsheng, et al.
Veröffentlicht: (2026)
von: Zhang, Chengsheng, et al.
Veröffentlicht: (2026)
Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal Models
von: Hao, Jitai, et al.
Veröffentlicht: (2025)
von: Hao, Jitai, et al.
Veröffentlicht: (2025)
CarbonNet: How Computer Vision Plays a Role in Climate Change? Application: Learning Geomechanics from Subsurface Geometry of CCS to Mitigate Global Warming
von: Chen, Wei, et al.
Veröffentlicht: (2024)
von: Chen, Wei, et al.
Veröffentlicht: (2024)
Trifuse: Enhancing Attention-Based GUI Grounding via Multimodal Fusion
von: Ma, Longhui, et al.
Veröffentlicht: (2026)
von: Ma, Longhui, et al.
Veröffentlicht: (2026)
DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception
von: Li, Xiaotong, et al.
Veröffentlicht: (2024)
von: Li, Xiaotong, et al.
Veröffentlicht: (2024)
ExpertSim: Fast Particle Detector Simulation Using Mixture-of-Generative-Experts
von: Będkowski, Patryk, et al.
Veröffentlicht: (2025)
von: Będkowski, Patryk, et al.
Veröffentlicht: (2025)
Spatio-Semantic Expert Routing Architecture with Mixture-of-Experts for Referring Image Segmentation
von: Dalaq, Alaa, et al.
Veröffentlicht: (2026)
von: Dalaq, Alaa, et al.
Veröffentlicht: (2026)
Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
AgroGPT: Efficient Agricultural Vision-Language Model with Expert Tuning
von: Awais, Muhammad, et al.
Veröffentlicht: (2024)
von: Awais, Muhammad, et al.
Veröffentlicht: (2024)
Parameter-Efficient Quantized Mixture-of-Experts Meets Vision-Language Instruction Tuning for Semiconductor Electron Micrograph Analysis
von: Srinivas, Sakhinana Sagar, et al.
Veröffentlicht: (2024)
von: Srinivas, Sakhinana Sagar, et al.
Veröffentlicht: (2024)
Learning Robust Stereo Matching in the Wild with Selective Mixture-of-Experts
von: Wang, Yun, et al.
Veröffentlicht: (2025)
von: Wang, Yun, et al.
Veröffentlicht: (2025)
TrajMoE: Scene-Adaptive Trajectory Planning with Mixture of Experts and Reinforcement Learning
von: Xing, Zebin, et al.
Veröffentlicht: (2025)
von: Xing, Zebin, et al.
Veröffentlicht: (2025)
EDIT: Enhancing Vision Transformers by Mitigating Attention Sink through an Encoder-Decoder Architecture
von: Feng, Wenfeng, et al.
Veröffentlicht: (2025)
von: Feng, Wenfeng, et al.
Veröffentlicht: (2025)
TAG-MoE: Task-Aware Gating for Unified Generative Mixture-of-Experts
von: Xu, Yu, et al.
Veröffentlicht: (2026)
von: Xu, Yu, et al.
Veröffentlicht: (2026)
Dynamic Pondering Sparsity-aware Mixture-of-Experts Transformer for Event Stream based Visual Object Tracking
von: Wang, Shiao, et al.
Veröffentlicht: (2026)
von: Wang, Shiao, et al.
Veröffentlicht: (2026)
RS-MoE: A Vision-Language Model with Mixture of Experts for Remote Sensing Image Captioning and Visual Question Answering
von: Lin, Hui, et al.
Veröffentlicht: (2024)
von: Lin, Hui, et al.
Veröffentlicht: (2024)
GazeMoE: Perception of Gaze Target with Mixture-of-Experts
von: Dai, Zhuangzhuang, et al.
Veröffentlicht: (2026)
von: Dai, Zhuangzhuang, et al.
Veröffentlicht: (2026)
VidPrism: Heterogeneous Mixture of Experts for Image-to-Video Transfer
von: Lin, Rui, et al.
Veröffentlicht: (2026)
von: Lin, Rui, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Incorporating Visual Experts to Resolve the Information Loss in Multimodal Large Language Models
von: He, Xin, et al.
Veröffentlicht: (2024) -
ViMoE: An Empirical Study of Designing Vision Mixture-of-Experts
von: Han, Xumeng, et al.
Veröffentlicht: (2024) -
EMMA: Efficient Multimodal Understanding, Generation, and Editing with a Unified Architecture
von: He, Xin, et al.
Veröffentlicht: (2025) -
DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
von: Wu, Zhiyu, et al.
Veröffentlicht: (2024) -
When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models
von: Yakun, Cui, et al.
Veröffentlicht: (2026)