Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Xu, Xiao, Qin, Libo, Che, Wanxiang, Kan, Min-Yen |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Exploring Multi-Grained Concept Annotations for Multimodal Large Language Models
di: Xu, Xiao, et al.
Pubblicazione: (2024)
di: Xu, Xiao, et al.
Pubblicazione: (2024)
BridgeTower: Building Bridges Between Encoders in Vision-Language Representation Learning
di: Xu, Xiao, et al.
Pubblicazione: (2022)
di: Xu, Xiao, et al.
Pubblicazione: (2022)
M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought
di: Chen, Qiguang, et al.
Pubblicazione: (2024)
di: Chen, Qiguang, et al.
Pubblicazione: (2024)
What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration
di: Qin, Libo, et al.
Pubblicazione: (2024)
di: Qin, Libo, et al.
Pubblicazione: (2024)
Leveraging NTPs for Efficient Hallucination Detection in VLMs
di: Azachi, Ofir, et al.
Pubblicazione: (2025)
di: Azachi, Ofir, et al.
Pubblicazione: (2025)
Understanding and Rectifying Safety Perception Distortion in VLMs
di: Zou, Xiaohan, et al.
Pubblicazione: (2025)
di: Zou, Xiaohan, et al.
Pubblicazione: (2025)
Sensitivity of Generative VLMs to Semantically and Lexically Altered Prompts
di: Dumpala, Sri Harsha, et al.
Pubblicazione: (2024)
di: Dumpala, Sri Harsha, et al.
Pubblicazione: (2024)
PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs
di: Nasiriany, Soroush, et al.
Pubblicazione: (2024)
di: Nasiriany, Soroush, et al.
Pubblicazione: (2024)
The Instinctive Bias: Spurious Images lead to Illusion in MLLMs
di: Han, Tianyang, et al.
Pubblicazione: (2024)
di: Han, Tianyang, et al.
Pubblicazione: (2024)
Have the VLMs Lost Confidence? A Study of Sycophancy in VLMs
di: Li, Shuo, et al.
Pubblicazione: (2024)
di: Li, Shuo, et al.
Pubblicazione: (2024)
Fine-tuning MLLMs Without Forgetting Is Easier Than You Think
di: Li, He, et al.
Pubblicazione: (2026)
di: Li, He, et al.
Pubblicazione: (2026)
TechING: Towards Real World Technical Image Understanding via VLMs
di: Nadeem, Tafazzul, et al.
Pubblicazione: (2026)
di: Nadeem, Tafazzul, et al.
Pubblicazione: (2026)
Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression
di: Du, Yao, et al.
Pubblicazione: (2026)
di: Du, Yao, et al.
Pubblicazione: (2026)
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models
di: Zhang, Yongheng, et al.
Pubblicazione: (2025)
di: Zhang, Yongheng, et al.
Pubblicazione: (2025)
Exploring How Generative MLLMs Perceive More Than CLIP with the Same Vision Encoder
di: Li, Siting, et al.
Pubblicazione: (2024)
di: Li, Siting, et al.
Pubblicazione: (2024)
VLMGuard-R1: Proactive Safety Alignment for VLMs via Reasoning-Driven Prompt Optimization
di: Chen, Menglan, et al.
Pubblicazione: (2025)
di: Chen, Menglan, et al.
Pubblicazione: (2025)
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?
di: Park, Simon, et al.
Pubblicazione: (2025)
di: Park, Simon, et al.
Pubblicazione: (2025)
D$^{3}$ToM: Decider-Guided Dynamic Token Merging for Accelerating Diffusion MLLMs
di: Chang, Shuochen, et al.
Pubblicazione: (2025)
di: Chang, Shuochen, et al.
Pubblicazione: (2025)
RBF++: Quantifying and Optimizing Reasoning Boundaries across Measurable and Unmeasurable Capabilities for Chain-of-Thought Reasoning
di: Chen, Qiguang, et al.
Pubblicazione: (2025)
di: Chen, Qiguang, et al.
Pubblicazione: (2025)
M2-Reasoning: Empowering MLLMs with Unified General and Spatial Reasoning
di: AI, Inclusion, et al.
Pubblicazione: (2025)
di: AI, Inclusion, et al.
Pubblicazione: (2025)
Visual Thoughts: A Unified Perspective of Understanding Multimodal Chain-of-Thought
di: Cheng, Zihui, et al.
Pubblicazione: (2025)
di: Cheng, Zihui, et al.
Pubblicazione: (2025)
MileBench: Benchmarking MLLMs in Long Context
di: Song, Dingjie, et al.
Pubblicazione: (2024)
di: Song, Dingjie, et al.
Pubblicazione: (2024)
MLLMs-Augmented Visual-Language Representation Learning
di: Liu, Yanqing, et al.
Pubblicazione: (2023)
di: Liu, Yanqing, et al.
Pubblicazione: (2023)
Can World Models Benefit VLMs for World Dynamics?
di: Zhang, Kevin, et al.
Pubblicazione: (2025)
di: Zhang, Kevin, et al.
Pubblicazione: (2025)
xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs
di: Ryoo, Michael S., et al.
Pubblicazione: (2024)
di: Ryoo, Michael S., et al.
Pubblicazione: (2024)
Mixture of Group Experts for Learning Invariant Representations
di: Kang, Lei, et al.
Pubblicazione: (2025)
di: Kang, Lei, et al.
Pubblicazione: (2025)
Through the Lens of Contrast: Self-Improving Visual Reasoning in VLMs
di: Pan, Zhiyu, et al.
Pubblicazione: (2026)
di: Pan, Zhiyu, et al.
Pubblicazione: (2026)
Sparse Autoencoders as Plug-and-Play Firewalls for Adversarial Attack Detection in VLMs
di: Wang, Hao, et al.
Pubblicazione: (2026)
di: Wang, Hao, et al.
Pubblicazione: (2026)
Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models
di: Karamcheti, Siddharth, et al.
Pubblicazione: (2024)
di: Karamcheti, Siddharth, et al.
Pubblicazione: (2024)
Scaling Up Temporal Domain Generalization via Temporal Experts Averaging
di: Liu, Aoming, et al.
Pubblicazione: (2025)
di: Liu, Aoming, et al.
Pubblicazione: (2025)
MemLoRA: Distilling Expert Adapters for On-Device Memory Systems
di: Bini, Massimo, et al.
Pubblicazione: (2025)
di: Bini, Massimo, et al.
Pubblicazione: (2025)
Tables as Texts or Images: Evaluating the Table Reasoning Ability of LLMs and MLLMs
di: Deng, Naihao, et al.
Pubblicazione: (2024)
di: Deng, Naihao, et al.
Pubblicazione: (2024)
Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities
di: Yang, Enneng, et al.
Pubblicazione: (2024)
di: Yang, Enneng, et al.
Pubblicazione: (2024)
Mixture of Experts Made Personalized: Federated Prompt Learning for Vision-Language Models
di: Luo, Jun, et al.
Pubblicazione: (2024)
di: Luo, Jun, et al.
Pubblicazione: (2024)
LiME: Lightweight Mixture of Experts for Efficient Multimodal Multi-task Learning
di: Kowsher, Md, et al.
Pubblicazione: (2026)
di: Kowsher, Md, et al.
Pubblicazione: (2026)
VideoScaffold: Elastic-Scale Visual Hierarchies for Streaming Video Understanding in MLLMs
di: Zheng, Naishan, et al.
Pubblicazione: (2025)
di: Zheng, Naishan, et al.
Pubblicazione: (2025)
Multimodal Representation Learning by Alternating Unimodal Adaptation
di: Zhang, Xiaohui, et al.
Pubblicazione: (2023)
di: Zhang, Xiaohui, et al.
Pubblicazione: (2023)
BEEM: Boosting Performance of Early Exit DNNs using Multi-Exit Classifiers as Experts
di: Bajpai, Divya Jyoti, et al.
Pubblicazione: (2025)
di: Bajpai, Divya Jyoti, et al.
Pubblicazione: (2025)
ChartREG++: Towards Benchmarking and Improving Chart Referring Expression Grounding under Diverse referring clues and Multi-Target Referring
di: Niu, Tianhao, et al.
Pubblicazione: (2026)
di: Niu, Tianhao, et al.
Pubblicazione: (2026)
When Does Sparse MoE Help in Vision? The Role of Backbone Compute Leverage in Sparse Routing
di: Sun, Libo, et al.
Pubblicazione: (2026)
di: Sun, Libo, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Exploring Multi-Grained Concept Annotations for Multimodal Large Language Models
di: Xu, Xiao, et al.
Pubblicazione: (2024) -
BridgeTower: Building Bridges Between Encoders in Vision-Language Representation Learning
di: Xu, Xiao, et al.
Pubblicazione: (2022) -
M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought
di: Chen, Qiguang, et al.
Pubblicazione: (2024) -
What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration
di: Qin, Libo, et al.
Pubblicazione: (2024) -
Leveraging NTPs for Efficient Hallucination Detection in VLMs
di: Azachi, Ofir, et al.
Pubblicazione: (2025)