Investigating Redundancy in Multimodal Large Language Models with Multiple Vision Encoders
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yizhou, Mao, Song, Chen, Yang, Shen, Yufan, Yan, Yinqiao, Cai, Pinlong, Wang, Ding, Yan, Guohang, Yu, Zhi, Hu, Xuming, Shi, Botian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DeepWriter: A Fact-Grounded Multimodal Writing Assistant Based On Offline Knowledge Base
by: Mao, Song, et al.
Published: (2025)
by: Mao, Song, et al.
Published: (2025)
Aligning Vision to Language: Annotation-Free Multimodal Knowledge Graph Construction for Enhanced LLMs Reasoning
by: Liu, Junming, et al.
Published: (2025)
by: Liu, Junming, et al.
Published: (2025)
LeanRAG: Knowledge-Graph-Based Generation with Semantic Aggregation and Hierarchical Retrieval
by: Zhang, Yaoze, et al.
Published: (2025)
by: Zhang, Yaoze, et al.
Published: (2025)
From Ranking to Selection: A Simple but Efficient Dynamic Passage Selector for Retrieval Augmented Generation
by: Meng, Siyuan, et al.
Published: (2025)
by: Meng, Siyuan, et al.
Published: (2025)
RAKG:Document-level Retrieval Augmented Knowledge Graph Construction
by: Zhang, Hairong, et al.
Published: (2025)
by: Zhang, Hairong, et al.
Published: (2025)
HetaRAG: Hybrid Deep Retrieval-Augmented Generation across Heterogeneous Data Stores
by: Yan, Guohang, et al.
Published: (2025)
by: Yan, Guohang, et al.
Published: (2025)
TrafficMCTS: A Closed-Loop Traffic Flow Generation Framework with Group-Based Monte Carlo Tree Search
by: Fu, Ze, et al.
Published: (2023)
by: Fu, Ze, et al.
Published: (2023)
LimSim++: A Closed-Loop Platform for Deploying Multimodal LLMs in Autonomous Driving
by: Fu, Daocheng, et al.
Published: (2024)
by: Fu, Daocheng, et al.
Published: (2024)
KG-TRACES: Enhancing Large Language Models with Knowledge Graph-constrained Trajectory Reasoning and Attribution Supervision
by: Wu, Rong, et al.
Published: (2025)
by: Wu, Rong, et al.
Published: (2025)
MELLA: Bridging Linguistic Capability and Cultural Groundedness for Low-Resource Language MLLMs
by: Gao, Yufei, et al.
Published: (2025)
by: Gao, Yufei, et al.
Published: (2025)
UR-Bench: A Benchmark for Multi-Hop Reasoning over Ultra-High-Resolution Images
by: Li, Siqi, et al.
Published: (2025)
by: Li, Siqi, et al.
Published: (2025)
LimSim Series: An Autonomous Driving Simulation Platform for Validation and Enhancement
by: Fu, Daocheng, et al.
Published: (2025)
by: Fu, Daocheng, et al.
Published: (2025)
GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling
by: Li, Siqi, et al.
Published: (2025)
by: Li, Siqi, et al.
Published: (2025)
EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle
by: Wu, Rong, et al.
Published: (2025)
by: Wu, Rong, et al.
Published: (2025)
MMNeuron: Discovering Neuron-Level Domain-Specific Interpretation in Multimodal Large Language Model
by: Huo, Jiahao, et al.
Published: (2024)
by: Huo, Jiahao, et al.
Published: (2024)
DiLu: A Knowledge-Driven Approach to Autonomous Driving with Large Language Models
by: Wen, Licheng, et al.
Published: (2023)
by: Wen, Licheng, et al.
Published: (2023)
MINER: Mining the Underlying Pattern of Modality-Specific Neurons in Multimodal Large Language Models
by: Huang, Kaichen, et al.
Published: (2024)
by: Huang, Kaichen, et al.
Published: (2024)
Mitigating Modality Prior-Induced Hallucinations in Multimodal Large Language Models via Deciphering Attention Causality
by: Zhou, Guanyu, et al.
Published: (2024)
by: Zhou, Guanyu, et al.
Published: (2024)
SymDrive: Realistic and Controllable Driving Simulator via Symmetric Auto-regressive Online Restoration
by: Liu, Zhiyuan, et al.
Published: (2025)
by: Liu, Zhiyuan, et al.
Published: (2025)
MMUnlearner: Reformulating Multimodal Machine Unlearning in the Era of Multimodal Large Language Models
by: Huo, Jiahao, et al.
Published: (2025)
by: Huo, Jiahao, et al.
Published: (2025)
Reefknot: A Comprehensive Benchmark for Relation Hallucination Evaluation, Analysis and Mitigation in Multimodal Large Language Models
by: Zheng, Kening, et al.
Published: (2024)
by: Zheng, Kening, et al.
Published: (2024)
MemVerse: Multimodal Memory for Lifelong Learning Agents
by: Liu, Junming, et al.
Published: (2025)
by: Liu, Junming, et al.
Published: (2025)
A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges
by: Yan, Yibo, et al.
Published: (2024)
by: Yan, Yibo, et al.
Published: (2024)
Look Twice Before You Answer: Memory-Space Visual Retracing for Hallucination Mitigation in Multimodal Large Language Models
by: Zou, Xin, et al.
Published: (2024)
by: Zou, Xin, et al.
Published: (2024)
Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning
by: Yan, Yibo, et al.
Published: (2025)
by: Yan, Yibo, et al.
Published: (2025)
DocGenome: An Open Large-scale Scientific Document Benchmark for Training and Testing Multi-modal Large Language Models
by: Xia, Renqiu, et al.
Published: (2024)
by: Xia, Renqiu, et al.
Published: (2024)
VEAttack: Downstream-agnostic Vision Encoder Attack against Large Vision Language Models
by: Mei, Hefei, et al.
Published: (2025)
by: Mei, Hefei, et al.
Published: (2025)
Learning on the Job: An Experience-Driven Self-Evolving Agent for Long-Horizon Tasks
by: Yang, Cheng, et al.
Published: (2025)
by: Yang, Cheng, et al.
Published: (2025)
RE-Searcher: Robust Agentic Search with Goal-oriented Planning and Self-reflection
by: Fu, Daocheng, et al.
Published: (2025)
by: Fu, Daocheng, et al.
Published: (2025)
A Redundancy Management Method Based on System State Evaluation for a Redundant Flight Control System
by: An Wang, et al.
Published: (2025)
by: An Wang, et al.
Published: (2025)
The Agent's First Day: Benchmarking Learning, Exploration, and Scheduling in the Workplace Scenarios
by: Fu, Daocheng, et al.
Published: (2026)
by: Fu, Daocheng, et al.
Published: (2026)
MathAgent: Leveraging a Mixture-of-Math-Agent Framework for Real-World Multimodal Mathematical Error Detection
by: Yan, Yibo, et al.
Published: (2025)
by: Yan, Yibo, et al.
Published: (2025)
Renaissance: Investigating the Pretraining of Vision-Language Encoders
by: Fields, Clayton, et al.
Published: (2024)
by: Fields, Clayton, et al.
Published: (2024)
OASim: an Open and Adaptive Simulator based on Neural Rendering for Autonomous Driving
by: Yan, Guohang, et al.
Published: (2024)
by: Yan, Guohang, et al.
Published: (2024)
Vision Function Layer in Multimodal LLMs
by: Shi, Cheng, et al.
Published: (2025)
by: Shi, Cheng, et al.
Published: (2025)
GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning
by: Zhang, Jianghangfan, et al.
Published: (2025)
by: Zhang, Jianghangfan, et al.
Published: (2025)
AlignMMBench: Evaluating Chinese Multimodal Alignment in Large Vision-Language Models
by: Wu, Yuhang, et al.
Published: (2024)
by: Wu, Yuhang, et al.
Published: (2024)
Dearomatic Sulfur‐Shifted Ene Reaction of 3‐Thiiranylbenzo[b]Thiophenes and Ketenes: Further Studies and Experimental Evidence
by: Yinqiao Wang, et al.
Published: (2025)
by: Yinqiao Wang, et al.
Published: (2025)
SafeEraser: Enhancing Safety in Multimodal Large Language Models through Multimodal Machine Unlearning
by: Chen, Junkai, et al.
Published: (2025)
by: Chen, Junkai, et al.
Published: (2025)
MultiFoodhat: A potential new paradigm for intelligent food quality inspection
by: Hu, Yue, et al.
Published: (2025)
by: Hu, Yue, et al.
Published: (2025)
Similar Items
-
DeepWriter: A Fact-Grounded Multimodal Writing Assistant Based On Offline Knowledge Base
by: Mao, Song, et al.
Published: (2025) -
Aligning Vision to Language: Annotation-Free Multimodal Knowledge Graph Construction for Enhanced LLMs Reasoning
by: Liu, Junming, et al.
Published: (2025) -
LeanRAG: Knowledge-Graph-Based Generation with Semantic Aggregation and Hierarchical Retrieval
by: Zhang, Yaoze, et al.
Published: (2025) -
From Ranking to Selection: A Simple but Efficient Dynamic Passage Selector for Retrieval Augmented Generation
by: Meng, Siyuan, et al.
Published: (2025) -
RAKG:Document-level Retrieval Augmented Knowledge Graph Construction
by: Zhang, Hairong, et al.
Published: (2025)