Mitigating Visual Context Degradation in Large Multimodal Models: A Training-Free Decoupled Agentic Framework
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jia, Hongrui, Jiang, Chaoya, Zhang, Shikun, Ye, Wei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
From Blind Spots to Gains: Diagnostic-Driven Iterative Training for Large Multimodal Models
von: Jia, Hongrui, et al.
Veröffentlicht: (2026)
von: Jia, Hongrui, et al.
Veröffentlicht: (2026)
MaVEn: An Effective Multi-granularity Hybrid Visual Encoding Framework for Multimodal Large Language Model
von: Jiang, Chaoya, et al.
Veröffentlicht: (2024)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2024)
SymDPO: Boosting In-Context Learning of Large Multimodal Models with Symbol Demonstration Direct Preference Optimization
von: Jia, Hongrui, et al.
Veröffentlicht: (2024)
von: Jia, Hongrui, et al.
Veröffentlicht: (2024)
EVE: Verifiable Self-Evolution of MLLMs via Executable Visual Transformations
von: Heng, Yongrui, et al.
Veröffentlicht: (2026)
von: Heng, Yongrui, et al.
Veröffentlicht: (2026)
Hallucination Augmented Contrastive Learning for Multimodal Large Language Model
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought
von: Jiang, Chaoya, et al.
Veröffentlicht: (2025)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2025)
TiMix: Text-aware Image Mixing for Effective Vision-Language Pre-training
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
IKOD: Mitigating Visual Attention Degradation in Large Vision-Language Models
von: Yang, Jiabing, et al.
Veröffentlicht: (2025)
von: Yang, Jiabing, et al.
Veröffentlicht: (2025)
OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents
von: Jia, Hongrui, et al.
Veröffentlicht: (2025)
von: Jia, Hongrui, et al.
Veröffentlicht: (2025)
ControlMLLM: Training-Free Visual Prompt Learning for Multimodal Large Language Models
von: Wu, Mingrui, et al.
Veröffentlicht: (2024)
von: Wu, Mingrui, et al.
Veröffentlicht: (2024)
BUS:Efficient and Effective Vision-language Pre-training with Bottom-Up Patch Summarization
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
Training-Free Mitigation of Language Reasoning Degradation After Multimodal Instruction Tuning
von: Ratzlaff, Neale, et al.
Veröffentlicht: (2024)
von: Ratzlaff, Neale, et al.
Veröffentlicht: (2024)
TRIPS: Efficient Vision-and-Language Pre-training with Text-Relevant Image Patch Selection
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
Efficient Vision-and-Language Pre-training with Text-Relevant Image Patch Selection
von: Ye, Wei, et al.
Veröffentlicht: (2024)
von: Ye, Wei, et al.
Veröffentlicht: (2024)
MIBench: Evaluating Multimodal Large Language Models over Multiple Images
von: Liu, Haowei, et al.
Veröffentlicht: (2024)
von: Liu, Haowei, et al.
Veröffentlicht: (2024)
Echo-α: Large Agentic Multimodal Reasoning Model for Ultrasound Interpretation
von: Zhang, Jing, et al.
Veröffentlicht: (2026)
von: Zhang, Jing, et al.
Veröffentlicht: (2026)
Personal Visual Context Learning in Large Multimodal Models
von: Xue, Zihui, et al.
Veröffentlicht: (2026)
von: Xue, Zihui, et al.
Veröffentlicht: (2026)
GAMBIT: A Gamified Jailbreak Framework for Multimodal Large Language Models
von: Hu, Xiangdong, et al.
Veröffentlicht: (2026)
von: Hu, Xiangdong, et al.
Veröffentlicht: (2026)
See It, Say It, Sorted: An Iterative Training-Free Framework for Visually-Grounded Multimodal Reasoning in LVLMs
von: Zhang, Yongchang, et al.
Veröffentlicht: (2026)
von: Zhang, Yongchang, et al.
Veröffentlicht: (2026)
Connecting the Dots: Training-Free Visual Grounding via Agentic Reasoning
von: Luo, Liqin, et al.
Veröffentlicht: (2025)
von: Luo, Liqin, et al.
Veröffentlicht: (2025)
Visual Attention Drifts,but Anchors Hold:Mitigating Hallucination in Multimodal Large Language Models via Cross-Layer Visual Anchors
von: Yang, Chengxu, et al.
Veröffentlicht: (2026)
von: Yang, Chengxu, et al.
Veröffentlicht: (2026)
Decoupling Training-Free Guided Diffusion by ADMM
von: Zhang, Youyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Youyuan, et al.
Veröffentlicht: (2024)
Look Twice Before You Answer: Memory-Space Visual Retracing for Hallucination Mitigation in Multimodal Large Language Models
von: Zou, Xin, et al.
Veröffentlicht: (2024)
von: Zou, Xin, et al.
Veröffentlicht: (2024)
DeepScan: A Training-Free Framework for Visually Grounded Reasoning in Large Vision-Language Models
von: Li, Yangfu, et al.
Veröffentlicht: (2026)
von: Li, Yangfu, et al.
Veröffentlicht: (2026)
Look Carefully: Adaptive Visual Reinforcements in Multimodal Large Language Models for Hallucination Mitigation
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
CRoPS: A Training-Free Hallucination Mitigation Framework for Vision-Language Models
von: Anand, Neeraj, et al.
Veröffentlicht: (2026)
von: Anand, Neeraj, et al.
Veröffentlicht: (2026)
Extracting Visual Facts from Intermediate Layers for Mitigating Hallucinations in Multimodal Large Language Models
von: Zhou, Haoran, et al.
Veröffentlicht: (2025)
von: Zhou, Haoran, et al.
Veröffentlicht: (2025)
Training-Free Image Editing with Visual Context Integration and Concept Alignment
von: Song, Rui, et al.
Veröffentlicht: (2026)
von: Song, Rui, et al.
Veröffentlicht: (2026)
Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models
von: Li, Hengzhuang, et al.
Veröffentlicht: (2025)
von: Li, Hengzhuang, et al.
Veröffentlicht: (2025)
DGL-RSIS: Decoupling Global Spatial Context and Local Class Semantics for Training-Free Remote Sensing Image Segmentation
von: Li, Boyi, et al.
Veröffentlicht: (2025)
von: Li, Boyi, et al.
Veröffentlicht: (2025)
Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models
von: Wang, Wenbin, et al.
Veröffentlicht: (2024)
von: Wang, Wenbin, et al.
Veröffentlicht: (2024)
Parameter-efficient Tuning of Large-scale Multimodal Foundation Model
von: Wang, Haixin, et al.
Veröffentlicht: (2023)
von: Wang, Haixin, et al.
Veröffentlicht: (2023)
DDFusion:Degradation-Decoupled Fusion Framework for Robust Infrared and Visible Images Fusion
von: Zhang, Tianpei, et al.
Veröffentlicht: (2025)
von: Zhang, Tianpei, et al.
Veröffentlicht: (2025)
ModuSeg: Decoupling Object Discovery and Semantic Retrieval for Training-Free Weakly Supervised Segmentation
von: He, Qingze, et al.
Veröffentlicht: (2026)
von: He, Qingze, et al.
Veröffentlicht: (2026)
Bridging the Gap Between Sparsity and Redundancy: A Dual-Decoding Framework with Global Context for Map Inference
von: Shen, Yudong, et al.
Veröffentlicht: (2025)
von: Shen, Yudong, et al.
Veröffentlicht: (2025)
FUSE: Label-Free Image-Event Joint Monocular Depth Estimation via Frequency-Decoupled Alignment and Degradation-Robust Fusion
von: Sun, Pihai, et al.
Veröffentlicht: (2025)
von: Sun, Pihai, et al.
Veröffentlicht: (2025)
Breaking Degradation Coupling: A Structural Entropy Guided Decoupled Framework and Benchmark for Infrared Enhancement
von: Li, Pu, et al.
Veröffentlicht: (2026)
von: Li, Pu, et al.
Veröffentlicht: (2026)
QVAD: A Question-Centric Agentic Framework for Efficient and Training-Free Video Anomaly Detection
von: Bekit, Lokman, et al.
Veröffentlicht: (2026)
von: Bekit, Lokman, et al.
Veröffentlicht: (2026)
Mitigating Hallucinations in Large Vision-Language Models without Performance Degradation
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
TrimTokenator-LC: Towards Adaptive Visual Token Pruning for Large Multimodal Models with Long Contexts
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
From Blind Spots to Gains: Diagnostic-Driven Iterative Training for Large Multimodal Models
von: Jia, Hongrui, et al.
Veröffentlicht: (2026) -
MaVEn: An Effective Multi-granularity Hybrid Visual Encoding Framework for Multimodal Large Language Model
von: Jiang, Chaoya, et al.
Veröffentlicht: (2024) -
SymDPO: Boosting In-Context Learning of Large Multimodal Models with Symbol Demonstration Direct Preference Optimization
von: Jia, Hongrui, et al.
Veröffentlicht: (2024) -
EVE: Verifiable Self-Evolution of MLLMs via Executable Visual Transformations
von: Heng, Yongrui, et al.
Veröffentlicht: (2026) -
Hallucination Augmented Contrastive Learning for Multimodal Large Language Model
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)