From Blind Spots to Gains: Diagnostic-Driven Iterative Training for Large Multimodal Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jia, Hongrui, Jiang, Chaoya, Heng, Yongrui, Zhang, Shikun, Ye, Wei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Mitigating Visual Context Degradation in Large Multimodal Models: A Training-Free Decoupled Agentic Framework
von: Jia, Hongrui, et al.
Veröffentlicht: (2025)
von: Jia, Hongrui, et al.
Veröffentlicht: (2025)
EVE: Verifiable Self-Evolution of MLLMs via Executable Visual Transformations
von: Heng, Yongrui, et al.
Veröffentlicht: (2026)
von: Heng, Yongrui, et al.
Veröffentlicht: (2026)
SymDPO: Boosting In-Context Learning of Large Multimodal Models with Symbol Demonstration Direct Preference Optimization
von: Jia, Hongrui, et al.
Veröffentlicht: (2024)
von: Jia, Hongrui, et al.
Veröffentlicht: (2024)
MaVEn: An Effective Multi-granularity Hybrid Visual Encoding Framework for Multimodal Large Language Model
von: Jiang, Chaoya, et al.
Veröffentlicht: (2024)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2024)
VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought
von: Jiang, Chaoya, et al.
Veröffentlicht: (2025)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2025)
Hallucination Augmented Contrastive Learning for Multimodal Large Language Model
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
TiMix: Text-aware Image Mixing for Effective Vision-Language Pre-training
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents
von: Jia, Hongrui, et al.
Veröffentlicht: (2025)
von: Jia, Hongrui, et al.
Veröffentlicht: (2025)
BUS:Efficient and Effective Vision-language Pre-training with Bottom-Up Patch Summarization
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
TRIPS: Efficient Vision-and-Language Pre-training with Text-Relevant Image Patch Selection
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
Efficient Vision-and-Language Pre-training with Text-Relevant Image Patch Selection
von: Ye, Wei, et al.
Veröffentlicht: (2024)
von: Ye, Wei, et al.
Veröffentlicht: (2024)
MIBench: Evaluating Multimodal Large Language Models over Multiple Images
von: Liu, Haowei, et al.
Veröffentlicht: (2024)
von: Liu, Haowei, et al.
Veröffentlicht: (2024)
Spot the Fake: Large Multimodal Model-Based Synthetic Image Detection with Artifact Explanation
von: Wen, Siwei, et al.
Veröffentlicht: (2025)
von: Wen, Siwei, et al.
Veröffentlicht: (2025)
The Blind Spot of Adaptation: Quantifying and Mitigating Forgetting in Fine-tuned Driving Models
von: Mao, Runhao, et al.
Veröffentlicht: (2026)
von: Mao, Runhao, et al.
Veröffentlicht: (2026)
Spatiotemporal Blind-Spot Network with Calibrated Flow Alignment for Self-Supervised Video Denoising
von: Chen, Zikang, et al.
Veröffentlicht: (2024)
von: Chen, Zikang, et al.
Veröffentlicht: (2024)
Temporal Gains, Spatial Costs: Revisiting Video Fine-Tuning in Multimodal Large Language Models
von: Zhang, Linghao, et al.
Veröffentlicht: (2026)
von: Zhang, Linghao, et al.
Veröffentlicht: (2026)
Parameter-efficient Tuning of Large-scale Multimodal Foundation Model
von: Wang, Haixin, et al.
Veröffentlicht: (2023)
von: Wang, Haixin, et al.
Veröffentlicht: (2023)
Single Image Unlearning: Efficient Machine Unlearning in Multimodal Large Language Models
von: Li, Jiaqi, et al.
Veröffentlicht: (2024)
von: Li, Jiaqi, et al.
Veröffentlicht: (2024)
SPHERE: Unveiling Spatial Blind Spots in Vision-Language Models Through Hierarchical Evaluation
von: Zhang, Wenyu, et al.
Veröffentlicht: (2024)
von: Zhang, Wenyu, et al.
Veröffentlicht: (2024)
From Learning to Unlearning: Biomedical Security Protection in Multimodal Large Language Models
von: Xu, Dunyuan, et al.
Veröffentlicht: (2025)
von: Xu, Dunyuan, et al.
Veröffentlicht: (2025)
When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding
von: Shu, Yan, et al.
Veröffentlicht: (2025)
von: Shu, Yan, et al.
Veröffentlicht: (2025)
Skinned Motion Retargeting with Dense Geometric Interaction Perception
von: Ye, Zijie, et al.
Veröffentlicht: (2024)
von: Ye, Zijie, et al.
Veröffentlicht: (2024)
Exploring Efficient Asymmetric Blind-Spots for Self-Supervised Denoising in Real-World Scenarios
von: Chen, Shiyan, et al.
Veröffentlicht: (2023)
von: Chen, Shiyan, et al.
Veröffentlicht: (2023)
Seeing Isn't Believing: Uncovering Blind Spots in Evaluator Vision-Language Models
von: Khan, Mohammed Safi Ur Rahman, et al.
Veröffentlicht: (2026)
von: Khan, Mohammed Safi Ur Rahman, et al.
Veröffentlicht: (2026)
Blind-Spot Guided Diffusion for Self-supervised Real-World Denoising
von: Cheng, Shen, et al.
Veröffentlicht: (2025)
von: Cheng, Shen, et al.
Veröffentlicht: (2025)
MedDiff-FM: A Diffusion-based Foundation Model for Versatile Medical Image Applications
von: Yu, Yongrui, et al.
Veröffentlicht: (2024)
von: Yu, Yongrui, et al.
Veröffentlicht: (2024)
SwinIA: Self-Supervised Blind-Spot Image Denoising without Convolutions
von: Papkov, Mikhail, et al.
Veröffentlicht: (2023)
von: Papkov, Mikhail, et al.
Veröffentlicht: (2023)
GAMBIT: A Gamified Jailbreak Framework for Multimodal Large Language Models
von: Hu, Xiangdong, et al.
Veröffentlicht: (2026)
von: Hu, Xiangdong, et al.
Veröffentlicht: (2026)
HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language Models
von: Kang, Zhaolu, et al.
Veröffentlicht: (2025)
von: Kang, Zhaolu, et al.
Veröffentlicht: (2025)
MIRROR: Multimodal Iterative Reasoning via Reflection on Visual Regions
von: Zhang, Haoyu, et al.
Veröffentlicht: (2026)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2026)
ISR-DPO: Aligning Large Multimodal Models for Videos by Iterative Self-Retrospective DPO
von: Ahn, Daechul, et al.
Veröffentlicht: (2024)
von: Ahn, Daechul, et al.
Veröffentlicht: (2024)
See It, Say It, Sorted: An Iterative Training-Free Framework for Visually-Grounded Multimodal Reasoning in LVLMs
von: Zhang, Yongchang, et al.
Veröffentlicht: (2026)
von: Zhang, Yongchang, et al.
Veröffentlicht: (2026)
LISA: A Layer-wise Integration and Suppression Approach for Hallucination Mitigation in Multimodal Large Language Models
von: Guo, Zhihui, et al.
Veröffentlicht: (2025)
von: Guo, Zhihui, et al.
Veröffentlicht: (2025)
Reflect to Inform: Boosting Multimodal Reasoning via Information-Gain-Driven Verification
von: Lv, Shuai, et al.
Veröffentlicht: (2026)
von: Lv, Shuai, et al.
Veröffentlicht: (2026)
Forgotten Polygons: Multimodal Large Language Models are Shape-Blind
von: Rudman, William, et al.
Veröffentlicht: (2025)
von: Rudman, William, et al.
Veröffentlicht: (2025)
Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models
von: Jung, Woojun, et al.
Veröffentlicht: (2025)
von: Jung, Woojun, et al.
Veröffentlicht: (2025)
NAUTILUS: A Large Multimodal Model for Underwater Scene Understanding
von: Xu, Wei, et al.
Veröffentlicht: (2025)
von: Xu, Wei, et al.
Veröffentlicht: (2025)
Visible Yet Unreadable: A Systematic Blind Spot of Vision Language Models Across Writing Systems
von: Zhang, Jie, et al.
Veröffentlicht: (2025)
von: Zhang, Jie, et al.
Veröffentlicht: (2025)
Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space
von: Trinh, Quoc-Huy, et al.
Veröffentlicht: (2026)
von: Trinh, Quoc-Huy, et al.
Veröffentlicht: (2026)
Focusing Where Vision Matters: Selective Training for Large Vision Language Models via Visual Information Gain
von: Lee, Seulbi, et al.
Veröffentlicht: (2026)
von: Lee, Seulbi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Mitigating Visual Context Degradation in Large Multimodal Models: A Training-Free Decoupled Agentic Framework
von: Jia, Hongrui, et al.
Veröffentlicht: (2025) -
EVE: Verifiable Self-Evolution of MLLMs via Executable Visual Transformations
von: Heng, Yongrui, et al.
Veröffentlicht: (2026) -
SymDPO: Boosting In-Context Learning of Large Multimodal Models with Symbol Demonstration Direct Preference Optimization
von: Jia, Hongrui, et al.
Veröffentlicht: (2024) -
MaVEn: An Effective Multi-granularity Hybrid Visual Encoding Framework for Multimodal Large Language Model
von: Jiang, Chaoya, et al.
Veröffentlicht: (2024) -
VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought
von: Jiang, Chaoya, et al.
Veröffentlicht: (2025)