Revisiting Visual Understanding in Multimodal Reasoning through a Lens of Image Perturbation
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Yuting, Wei, Lai, Zheng, Kaipeng, Huang, Jingyuan, Li, Guilin, Wang, Bo, Kong, Linghe, Sun, Lichao, Huang, Weiran |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start
by: Wei, Lai, et al.
Published: (2025)
by: Wei, Lai, et al.
Published: (2025)
TransMed: Large Language Models Enhance Vision Transformer for Biomedical Image Classification
by: Zheng, Kaipeng, et al.
Published: (2023)
by: Zheng, Kaipeng, et al.
Published: (2023)
First SFT, Second RL, Third UPT: Continual Improving Multi-Modal LLM Reasoning via Unsupervised Post-Training
by: Wei, Lai, et al.
Published: (2025)
by: Wei, Lai, et al.
Published: (2025)
IDER: IDempotent Experience Replay for Reliable Continual Learning
by: Liu, Zhanwang, et al.
Published: (2026)
by: Liu, Zhanwang, et al.
Published: (2026)
Enhanced Continual Learning of Vision-Language Models with Model Fusion
by: Gao, Haoyuan, et al.
Published: (2025)
by: Gao, Haoyuan, et al.
Published: (2025)
OTMatch: Improving Semi-Supervised Learning with Optimal Transport
by: Tan, Zhiquan, et al.
Published: (2023)
by: Tan, Zhiquan, et al.
Published: (2023)
MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets
by: Wei, Lai, et al.
Published: (2023)
by: Wei, Lai, et al.
Published: (2023)
TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones
by: Yuan, Zhengqing, et al.
Published: (2023)
by: Yuan, Zhengqing, et al.
Published: (2023)
Improved Iterative Refinement for Chart-to-Code Generation via Structured Instruction
by: Xu, Chengzhi, et al.
Published: (2025)
by: Xu, Chengzhi, et al.
Published: (2025)
Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception
by: Wei, Lai, et al.
Published: (2026)
by: Wei, Lai, et al.
Published: (2026)
Revisiting MLLM Token Technology through the Lens of Classical Visual Coding
by: Liu, Jinming, et al.
Published: (2025)
by: Liu, Jinming, et al.
Published: (2025)
Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination
by: Zheng, Haojie, et al.
Published: (2024)
by: Zheng, Haojie, et al.
Published: (2024)
Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models
by: Huang, Xin, et al.
Published: (2025)
by: Huang, Xin, et al.
Published: (2025)
Provable Contrastive Continual Learning
by: Wen, Yichen, et al.
Published: (2024)
by: Wen, Yichen, et al.
Published: (2024)
LLM4Brain: Training a Large Language Model for Brain Video Understanding
by: Zheng, Ruizhe, et al.
Published: (2024)
by: Zheng, Ruizhe, et al.
Published: (2024)
Recursive Generalization Transformer for Image Super-Resolution
by: Chen, Zheng, et al.
Published: (2023)
by: Chen, Zheng, et al.
Published: (2023)
ChartSketcher: Reasoning with Multimodal Feedback and Reflection for Chart Understanding
by: Huang, Muye, et al.
Published: (2025)
by: Huang, Muye, et al.
Published: (2025)
On Revisiting Entropy for Identifying Mislabeled Images
by: Li, Chunlei, et al.
Published: (2026)
by: Li, Chunlei, et al.
Published: (2026)
Understanding Visual Feature Reliance through the Lens of Complexity
by: Fel, Thomas, et al.
Published: (2024)
by: Fel, Thomas, et al.
Published: (2024)
TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
by: Wang, Shihao, et al.
Published: (2025)
by: Wang, Shihao, et al.
Published: (2025)
ARMOR: Empowering Multimodal Understanding Model with Interleaved Multimodal Generation Capability
by: Sun, Jianwen, et al.
Published: (2025)
by: Sun, Jianwen, et al.
Published: (2025)
AdaSVD: Adaptive Singular Value Decomposition for Large Language Models
by: Li, Zhiteng, et al.
Published: (2025)
by: Li, Zhiteng, et al.
Published: (2025)
PANTHER: Generative Pretraining Beyond Language for Sequential User Behavior Modeling
by: Li, Guilin, et al.
Published: (2025)
by: Li, Guilin, et al.
Published: (2025)
CondiQuant: Condition Number Based Low-Bit Quantization for Image Super-Resolution
by: Liu, Kai, et al.
Published: (2025)
by: Liu, Kai, et al.
Published: (2025)
UM-Text: A Unified Multimodal Model for Image Understanding and Visual Text Editing
by: Ma, Lichen, et al.
Published: (2026)
by: Ma, Lichen, et al.
Published: (2026)
Improving SAM for Camouflaged Object Detection via Dual Stream Adapters
by: Liu, Jiaming, et al.
Published: (2025)
by: Liu, Jiaming, et al.
Published: (2025)
VRSA: Jailbreaking Multimodal Large Language Models through Visual Reasoning Sequential Attack
by: Zhao, Shiji, et al.
Published: (2025)
by: Zhao, Shiji, et al.
Published: (2025)
MoviePuzzle: Visual Narrative Reasoning through Multimodal Order Learning
by: Wang, Jianghui, et al.
Published: (2023)
by: Wang, Jianghui, et al.
Published: (2023)
Both Text and Images Leaked! A Systematic Analysis of Data Contamination in Multimodal LLM
by: Song, Dingjie, et al.
Published: (2024)
by: Song, Dingjie, et al.
Published: (2024)
RADAR: A Risk-Aware Dynamic Multi-Agent Framework for LLM Safety Evaluation via Role-Specialized Collaboration
by: Chen, Xiuyuan, et al.
Published: (2025)
by: Chen, Xiuyuan, et al.
Published: (2025)
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models
by: Zhan, Yufei, et al.
Published: (2025)
by: Zhan, Yufei, et al.
Published: (2025)
VisualLens: Personalization through Task-Agnostic Visual History
by: Zhu, Wang Bill, et al.
Published: (2024)
by: Zhu, Wang Bill, et al.
Published: (2024)
Image Super-Resolution with Text Prompt Diffusion
by: Chen, Zheng, et al.
Published: (2023)
by: Chen, Zheng, et al.
Published: (2023)
Binarized Diffusion Model for Image Super-Resolution
by: Chen, Zheng, et al.
Published: (2024)
by: Chen, Zheng, et al.
Published: (2024)
Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval
by: Tu, Rong-Cheng, et al.
Published: (2025)
by: Tu, Rong-Cheng, et al.
Published: (2025)
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding
by: Yu, Zhuoran, et al.
Published: (2025)
by: Yu, Zhuoran, et al.
Published: (2025)
QuantCache: Adaptive Importance-Guided Quantization with Hierarchical Latent and Layer Caching for Video Generation
by: Wu, Junyi, et al.
Published: (2025)
by: Wu, Junyi, et al.
Published: (2025)
Hierarchical Fusion and Joint Aggregation: A Multi-Level Feature Representation Method for AIGC Image Quality Assessment
by: Meng, Linghe, et al.
Published: (2025)
by: Meng, Linghe, et al.
Published: (2025)
UniModel: A Visual-Only Framework for Unified Multimodal Understanding and Generation
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
Similar Items
-
Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start
by: Wei, Lai, et al.
Published: (2025) -
TransMed: Large Language Models Enhance Vision Transformer for Biomedical Image Classification
by: Zheng, Kaipeng, et al.
Published: (2023) -
First SFT, Second RL, Third UPT: Continual Improving Multi-Modal LLM Reasoning via Unsupervised Post-Training
by: Wei, Lai, et al.
Published: (2025) -
IDER: IDempotent Experience Replay for Reliable Continual Learning
by: Liu, Zhanwang, et al.
Published: (2026) -
Enhanced Continual Learning of Vision-Language Models with Model Fusion
by: Gao, Haoyuan, et al.
Published: (2025)