Seeing Right but Saying Wrong: Inter- and Intra-Layer Refinement in MLLMs without Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Song, Shezheng, Li, Shasha, Yu, Jie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Where Does Vision Meet Language? Understanding and Refining Visual Fusion in MLLMs via Contrastive Attention
von: Song, Shezheng, et al.
Veröffentlicht: (2026)
von: Song, Shezheng, et al.
Veröffentlicht: (2026)
A Dual-way Enhanced Framework from Text Matching Point of View for Multimodal Entity Linking
von: Song, Shezheng, et al.
Veröffentlicht: (2023)
von: Song, Shezheng, et al.
Veröffentlicht: (2023)
How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model
von: Song, Shezheng, et al.
Veröffentlicht: (2023)
von: Song, Shezheng, et al.
Veröffentlicht: (2023)
DWE+: Dual-Way Matching Enhanced Framework for Multimodal Entity Linking
von: Song, Shezheng, et al.
Veröffentlicht: (2024)
von: Song, Shezheng, et al.
Veröffentlicht: (2024)
Do MLLMs Really See It: Reinforcing Visual Attention in Multimodal LLMs
von: Ou, Siqu, et al.
Veröffentlicht: (2026)
von: Ou, Siqu, et al.
Veröffentlicht: (2026)
Right Looks, Wrong Reasons: Compositional Fidelity in Text-to-Image Generation
von: Vatsa, Mayank, et al.
Veröffentlicht: (2025)
von: Vatsa, Mayank, et al.
Veröffentlicht: (2025)
Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes
von: Ling, Chen, et al.
Veröffentlicht: (2026)
von: Ling, Chen, et al.
Veröffentlicht: (2026)
See it. Say it. Sorted: Agentic System for Compositional Diagram Generation
von: Zhang, Hantao, et al.
Veröffentlicht: (2025)
von: Zhang, Hantao, et al.
Veröffentlicht: (2025)
Can NeRFs See without Cameras?
von: Amballa, Chaitanya, et al.
Veröffentlicht: (2025)
von: Amballa, Chaitanya, et al.
Veröffentlicht: (2025)
I Am Big, You Are Little; I Am Right, You Are Wrong
von: Kelly, David A., et al.
Veröffentlicht: (2025)
von: Kelly, David A., et al.
Veröffentlicht: (2025)
DuetFair: Coupling Inter- and Intra-Subgroup Robustness for Fair Medical Image Segmentation
von: Tian, Yiqi, et al.
Veröffentlicht: (2026)
von: Tian, Yiqi, et al.
Veröffentlicht: (2026)
Dense Connector for MLLMs
von: Yao, Huanjin, et al.
Veröffentlicht: (2024)
von: Yao, Huanjin, et al.
Veröffentlicht: (2024)
RASA: Replace Anyone, Say Anything -- A Training-Free Framework for Audio-Driven and Universal Portrait Video Editing
von: Pan, Tianrui, et al.
Veröffentlicht: (2025)
von: Pan, Tianrui, et al.
Veröffentlicht: (2025)
Intra-view and Inter-view Correlation Guided Multi-view Novel Class Discovery
von: Wan, Xinhang, et al.
Veröffentlicht: (2025)
von: Wan, Xinhang, et al.
Veröffentlicht: (2025)
See&Trek: Training-Free Spatial Prompting for Multimodal Large Language Model
von: Li, Pengteng, et al.
Veröffentlicht: (2025)
von: Li, Pengteng, et al.
Veröffentlicht: (2025)
Seeing is Improving: Visual Feedback for Iterative Text Layout Refinement
von: Guo, Junrong, et al.
Veröffentlicht: (2026)
von: Guo, Junrong, et al.
Veröffentlicht: (2026)
Learning with Alignments: Tackling the Inter- and Intra-domain Shifts for Cross-multidomain Facial Expression Recognition
von: Yang, Yuxiang, et al.
Veröffentlicht: (2024)
von: Yang, Yuxiang, et al.
Veröffentlicht: (2024)
Leveraging Static Relationships for Intra-Type and Inter-Type Message Passing in Video Question Answering
von: Liang, Lili, et al.
Veröffentlicht: (2025)
von: Liang, Lili, et al.
Veröffentlicht: (2025)
From Training-Free to Adaptive: Empirical Insights into MLLMs' Understanding of Detection Information
von: Jiao, Qirui, et al.
Veröffentlicht: (2024)
von: Jiao, Qirui, et al.
Veröffentlicht: (2024)
See Detail Say Clear: Towards Brain CT Report Generation via Pathological Clue-driven Representation Learning
von: Zheng, Chengxin, et al.
Veröffentlicht: (2024)
von: Zheng, Chengxin, et al.
Veröffentlicht: (2024)
MOSABench: Multi-Object Sentiment Analysis Benchmark for Evaluating Multimodal Large Language Models Understanding of Complex Image
von: Song, Shezheng, et al.
Veröffentlicht: (2024)
von: Song, Shezheng, et al.
Veröffentlicht: (2024)
PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension
von: Ouyang, Kun, et al.
Veröffentlicht: (2024)
von: Ouyang, Kun, et al.
Veröffentlicht: (2024)
KAN See In the Dark
von: Ning, Aoxiang, et al.
Veröffentlicht: (2024)
von: Ning, Aoxiang, et al.
Veröffentlicht: (2024)
360° Image Perception with MLLMs: A Comprehensive Benchmark and a Training-Free Method
von: Tran, Huyen T. T., et al.
Veröffentlicht: (2026)
von: Tran, Huyen T. T., et al.
Veröffentlicht: (2026)
Unlocking the Forgery Detection Potential of Vanilla MLLMs: A Novel Training-Free Pipeline
von: Zuo, Rui, et al.
Veröffentlicht: (2025)
von: Zuo, Rui, et al.
Veröffentlicht: (2025)
Refining Pre-Trained Motion Models
von: Sun, Xinglong, et al.
Veröffentlicht: (2024)
von: Sun, Xinglong, et al.
Veröffentlicht: (2024)
Seeing Clearly by Layer Two: Enhancing Attention Heads to Alleviate Hallucination in LVLMs
von: Zhang, Xiaofeng, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaofeng, et al.
Veröffentlicht: (2024)
Exploring the Underwater World Segmentation without Extra Training
von: Li, Bingyu, et al.
Veröffentlicht: (2025)
von: Li, Bingyu, et al.
Veröffentlicht: (2025)
CrystaL: Spontaneous Emergence of Visual Latents in MLLMs
von: Zhang, Yang, et al.
Veröffentlicht: (2026)
von: Zhang, Yang, et al.
Veröffentlicht: (2026)
Fast Wrong-way Cycling Detection in CCTV Videos: Sparse Sampling is All You Need
von: Xu, Jing, et al.
Veröffentlicht: (2024)
von: Xu, Jing, et al.
Veröffentlicht: (2024)
Automated Radiographic Total Sharp Score (ARTSS) in Rheumatoid Arthritis: A Solution to Reduce Inter-Intra Reader Variation and Enhancing Clinical Practice
von: Moradmand, Hajar, et al.
Veröffentlicht: (2025)
von: Moradmand, Hajar, et al.
Veröffentlicht: (2025)
Decomposing and Fusing Intra- and Inter-Sensor Spatio-Temporal Signal for Multi-Sensor Wearable Human Activity Recognition
von: Xie, Haoyu, et al.
Veröffentlicht: (2025)
von: Xie, Haoyu, et al.
Veröffentlicht: (2025)
Inter- and Intra-image Refinement for Few Shot Segmentation
von: Fu, Ourui, et al.
Veröffentlicht: (2025)
von: Fu, Ourui, et al.
Veröffentlicht: (2025)
See in Depth: Training-Free Surgical Scene Segmentation with Monocular Depth Priors
von: Yang, Kunyi, et al.
Veröffentlicht: (2025)
von: Yang, Kunyi, et al.
Veröffentlicht: (2025)
SOWing Information: Cultivating Contextual Coherence with MLLMs in Image Generation
von: Pei, Yuhan, et al.
Veröffentlicht: (2024)
von: Pei, Yuhan, et al.
Veröffentlicht: (2024)
Right this way: Can VLMs Guide Us to See More to Answer Questions?
von: Liu, Li, et al.
Veröffentlicht: (2024)
von: Liu, Li, et al.
Veröffentlicht: (2024)
Med-Scout: Curing MLLMs' Geometric Blindness in Medical Perception via Geometry-Aware RL Post-Training
von: Liu, Anglin, et al.
Veröffentlicht: (2026)
von: Liu, Anglin, et al.
Veröffentlicht: (2026)
Thinking Ahead: Foresight Intelligence in MLLMs and World Models
von: Gong, Zhantao, et al.
Veröffentlicht: (2025)
von: Gong, Zhantao, et al.
Veröffentlicht: (2025)
Unmasking Bias in Diffusion Model Training
von: Yu, Hu, et al.
Veröffentlicht: (2023)
von: Yu, Hu, et al.
Veröffentlicht: (2023)
IDPruner: Harmonizing Importance and Diversity in Visual Token Pruning for MLLMs
von: Tan, Yifan, et al.
Veröffentlicht: (2026)
von: Tan, Yifan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Where Does Vision Meet Language? Understanding and Refining Visual Fusion in MLLMs via Contrastive Attention
von: Song, Shezheng, et al.
Veröffentlicht: (2026) -
A Dual-way Enhanced Framework from Text Matching Point of View for Multimodal Entity Linking
von: Song, Shezheng, et al.
Veröffentlicht: (2023) -
How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model
von: Song, Shezheng, et al.
Veröffentlicht: (2023) -
DWE+: Dual-Way Matching Enhanced Framework for Multimodal Entity Linking
von: Song, Shezheng, et al.
Veröffentlicht: (2024) -
Do MLLMs Really See It: Reinforcing Visual Attention in Multimodal LLMs
von: Ou, Siqu, et al.
Veröffentlicht: (2026)