Seeing Clearly without Training: Mitigating Hallucinations in Multimodal LLMs for Remote Sensing
Fuente:
arXiv
Guardado en:
| Autores principales: | Liu, Yi, Zhang, Jing, Wang, Di, Tian, Xiaoyu, Guo, Haonan, Du, Bo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
VLRS-Bench: A Vision-Language Reasoning Benchmark for Remote Sensing
por: Luo, Zhiming, et al.
Publicado: (2026)
por: Luo, Zhiming, et al.
Publicado: (2026)
Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding
por: Tang, Feilong, et al.
Publicado: (2025)
por: Tang, Feilong, et al.
Publicado: (2025)
Remote Sensing ChatGPT: Solving Remote Sensing Tasks with ChatGPT and Visual Models
por: Guo, Haonan, et al.
Publicado: (2024)
por: Guo, Haonan, et al.
Publicado: (2024)
Any2Any: Unified Arbitrary Modality Translation for Remote Sensing
por: Chen, Haoyang, et al.
Publicado: (2026)
por: Chen, Haoyang, et al.
Publicado: (2026)
MTP: Advancing Remote Sensing Foundation Model via Multi-Task Pretraining
por: Wang, Di, et al.
Publicado: (2024)
por: Wang, Di, et al.
Publicado: (2024)
Can Multimodal LLMs See Materials Clearly? A Multimodal Benchmark on Materials Characterization
por: Lai, Zhengzhao, et al.
Publicado: (2025)
por: Lai, Zhengzhao, et al.
Publicado: (2025)
Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language Models
por: He, Zhentao, et al.
Publicado: (2025)
por: He, Zhentao, et al.
Publicado: (2025)
Tracing and Mitigating Hallucinations in Multimodal LLMs via Dynamic Attention Localization
por: Yang, Tiancheng, et al.
Publicado: (2025)
por: Yang, Tiancheng, et al.
Publicado: (2025)
Building-road Collaborative Extraction from Remotely Sensed Images via Cross-Interaction
por: Guo, Haonan, et al.
Publicado: (2023)
por: Guo, Haonan, et al.
Publicado: (2023)
XLRS-Bench: Could Your Multimodal LLMs Understand Extremely Large Ultra-High-Resolution Remote Sensing Imagery?
por: Wang, Fengxiang, et al.
Publicado: (2025)
por: Wang, Fengxiang, et al.
Publicado: (2025)
Measuring and Mitigating Hallucinations in Vision-Language Dataset Generation for Remote Sensing
por: Anderson, Madeline, et al.
Publicado: (2025)
por: Anderson, Madeline, et al.
Publicado: (2025)
See Further When Clear: Curriculum Consistency Model
por: Liu, Yunpeng, et al.
Publicado: (2024)
por: Liu, Yunpeng, et al.
Publicado: (2024)
LeMeViT: Efficient Vision Transformer with Learnable Meta Tokens for Remote Sensing Image Interpretation
por: Jiang, Wentao, et al.
Publicado: (2024)
por: Jiang, Wentao, et al.
Publicado: (2024)
FUSAR-KLIP: Towards Multimodal Foundation Models for Remote Sensing
por: Yang, Yi, et al.
Publicado: (2025)
por: Yang, Yi, et al.
Publicado: (2025)
Expediting Building Footprint Extraction from High-resolution Remote Sensing Images via progressive lenient supervision
por: Guo, Haonan, et al.
Publicado: (2023)
por: Guo, Haonan, et al.
Publicado: (2023)
ClearSight: Visual Signal Enhancement for Object Hallucination Mitigation in Multimodal Large language Models
por: Yin, Hao, et al.
Publicado: (2025)
por: Yin, Hao, et al.
Publicado: (2025)
Mitigating Hallucination in Multimodal LLMs with Layer Contrastive Decoding
por: Tong, Bingkui, et al.
Publicado: (2025)
por: Tong, Bingkui, et al.
Publicado: (2025)
Unveiling the Ignorance of MLLMs: Seeing Clearly, Answering Incorrectly
por: Liu, Yexin, et al.
Publicado: (2024)
por: Liu, Yexin, et al.
Publicado: (2024)
UHR-Micro: Diagnosing and Mitigating the Resolution Illusion in Earth Observation VLMs
por: Ni, Shuo, et al.
Publicado: (2026)
por: Ni, Shuo, et al.
Publicado: (2026)
GeoLLaVA-8K: Scaling Remote-Sensing Multimodal Large Language Models to 8K Resolution
por: Wang, Fengxiang, et al.
Publicado: (2025)
por: Wang, Fengxiang, et al.
Publicado: (2025)
Seeing Clearly by Layer Two: Enhancing Attention Heads to Alleviate Hallucination in LVLMs
por: Zhang, Xiaofeng, et al.
Publicado: (2024)
por: Zhang, Xiaofeng, et al.
Publicado: (2024)
ReLoop: "Seeing Twice and Thinking Backwards" via Closed-loop Training to Mitigate Hallucinations in Multimodal understanding
por: Yang, Jianjiang, et al.
Publicado: (2025)
por: Yang, Jianjiang, et al.
Publicado: (2025)
RSHallu: Dual-Mode Hallucination Evaluation for Remote-Sensing Multimodal Large Language Models with Domain-Tailored Mitigation
por: Zhou, Zihui, et al.
Publicado: (2026)
por: Zhou, Zihui, et al.
Publicado: (2026)
DGSolver: Diffusion Generalist Solver with Universal Posterior Sampling for Image Restoration
por: Wang, Hebaixu, et al.
Publicado: (2025)
por: Wang, Hebaixu, et al.
Publicado: (2025)
See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs
por: Dai, Ziyun, et al.
Publicado: (2025)
por: Dai, Ziyun, et al.
Publicado: (2025)
Boosting Semi-Supervised Object Detection in Remote Sensing Images With Active Teaching
por: Zhang, Boxuan, et al.
Publicado: (2024)
por: Zhang, Boxuan, et al.
Publicado: (2024)
S5: Scalable Semi-Supervised Semantic Segmentation in Remote Sensing
por: Lv, Liang, et al.
Publicado: (2025)
por: Lv, Liang, et al.
Publicado: (2025)
Text-Guided Layer Fusion Mitigates Hallucination in Multimodal LLMs
por: Lin, Chenchen, et al.
Publicado: (2026)
por: Lin, Chenchen, et al.
Publicado: (2026)
Understanding and Mitigating Hallucinations in Multimodal Chain-of-Thought Models
por: Ma, Ji, et al.
Publicado: (2026)
por: Ma, Ji, et al.
Publicado: (2026)
SPEX: A Vision-Language Model for Land Cover Extraction on Spectral Remote Sensing Images
por: Si, Dongchen, et al.
Publicado: (2025)
por: Si, Dongchen, et al.
Publicado: (2025)
MM-Snowball: Evaluating and Mitigating Hallucination Snowballing in Multimodal Multi-Turn Dialogue
por: Jiang, Yue, et al.
Publicado: (2026)
por: Jiang, Yue, et al.
Publicado: (2026)
Seeing Through the Chain: Mitigate Hallucination in Multimodal Reasoning Models via CoT Compression and Contrastive Preference Optimization
por: Fang, Hao, et al.
Publicado: (2026)
por: Fang, Hao, et al.
Publicado: (2026)
A Survey on Remote Sensing Foundation Models: From Vision to Multimodality
por: Huang, Ziyue, et al.
Publicado: (2025)
por: Huang, Ziyue, et al.
Publicado: (2025)
MapGlue: Multimodal Remote Sensing Image Matching
por: Wu, Peihao, et al.
Publicado: (2025)
por: Wu, Peihao, et al.
Publicado: (2025)
Universal Pansharpening Foundation Model
por: Wang, Hebaixu, et al.
Publicado: (2026)
por: Wang, Hebaixu, et al.
Publicado: (2026)
Look Carefully: Adaptive Visual Reinforcements in Multimodal Large Language Models for Hallucination Mitigation
por: Zhu, Xingyu, et al.
Publicado: (2026)
por: Zhu, Xingyu, et al.
Publicado: (2026)
PerturboLLaVA: Reducing Multimodal Hallucinations with Perturbative Visual Training
por: Chen, Cong, et al.
Publicado: (2025)
por: Chen, Cong, et al.
Publicado: (2025)
Boosting Multimodal Remote Sensing Image Classification with Transformer-based Heterogeneously Salient Graph Representation
por: Yang, Jiaqi, et al.
Publicado: (2023)
por: Yang, Jiaqi, et al.
Publicado: (2023)
Residual Diffusion Bridge Model for Image Restoration
por: Wang, Hebaixu, et al.
Publicado: (2025)
por: Wang, Hebaixu, et al.
Publicado: (2025)
ChangeBridge: Spatiotemporal Image Generation with Multimodal Controls for Remote Sensing
por: Zhao, Zhenghui, et al.
Publicado: (2025)
por: Zhao, Zhenghui, et al.
Publicado: (2025)
Ejemplares similares
-
VLRS-Bench: A Vision-Language Reasoning Benchmark for Remote Sensing
por: Luo, Zhiming, et al.
Publicado: (2026) -
Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding
por: Tang, Feilong, et al.
Publicado: (2025) -
Remote Sensing ChatGPT: Solving Remote Sensing Tasks with ChatGPT and Visual Models
por: Guo, Haonan, et al.
Publicado: (2024) -
Any2Any: Unified Arbitrary Modality Translation for Remote Sensing
por: Chen, Haoyang, et al.
Publicado: (2026) -
MTP: Advancing Remote Sensing Foundation Model via Multi-Task Pretraining
por: Wang, Di, et al.
Publicado: (2024)