GHOST: Hallucination-Inducing Image Generation for Multimodal LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Parast, Aryan Yazdan, Hosseini, Parsa, Asadollahzadeh, Hesam, Moakhar, Arshia Soltani, Azam, Basim, Feizi, Soheil, Akhtar, Naveed |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DDB: Diffusion Driven Balancing to Address Spurious Correlations
by: Parast, Aryan Yazdan, et al.
Published: (2025)
by: Parast, Aryan Yazdan, et al.
Published: (2025)
HSFM: Hard-Set-Guided Feature-Space Meta-Learning for Robust Classification under Spurious Correlations
by: Parast, Aryan Yazdan, et al.
Published: (2026)
by: Parast, Aryan Yazdan, et al.
Published: (2026)
Latent Video Prediction Learns Better World Models
by: Alrasheed, Ali J, et al.
Published: (2026)
by: Alrasheed, Ali J, et al.
Published: (2026)
Plug-and-Play Interpretable Responsible Text-to-Image Generation via Dual-Space Multi-facet Concept Control
by: Azam, Basim, et al.
Published: (2025)
by: Azam, Basim, et al.
Published: (2025)
Suitability of KANs for Computer Vision: A preliminary investigation
by: Azam, Basim, et al.
Published: (2024)
by: Azam, Basim, et al.
Published: (2024)
Decompose-and-Compose: A Compositional Approach to Mitigating Spurious Correlation
by: Noohdani, Fahimeh Hosseini, et al.
Published: (2024)
by: Noohdani, Fahimeh Hosseini, et al.
Published: (2024)
SpurLens: Automatic Detection of Spurious Cues in Multimodal LLMs
by: Hosseini, Parsa, et al.
Published: (2025)
by: Hosseini, Parsa, et al.
Published: (2025)
Shortest-Path Flow Matching with Mixture-Conditioned Bases for OOD Generalization to Unseen Conditions
by: Rubbi, Andrea, et al.
Published: (2026)
by: Rubbi, Andrea, et al.
Published: (2026)
GO-N3RDet: Geometry Optimized NeRF-enhanced 3D Object Detector
by: Li, Zechuan, et al.
Published: (2025)
by: Li, Zechuan, et al.
Published: (2025)
Your LLM Agents are Temporally Blind: The Misalignment Between Tool Use Decisions and Human Time Perception
by: Cheng, Yize, et al.
Published: (2025)
by: Cheng, Yize, et al.
Published: (2025)
TRACER: Persistent Regularization for Robust Multimodal Finetuning
by: Asadollahzadeh, Hesam, et al.
Published: (2026)
by: Asadollahzadeh, Hesam, et al.
Published: (2026)
Failing to Explore: Language Models on Interactive Tasks
by: JafariRaviz, Mahdi, et al.
Published: (2026)
by: JafariRaviz, Mahdi, et al.
Published: (2026)
Decomposing and Interpreting Image Representations via Text in ViTs Beyond CLIP
by: Balasubramanian, Sriram, et al.
Published: (2024)
by: Balasubramanian, Sriram, et al.
Published: (2024)
Distilling Knowledge from Text-to-Image Generative Models Improves Visio-Linguistic Reasoning in CLIP
by: Basu, Samyadeep, et al.
Published: (2023)
by: Basu, Samyadeep, et al.
Published: (2023)
IConMark: Robust Interpretable Concept-Based Watermark For AI Images
by: Sadasivan, Vinu Sankar, et al.
Published: (2025)
by: Sadasivan, Vinu Sankar, et al.
Published: (2025)
Exploring Bias in over 100 Text-to-Image Generative Models
by: Vice, Jordan, et al.
Published: (2025)
by: Vice, Jordan, et al.
Published: (2025)
Context-guided Responsible Data Augmentation with Diffusion Models
by: Islam, Khawar, et al.
Published: (2025)
by: Islam, Khawar, et al.
Published: (2025)
AgentComp: From Agentic Reasoning to Compositional Mastery in Text-to-Image Models
by: Zarei, Arman, et al.
Published: (2025)
by: Zarei, Arman, et al.
Published: (2025)
Chain-of-Defensive-Thought: Structured Reasoning Elicits Robustness in Large Language Models against Reference Corruption
by: Wang, Wenxiao, et al.
Published: (2025)
by: Wang, Wenxiao, et al.
Published: (2025)
Mitigating Memorization in Text-to-Image Diffusion via Region-Aware Prompt Augmentation and Multimodal Copy Detection
by: Chen, Yunzhuo, et al.
Published: (2026)
by: Chen, Yunzhuo, et al.
Published: (2026)
Improving Compositional Attribute Binding in Text-to-Image Generative Models via Enhanced Text Embeddings
by: Zarei, Arman, et al.
Published: (2024)
by: Zarei, Arman, et al.
Published: (2024)
On Mechanistic Knowledge Localization in Text-to-Image Generative Models
by: Basu, Samyadeep, et al.
Published: (2024)
by: Basu, Samyadeep, et al.
Published: (2024)
PRIME: Prioritizing Interpretability in Failure Mode Extraction
by: Rezaei, Keivan, et al.
Published: (2023)
by: Rezaei, Keivan, et al.
Published: (2023)
Efficient Diffusion Models for Vision: A Survey
by: Ulhaq, Anwaar, et al.
Published: (2022)
by: Ulhaq, Anwaar, et al.
Published: (2022)
On the Fairness, Diversity and Reliability of Text-to-Image Generative Models
by: Vice, Jordan, et al.
Published: (2024)
by: Vice, Jordan, et al.
Published: (2024)
GenMix: Effective Data Augmentation with Generative Diffusion Model Image Editing
by: Islam, Khawar, et al.
Published: (2024)
by: Islam, Khawar, et al.
Published: (2024)
Hallucination Behavior in Multimodal LLMs Across Agricultural Image Interpretation and Generation Tasks
by: Ghose, Partho, et al.
Published: (2026)
by: Ghose, Partho, et al.
Published: (2026)
SliderEdit: Continuous Image Editing with Fine-Grained Instruction Control
by: Zarei, Arman, et al.
Published: (2025)
by: Zarei, Arman, et al.
Published: (2025)
Rethinking Artistic Copyright Infringements in the Era of Text-to-Image Generative Models
by: Moayeri, Mazda, et al.
Published: (2024)
by: Moayeri, Mazda, et al.
Published: (2024)
Altered Thoughts, Altered Actions: Probing Chain-of-Thought Vulnerabilities in VLA Robotic Manipulation
by: Trinh, Tuan Duong, et al.
Published: (2026)
by: Trinh, Tuan Duong, et al.
Published: (2026)
How Learnable Grids Recover Fine Detail in Low Dimensions: A Neural Tangent Kernel Analysis of Multigrid Parametric Encodings
by: Audia, Samuel, et al.
Published: (2025)
by: Audia, Samuel, et al.
Published: (2025)
GLMHA A Guided Low-rank Multi-Head Self-Attention for Efficient Image Restoration and Spectral Reconstruction
by: Ilyas, Zaid, et al.
Published: (2024)
by: Ilyas, Zaid, et al.
Published: (2024)
Safety Without Semantic Disruptions: Editing-free Safe Image Generation via Context-preserving Dual Latent Reconstruction
by: Vice, Jordan, et al.
Published: (2024)
by: Vice, Jordan, et al.
Published: (2024)
Manipulating and Mitigating Generative Model Biases without Retraining
by: Vice, Jordan, et al.
Published: (2024)
by: Vice, Jordan, et al.
Published: (2024)
Trained Models Tell Us How to Make Them Robust to Spurious Correlation without Group Annotation
by: Ghaznavi, Mahdi, et al.
Published: (2024)
by: Ghaznavi, Mahdi, et al.
Published: (2024)
RODEO: Robust Outlier Detection via Exposing Adaptive Out-of-Distribution Samples
by: Mirzaei, Hossein, et al.
Published: (2025)
by: Mirzaei, Hossein, et al.
Published: (2025)
Robustness of AI-Image Detectors: Fundamental Limits and Practical Attacks
by: Saberi, Mehrdad, et al.
Published: (2023)
by: Saberi, Mehrdad, et al.
Published: (2023)
IntCoOp: Interpretability-Aware Vision-Language Prompt Tuning
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
Mitigating Hallucination in Multimodal LLMs with Layer Contrastive Decoding
by: Tong, Bingkui, et al.
Published: (2025)
by: Tong, Bingkui, et al.
Published: (2025)
HII-DPO: Eliminate Hallucination via Accurate Hallucination-Inducing Counterfactual Images
by: Yang, Yilin, et al.
Published: (2026)
by: Yang, Yilin, et al.
Published: (2026)
Similar Items
-
DDB: Diffusion Driven Balancing to Address Spurious Correlations
by: Parast, Aryan Yazdan, et al.
Published: (2025) -
HSFM: Hard-Set-Guided Feature-Space Meta-Learning for Robust Classification under Spurious Correlations
by: Parast, Aryan Yazdan, et al.
Published: (2026) -
Latent Video Prediction Learns Better World Models
by: Alrasheed, Ali J, et al.
Published: (2026) -
Plug-and-Play Interpretable Responsible Text-to-Image Generation via Dual-Space Multi-facet Concept Control
by: Azam, Basim, et al.
Published: (2025) -
Suitability of KANs for Computer Vision: A preliminary investigation
by: Azam, Basim, et al.
Published: (2024)