Mechanisms of Prompt-Induced Hallucination in Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rudman, William, Golovanevsky, Michal, Arad, Dana, Belinkov, Yonatan, Singh, Ritambhara, Eickhoff, Carsten, Mahowald, Kyle |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't
von: Nemitz, Jonathan, et al.
Veröffentlicht: (2026)
von: Nemitz, Jonathan, et al.
Veröffentlicht: (2026)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images
von: Chen, Yuangong, et al.
Veröffentlicht: (2026)
von: Chen, Yuangong, et al.
Veröffentlicht: (2026)
WildRoadBench: A Wild Aerial Road-Damage Grounding Benchmark for Vision-Language Models and Autonomous Agents
von: Liu, Bingnan, et al.
Veröffentlicht: (2026)
von: Liu, Bingnan, et al.
Veröffentlicht: (2026)
Leum-VL Technical Report
von: He, Yuxuan, et al.
Veröffentlicht: (2026)
von: He, Yuxuan, et al.
Veröffentlicht: (2026)
DriveMRP: Enhancing Vision-Language Models with Synthetic Motion Data for Motion Risk Prediction
von: Hou, Zhiyi, et al.
Veröffentlicht: (2025)
von: Hou, Zhiyi, et al.
Veröffentlicht: (2025)
Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning
von: Ji, Binbin, et al.
Veröffentlicht: (2025)
von: Ji, Binbin, et al.
Veröffentlicht: (2025)
StoryMovie: A Dataset for Semantic Alignment of Visual Stories with Movie Scripts and Subtitles
von: Oliveira, Daniel, et al.
Veröffentlicht: (2026)
von: Oliveira, Daniel, et al.
Veröffentlicht: (2026)
GeoVision Labeler: Zero-Shot Geospatial Classification with Vision and Language Models
von: Hacheme, Gilles Quentin, et al.
Veröffentlicht: (2025)
von: Hacheme, Gilles Quentin, et al.
Veröffentlicht: (2025)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025)
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025)
Lifelong Learning in Vision-Language Models: Enhanced EWC with Cross-Modal Knowledge Retention
von: Durrani, Hamza Ahmed, et al.
Veröffentlicht: (2026)
von: Durrani, Hamza Ahmed, et al.
Veröffentlicht: (2026)
GLoT: A Novel Gated-Logarithmic Transformer for Efficient Sign Language Translation
von: Shahin, Nada, et al.
Veröffentlicht: (2025)
von: Shahin, Nada, et al.
Veröffentlicht: (2025)
Reducing Object Hallucination in LVLMs via Emphasizing Image-negative Tokens
von: Shen, Meng, et al.
Veröffentlicht: (2026)
von: Shen, Meng, et al.
Veröffentlicht: (2026)
Decoupling Vision and Language: Codebook Anchored Visual Adaptation
von: Wu, Jason, et al.
Veröffentlicht: (2026)
von: Wu, Jason, et al.
Veröffentlicht: (2026)
NOAH: Benchmarking Narrative Prior driven Hallucination and Omission in Video Large Language Models
von: Lee, Kyuho, et al.
Veröffentlicht: (2025)
von: Lee, Kyuho, et al.
Veröffentlicht: (2025)
FlexDoc: Parameterized Sampling for Diverse Multilingual Synthetic Documents for Training Document Understanding Models
von: Dua, Karan, et al.
Veröffentlicht: (2025)
von: Dua, Karan, et al.
Veröffentlicht: (2025)
Transfer-learning for video classification: Video Swin Transformer on multiple domains
von: Oliveira, Daniel A. P., et al.
Veröffentlicht: (2022)
von: Oliveira, Daniel A. P., et al.
Veröffentlicht: (2022)
Prompt Sensitivity in Vision-Language Grounding: How Small Changes in Wording Affect Object Detection
von: Deka, Dawar Jyoti, et al.
Veröffentlicht: (2026)
von: Deka, Dawar Jyoti, et al.
Veröffentlicht: (2026)
Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy
von: Mehta, Vinit, et al.
Veröffentlicht: (2025)
von: Mehta, Vinit, et al.
Veröffentlicht: (2025)
A Vision-Language Model for Focal Liver Lesion Classification
von: Jian, Song, et al.
Veröffentlicht: (2025)
von: Jian, Song, et al.
Veröffentlicht: (2025)
From eye to AI: studying rodent social behavior in the era of machine Learning
von: Chindemi, Giuseppe, et al.
Veröffentlicht: (2025)
von: Chindemi, Giuseppe, et al.
Veröffentlicht: (2025)
Selection, Not Fusion: Radar-Modulated State Space Models for Radar-Camera Depth Estimation
von: Hou, Zhangcheng, et al.
Veröffentlicht: (2026)
von: Hou, Zhangcheng, et al.
Veröffentlicht: (2026)
MSPCaps: A Multi-Scale Patchify Capsule Network with Cross-Agreement Routing for Visual Recognition
von: Hu, Yudong, et al.
Veröffentlicht: (2025)
von: Hu, Yudong, et al.
Veröffentlicht: (2025)
Single-Shot Metric Depth from Focused Plenoptic Cameras
von: Lasheras-Hernandez, Blanca, et al.
Veröffentlicht: (2024)
von: Lasheras-Hernandez, Blanca, et al.
Veröffentlicht: (2024)
Smooth regularization for efficient video recognition
von: Goldman, Gil, et al.
Veröffentlicht: (2025)
von: Goldman, Gil, et al.
Veröffentlicht: (2025)
Evaluating the Impact of Synthetic Data on Object Detection Tasks in Autonomous Driving
von: Özeren, Enes, et al.
Veröffentlicht: (2025)
von: Özeren, Enes, et al.
Veröffentlicht: (2025)
Neuromorphic Monocular Depth Estimation with Uncertainty Modeling
von: Bergkvist, Viktor, et al.
Veröffentlicht: (2026)
von: Bergkvist, Viktor, et al.
Veröffentlicht: (2026)
ADAT: Time-Series-Aware Adaptive Transformer Architecture for Sign Language Translation
von: Shahin, Nada, et al.
Veröffentlicht: (2025)
von: Shahin, Nada, et al.
Veröffentlicht: (2025)
IMASHRIMP: Automatic White Shrimp (Penaeus vannamei) Biometrical Analysis from Laboratory Images Using Computer Vision and Deep Learning
von: González, Abiam Remache, et al.
Veröffentlicht: (2025)
von: González, Abiam Remache, et al.
Veröffentlicht: (2025)
Beyond RNNs: Benchmarking Attention-Based Image Captioning Models
von: Yanambakkam, Hemanth Teja, et al.
Veröffentlicht: (2025)
von: Yanambakkam, Hemanth Teja, et al.
Veröffentlicht: (2025)
Motion-Guided Semantic Alignment with Negative Prompts for Zero-Shot Video Action Recognition
von: Wang, Yiming, et al.
Veröffentlicht: (2026)
von: Wang, Yiming, et al.
Veröffentlicht: (2026)
OmniAcc: Personalized Accessibility Assistant Using Generative AI
von: Karki, Siddhant, et al.
Veröffentlicht: (2025)
von: Karki, Siddhant, et al.
Veröffentlicht: (2025)
Domain-Adaptive Pretraining Improves Primate Behavior Recognition
von: Mueller, Felix B., et al.
Veröffentlicht: (2025)
von: Mueller, Felix B., et al.
Veröffentlicht: (2025)
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation
von: Xiao, Jiasong, et al.
Veröffentlicht: (2026)
von: Xiao, Jiasong, et al.
Veröffentlicht: (2026)
Spatially-Aware Speaker for Vision-and-Language Navigation Instruction Generation
von: Gopinathan, Muraleekrishna, et al.
Veröffentlicht: (2024)
von: Gopinathan, Muraleekrishna, et al.
Veröffentlicht: (2024)
Unveiling the Potential of iMarkers: Invisible Fiducial Markers for Advanced Robotics
von: Tourani, Ali, et al.
Veröffentlicht: (2025)
von: Tourani, Ali, et al.
Veröffentlicht: (2025)
SuperPoint-SLAM3: Augmenting ORB-SLAM3 with Deep Features, Adaptive NMS, and Learning-Based Loop Closure
von: Syed, Shahram Najam, et al.
Veröffentlicht: (2025)
von: Syed, Shahram Najam, et al.
Veröffentlicht: (2025)
vS-Graphs: Tightly Coupling Visual SLAM and 3D Scene Graphs Exploiting Hierarchical Scene Understanding
von: Tourani, Ali, et al.
Veröffentlicht: (2025)
von: Tourani, Ali, et al.
Veröffentlicht: (2025)
Temporally Consistent Object 6D Pose Estimation for Robot Control
von: Zorina, Kateryna, et al.
Veröffentlicht: (2026)
von: Zorina, Kateryna, et al.
Veröffentlicht: (2026)
Caption-Driven Explainability: Probing CNNs for Bias via CLIP
von: Koller, Patrick, et al.
Veröffentlicht: (2025)
von: Koller, Patrick, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't
von: Nemitz, Jonathan, et al.
Veröffentlicht: (2026) -
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025) -
Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images
von: Chen, Yuangong, et al.
Veröffentlicht: (2026) -
WildRoadBench: A Wild Aerial Road-Damage Grounding Benchmark for Vision-Language Models and Autonomous Agents
von: Liu, Bingnan, et al.
Veröffentlicht: (2026) -
Leum-VL Technical Report
von: He, Yuxuan, et al.
Veröffentlicht: (2026)