Verb Mirage: Unveiling and Assessing Verb Concept Hallucinations in Multimodal Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Zehao, Liu, Xinpeng, Zhang, Yudonglin, Wu, Xiaoqian, Fang, Zhou, Fang, Yifan, Pu, Junfu, Lu, Cewu, Li, Yong-Lu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Isolated Islands to Pangea: Unifying Semantic Space for Human Action Understanding
by: Li, Yong-Lu, et al.
Published: (2023)
by: Li, Yong-Lu, et al.
Published: (2023)
Dynamic Relation Inference via Verb Embeddings
by: Suissa, Omri, et al.
Published: (2025)
by: Suissa, Omri, et al.
Published: (2025)
Mirage: Unveiling Hidden Artifacts in Synthetic Images with Large Vision-Language Models
by: Sharma, Pranav, et al.
Published: (2025)
by: Sharma, Pranav, et al.
Published: (2025)
Bridging the Gap between Human Motion and Action Semantics via Kinematic Phrases
by: Liu, Xinpeng, et al.
Published: (2023)
by: Liu, Xinpeng, et al.
Published: (2023)
Beyond Image-Text Matching: Verb Understanding in Multimodal Transformers Using Guided Masking
by: Beňová, Ivana, et al.
Published: (2024)
by: Beňová, Ivana, et al.
Published: (2024)
VerbDiff: Text-Only Diffusion Models with Enhanced Interaction Awareness
by: Cha, SeungJu, et al.
Published: (2025)
by: Cha, SeungJu, et al.
Published: (2025)
Revisit Human-Scene Interaction via Space Occupancy
by: Liu, Xinpeng, et al.
Published: (2023)
by: Liu, Xinpeng, et al.
Published: (2023)
Digital Gene: Learning about the Physical World through Analytic Concepts
by: Sun, Jianhua, et al.
Published: (2025)
by: Sun, Jianhua, et al.
Published: (2025)
Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language Models
by: He, Zhentao, et al.
Published: (2025)
by: He, Zhentao, et al.
Published: (2025)
Mitigating Hallucinations in Large Vision-Language Models by Self-Injecting Hallucinations
by: Lu, Yifan, et al.
Published: (2025)
by: Lu, Yifan, et al.
Published: (2025)
Mirage in the Eyes: Hallucination Attack on Multi-modal Large Language Models with Only Attention Sink
by: Wang, Yining, et al.
Published: (2025)
by: Wang, Yining, et al.
Published: (2025)
Low-Rank Similarity Mining for Multimodal Dataset Distillation
by: Xu, Yue, et al.
Published: (2024)
by: Xu, Yue, et al.
Published: (2024)
Homogeneous Dynamics Space for Heterogeneous Humans
by: Liu, Xinpeng, et al.
Published: (2024)
by: Liu, Xinpeng, et al.
Published: (2024)
ImDy: Human Inverse Dynamics from Imitated Observations
by: Liu, Xinpeng, et al.
Published: (2024)
by: Liu, Xinpeng, et al.
Published: (2024)
The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio
by: Leng, Sicong, et al.
Published: (2024)
by: Leng, Sicong, et al.
Published: (2024)
Physically Ground Commonsense Knowledge for Articulated Object Manipulation with Analytic Concepts
by: Wei, Jiude, et al.
Published: (2025)
by: Wei, Jiude, et al.
Published: (2025)
Hallucination of Multimodal Large Language Models: A Survey
by: Bai, Zechen, et al.
Published: (2024)
by: Bai, Zechen, et al.
Published: (2024)
DeformPAM: Data-Efficient Learning for Long-horizon Deformable Object Manipulation via Preference-based Action Alignment
by: Chen, Wendi, et al.
Published: (2024)
by: Chen, Wendi, et al.
Published: (2024)
Dynamic Multimodal Activation Steering for Hallucination Mitigation in Large Vision-Language Models
by: Yin, Jianghao, et al.
Published: (2026)
by: Yin, Jianghao, et al.
Published: (2026)
Grounded Chain-of-Thought for Multimodal Large Language Models
by: Wu, Qiong, et al.
Published: (2025)
by: Wu, Qiong, et al.
Published: (2025)
Towards Robust Evaluation of Visual Activity Recognition: Resolving Verb Ambiguity with Sense Clustering
by: Yao, Louie Hong, et al.
Published: (2025)
by: Yao, Louie Hong, et al.
Published: (2025)
LIDEA: Human-to-Robot Imitation Learning via Implicit Feature Distillation and Explicit Geometry Alignment
by: Xu, Yifu, et al.
Published: (2026)
by: Xu, Yifu, et al.
Published: (2026)
Hallucination Augmented Contrastive Learning for Multimodal Large Language Model
by: Jiang, Chaoya, et al.
Published: (2023)
by: Jiang, Chaoya, et al.
Published: (2023)
Graspness Discovery in Clutters for Fast and Accurate Grasp Detection
by: Wang, Chenxi, et al.
Published: (2024)
by: Wang, Chenxi, et al.
Published: (2024)
AnyDexGrasp: General Dexterous Grasping for Different Hands with Human-level Learning Efficiency
by: Fang, Hao-Shu, et al.
Published: (2025)
by: Fang, Hao-Shu, et al.
Published: (2025)
MIHBench: Benchmarking and Mitigating Multi-Image Hallucinations in Multimodal Large Language Models
by: Li, Jiale, et al.
Published: (2025)
by: Li, Jiale, et al.
Published: (2025)
Mitigating Hallucinations in Large Vision-Language Models without Performance Degradation
by: Zhu, Xingyu, et al.
Published: (2026)
by: Zhu, Xingyu, et al.
Published: (2026)
RSHallu: Dual-Mode Hallucination Evaluation for Remote-Sensing Multimodal Large Language Models with Domain-Tailored Mitigation
by: Zhou, Zihui, et al.
Published: (2026)
by: Zhou, Zihui, et al.
Published: (2026)
Extracting Visual Facts from Intermediate Layers for Mitigating Hallucinations in Multimodal Large Language Models
by: Zhou, Haoran, et al.
Published: (2025)
by: Zhou, Haoran, et al.
Published: (2025)
Reflective Instruction Tuning: Mitigating Hallucinations in Large Vision-Language Models
by: Zhang, Jinrui, et al.
Published: (2024)
by: Zhang, Jinrui, et al.
Published: (2024)
Dense Policy: Bidirectional Autoregressive Learning of Actions
by: Su, Yue, et al.
Published: (2025)
by: Su, Yue, et al.
Published: (2025)
Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models
by: Dong, Xinpeng, et al.
Published: (2026)
by: Dong, Xinpeng, et al.
Published: (2026)
Graph Integrated Multimodal Concept Bottleneck Model
by: Lin, Jiakai, et al.
Published: (2025)
by: Lin, Jiakai, et al.
Published: (2025)
Dancing with Still Images: Video Distillation via Static-Dynamic Disentanglement
by: Wang, Ziyu, et al.
Published: (2023)
by: Wang, Ziyu, et al.
Published: (2023)
Interacted Object Grounding in Spatio-Temporal Human-Object Interactions
by: Liu, Xiaoyang, et al.
Published: (2024)
by: Liu, Xiaoyang, et al.
Published: (2024)
DISCO: Embodied Navigation and Interaction via Differentiable Scene Semantics and Dual-level Control
by: Xu, Xinyu, et al.
Published: (2024)
by: Xu, Xinyu, et al.
Published: (2024)
Decoupling Perception from Reasoning for Hallucination-Resistant Video Understanding
by: Pu, Bowei, et al.
Published: (2025)
by: Pu, Bowei, et al.
Published: (2025)
SSL-OTA: Unveiling Backdoor Threats in Self-Supervised Learning for Object Detection
by: Wang, Qiannan, et al.
Published: (2023)
by: Wang, Qiannan, et al.
Published: (2023)
PACE: A Large-Scale Dataset with Pose Annotations in Cluttered Environments
by: You, Yang, et al.
Published: (2023)
by: You, Yang, et al.
Published: (2023)
NoiseBoost: Alleviating Hallucination with Noise Perturbation for Multimodal Large Language Models
by: Wu, Kai, et al.
Published: (2024)
by: Wu, Kai, et al.
Published: (2024)
Similar Items
-
From Isolated Islands to Pangea: Unifying Semantic Space for Human Action Understanding
by: Li, Yong-Lu, et al.
Published: (2023) -
Dynamic Relation Inference via Verb Embeddings
by: Suissa, Omri, et al.
Published: (2025) -
Mirage: Unveiling Hidden Artifacts in Synthetic Images with Large Vision-Language Models
by: Sharma, Pranav, et al.
Published: (2025) -
Bridging the Gap between Human Motion and Action Semantics via Kinematic Phrases
by: Liu, Xinpeng, et al.
Published: (2023) -
Beyond Image-Text Matching: Verb Understanding in Multimodal Transformers Using Guided Masking
by: Beňová, Ivana, et al.
Published: (2024)