Verb Mirage: Unveiling and Assessing Verb Concept Hallucinations in Multimodal Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Zehao, Liu, Xinpeng, Zhang, Yudonglin, Wu, Xiaoqian, Fang, Zhou, Fang, Yifan, Pu, Junfu, Lu, Cewu, Li, Yong-Lu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
From Isolated Islands to Pangea: Unifying Semantic Space for Human Action Understanding
di: Li, Yong-Lu, et al.
Pubblicazione: (2023)
di: Li, Yong-Lu, et al.
Pubblicazione: (2023)
Dynamic Relation Inference via Verb Embeddings
di: Suissa, Omri, et al.
Pubblicazione: (2025)
di: Suissa, Omri, et al.
Pubblicazione: (2025)
Mirage: Unveiling Hidden Artifacts in Synthetic Images with Large Vision-Language Models
di: Sharma, Pranav, et al.
Pubblicazione: (2025)
di: Sharma, Pranav, et al.
Pubblicazione: (2025)
Bridging the Gap between Human Motion and Action Semantics via Kinematic Phrases
di: Liu, Xinpeng, et al.
Pubblicazione: (2023)
di: Liu, Xinpeng, et al.
Pubblicazione: (2023)
Beyond Image-Text Matching: Verb Understanding in Multimodal Transformers Using Guided Masking
di: Beňová, Ivana, et al.
Pubblicazione: (2024)
di: Beňová, Ivana, et al.
Pubblicazione: (2024)
VerbDiff: Text-Only Diffusion Models with Enhanced Interaction Awareness
di: Cha, SeungJu, et al.
Pubblicazione: (2025)
di: Cha, SeungJu, et al.
Pubblicazione: (2025)
Revisit Human-Scene Interaction via Space Occupancy
di: Liu, Xinpeng, et al.
Pubblicazione: (2023)
di: Liu, Xinpeng, et al.
Pubblicazione: (2023)
Digital Gene: Learning about the Physical World through Analytic Concepts
di: Sun, Jianhua, et al.
Pubblicazione: (2025)
di: Sun, Jianhua, et al.
Pubblicazione: (2025)
Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language Models
di: He, Zhentao, et al.
Pubblicazione: (2025)
di: He, Zhentao, et al.
Pubblicazione: (2025)
Mitigating Hallucinations in Large Vision-Language Models by Self-Injecting Hallucinations
di: Lu, Yifan, et al.
Pubblicazione: (2025)
di: Lu, Yifan, et al.
Pubblicazione: (2025)
Mirage in the Eyes: Hallucination Attack on Multi-modal Large Language Models with Only Attention Sink
di: Wang, Yining, et al.
Pubblicazione: (2025)
di: Wang, Yining, et al.
Pubblicazione: (2025)
Low-Rank Similarity Mining for Multimodal Dataset Distillation
di: Xu, Yue, et al.
Pubblicazione: (2024)
di: Xu, Yue, et al.
Pubblicazione: (2024)
Homogeneous Dynamics Space for Heterogeneous Humans
di: Liu, Xinpeng, et al.
Pubblicazione: (2024)
di: Liu, Xinpeng, et al.
Pubblicazione: (2024)
ImDy: Human Inverse Dynamics from Imitated Observations
di: Liu, Xinpeng, et al.
Pubblicazione: (2024)
di: Liu, Xinpeng, et al.
Pubblicazione: (2024)
The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio
di: Leng, Sicong, et al.
Pubblicazione: (2024)
di: Leng, Sicong, et al.
Pubblicazione: (2024)
Physically Ground Commonsense Knowledge for Articulated Object Manipulation with Analytic Concepts
di: Wei, Jiude, et al.
Pubblicazione: (2025)
di: Wei, Jiude, et al.
Pubblicazione: (2025)
Hallucination of Multimodal Large Language Models: A Survey
di: Bai, Zechen, et al.
Pubblicazione: (2024)
di: Bai, Zechen, et al.
Pubblicazione: (2024)
DeformPAM: Data-Efficient Learning for Long-horizon Deformable Object Manipulation via Preference-based Action Alignment
di: Chen, Wendi, et al.
Pubblicazione: (2024)
di: Chen, Wendi, et al.
Pubblicazione: (2024)
Dynamic Multimodal Activation Steering for Hallucination Mitigation in Large Vision-Language Models
di: Yin, Jianghao, et al.
Pubblicazione: (2026)
di: Yin, Jianghao, et al.
Pubblicazione: (2026)
Grounded Chain-of-Thought for Multimodal Large Language Models
di: Wu, Qiong, et al.
Pubblicazione: (2025)
di: Wu, Qiong, et al.
Pubblicazione: (2025)
Towards Robust Evaluation of Visual Activity Recognition: Resolving Verb Ambiguity with Sense Clustering
di: Yao, Louie Hong, et al.
Pubblicazione: (2025)
di: Yao, Louie Hong, et al.
Pubblicazione: (2025)
LIDEA: Human-to-Robot Imitation Learning via Implicit Feature Distillation and Explicit Geometry Alignment
di: Xu, Yifu, et al.
Pubblicazione: (2026)
di: Xu, Yifu, et al.
Pubblicazione: (2026)
Hallucination Augmented Contrastive Learning for Multimodal Large Language Model
di: Jiang, Chaoya, et al.
Pubblicazione: (2023)
di: Jiang, Chaoya, et al.
Pubblicazione: (2023)
Graspness Discovery in Clutters for Fast and Accurate Grasp Detection
di: Wang, Chenxi, et al.
Pubblicazione: (2024)
di: Wang, Chenxi, et al.
Pubblicazione: (2024)
AnyDexGrasp: General Dexterous Grasping for Different Hands with Human-level Learning Efficiency
di: Fang, Hao-Shu, et al.
Pubblicazione: (2025)
di: Fang, Hao-Shu, et al.
Pubblicazione: (2025)
MIHBench: Benchmarking and Mitigating Multi-Image Hallucinations in Multimodal Large Language Models
di: Li, Jiale, et al.
Pubblicazione: (2025)
di: Li, Jiale, et al.
Pubblicazione: (2025)
Mitigating Hallucinations in Large Vision-Language Models without Performance Degradation
di: Zhu, Xingyu, et al.
Pubblicazione: (2026)
di: Zhu, Xingyu, et al.
Pubblicazione: (2026)
RSHallu: Dual-Mode Hallucination Evaluation for Remote-Sensing Multimodal Large Language Models with Domain-Tailored Mitigation
di: Zhou, Zihui, et al.
Pubblicazione: (2026)
di: Zhou, Zihui, et al.
Pubblicazione: (2026)
Extracting Visual Facts from Intermediate Layers for Mitigating Hallucinations in Multimodal Large Language Models
di: Zhou, Haoran, et al.
Pubblicazione: (2025)
di: Zhou, Haoran, et al.
Pubblicazione: (2025)
Reflective Instruction Tuning: Mitigating Hallucinations in Large Vision-Language Models
di: Zhang, Jinrui, et al.
Pubblicazione: (2024)
di: Zhang, Jinrui, et al.
Pubblicazione: (2024)
Dense Policy: Bidirectional Autoregressive Learning of Actions
di: Su, Yue, et al.
Pubblicazione: (2025)
di: Su, Yue, et al.
Pubblicazione: (2025)
Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models
di: Dong, Xinpeng, et al.
Pubblicazione: (2026)
di: Dong, Xinpeng, et al.
Pubblicazione: (2026)
Graph Integrated Multimodal Concept Bottleneck Model
di: Lin, Jiakai, et al.
Pubblicazione: (2025)
di: Lin, Jiakai, et al.
Pubblicazione: (2025)
Dancing with Still Images: Video Distillation via Static-Dynamic Disentanglement
di: Wang, Ziyu, et al.
Pubblicazione: (2023)
di: Wang, Ziyu, et al.
Pubblicazione: (2023)
Interacted Object Grounding in Spatio-Temporal Human-Object Interactions
di: Liu, Xiaoyang, et al.
Pubblicazione: (2024)
di: Liu, Xiaoyang, et al.
Pubblicazione: (2024)
DISCO: Embodied Navigation and Interaction via Differentiable Scene Semantics and Dual-level Control
di: Xu, Xinyu, et al.
Pubblicazione: (2024)
di: Xu, Xinyu, et al.
Pubblicazione: (2024)
Decoupling Perception from Reasoning for Hallucination-Resistant Video Understanding
di: Pu, Bowei, et al.
Pubblicazione: (2025)
di: Pu, Bowei, et al.
Pubblicazione: (2025)
SSL-OTA: Unveiling Backdoor Threats in Self-Supervised Learning for Object Detection
di: Wang, Qiannan, et al.
Pubblicazione: (2023)
di: Wang, Qiannan, et al.
Pubblicazione: (2023)
PACE: A Large-Scale Dataset with Pose Annotations in Cluttered Environments
di: You, Yang, et al.
Pubblicazione: (2023)
di: You, Yang, et al.
Pubblicazione: (2023)
NoiseBoost: Alleviating Hallucination with Noise Perturbation for Multimodal Large Language Models
di: Wu, Kai, et al.
Pubblicazione: (2024)
di: Wu, Kai, et al.
Pubblicazione: (2024)
Documenti analoghi
-
From Isolated Islands to Pangea: Unifying Semantic Space for Human Action Understanding
di: Li, Yong-Lu, et al.
Pubblicazione: (2023) -
Dynamic Relation Inference via Verb Embeddings
di: Suissa, Omri, et al.
Pubblicazione: (2025) -
Mirage: Unveiling Hidden Artifacts in Synthetic Images with Large Vision-Language Models
di: Sharma, Pranav, et al.
Pubblicazione: (2025) -
Bridging the Gap between Human Motion and Action Semantics via Kinematic Phrases
di: Liu, Xinpeng, et al.
Pubblicazione: (2023) -
Beyond Image-Text Matching: Verb Understanding in Multimodal Transformers Using Guided Masking
di: Beňová, Ivana, et al.
Pubblicazione: (2024)