HallE-Control: Controlling Object Hallucination in Large Multimodal Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zhai, Bohan, Yang, Shijia, Xu, Chenfeng, Shen, Sheng, Keutzer, Kurt, Li, Chunyuan, Li, Manling |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CaptionQA: Is Your Caption as Useful as the Image Itself?
by: Yang, Shijia, et al.
Published: (2025)
by: Yang, Shijia, et al.
Published: (2025)
Law of Vision Representation in MLLMs
by: Yang, Shijia, et al.
Published: (2024)
by: Yang, Shijia, et al.
Published: (2024)
Immiscible Diffusion: Accelerating Diffusion Training with Noise Assignment
by: Li, Yiheng, et al.
Published: (2024)
by: Li, Yiheng, et al.
Published: (2024)
Improved Immiscible Diffusion: Accelerate Diffusion Training by Reducing Its Miscibility
by: Li, Yiheng, et al.
Published: (2025)
by: Li, Yiheng, et al.
Published: (2025)
A Lesson in Splats: Teacher-Guided Diffusion for 3D Gaussian Splats Generation with 2D Supervision
by: Peng, Chensheng, et al.
Published: (2024)
by: Peng, Chensheng, et al.
Published: (2024)
Looking Backward: Streaming Video-to-Video Translation with Feature Banks
by: Liang, Feng, et al.
Published: (2024)
by: Liang, Feng, et al.
Published: (2024)
Segment Any Motion in Videos
by: Huang, Nan, et al.
Published: (2025)
by: Huang, Nan, et al.
Published: (2025)
Sparse Refinement for Efficient High-Resolution Semantic Segmentation
by: Liu, Zhijian, et al.
Published: (2024)
by: Liu, Zhijian, et al.
Published: (2024)
ReactBench: A Cause-Driven Benchmark for Multimodal Hallucination via Systematic Evaluation
by: Zhou, Shizhe, et al.
Published: (2026)
by: Zhou, Shizhe, et al.
Published: (2026)
Q-SLAM: Quadric Representations for Monocular SLAM
by: Peng, Chensheng, et al.
Published: (2024)
by: Peng, Chensheng, et al.
Published: (2024)
UniDrive: Towards Universal Driving Perception Across Camera Configurations
by: Li, Ye, et al.
Published: (2024)
by: Li, Ye, et al.
Published: (2024)
DeSiRe-GS: 4D Street Gaussians for Static-Dynamic Decomposition and Surface Reconstruction for Urban Driving Scenes
by: Peng, Chensheng, et al.
Published: (2024)
by: Peng, Chensheng, et al.
Published: (2024)
Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Attention Lens
by: Jiang, Zhangqi, et al.
Published: (2024)
by: Jiang, Zhangqi, et al.
Published: (2024)
Nullu: Mitigating Object Hallucinations in Large Vision-Language Models via HalluSpace Projection
by: Yang, Le, et al.
Published: (2024)
by: Yang, Le, et al.
Published: (2024)
LDGNet: A Lightweight Difference Guiding Network for Remote Sensing Change Detection
by: Xu, Chenfeng
Published: (2025)
by: Xu, Chenfeng
Published: (2025)
Graphic Design with Large Multimodal Model
by: Cheng, Yutao, et al.
Published: (2024)
by: Cheng, Yutao, et al.
Published: (2024)
GeoDrive: 3D Geometry-Informed Driving World Model with Precise Action Control
by: Chen, Anthony, et al.
Published: (2025)
by: Chen, Anthony, et al.
Published: (2025)
Exploring Hallucination of Large Multimodal Models in Video Understanding: Benchmark, Analysis and Mitigation
by: Gao, Hongcheng, et al.
Published: (2025)
by: Gao, Hongcheng, et al.
Published: (2025)
Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware Permutation
by: Yang, Shuo, et al.
Published: (2025)
by: Yang, Shuo, et al.
Published: (2025)
Hallucination Augmented Contrastive Learning for Multimodal Large Language Model
by: Jiang, Chaoya, et al.
Published: (2023)
by: Jiang, Chaoya, et al.
Published: (2023)
RealCamo: Boosting Real Camouflage Synthesis with Layout Controls and Textual-Visual Guidance
by: Chen, Chunyuan, et al.
Published: (2025)
by: Chen, Chunyuan, et al.
Published: (2025)
Beyond Raw Videos: Understanding Edited Videos with Large Multimodal Model
by: Xu, Lu, et al.
Published: (2024)
by: Xu, Lu, et al.
Published: (2024)
Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language Models
by: He, Zhentao, et al.
Published: (2025)
by: He, Zhentao, et al.
Published: (2025)
Reminding Multimodal Large Language Models of Object-aware Knowledge with Retrieved Tags
by: Qi, Daiqing, et al.
Published: (2024)
by: Qi, Daiqing, et al.
Published: (2024)
VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding
by: Li, Chaoyu, et al.
Published: (2024)
by: Li, Chaoyu, et al.
Published: (2024)
LISA: A Layer-wise Integration and Suppression Approach for Hallucination Mitigation in Multimodal Large Language Models
by: Guo, Zhihui, et al.
Published: (2025)
by: Guo, Zhihui, et al.
Published: (2025)
Hallucination of Multimodal Large Language Models: A Survey
by: Bai, Zechen, et al.
Published: (2024)
by: Bai, Zechen, et al.
Published: (2024)
IPDreamer: Appearance-Controllable 3D Object Generation with Complex Image Prompts
by: Zeng, Bohan, et al.
Published: (2023)
by: Zeng, Bohan, et al.
Published: (2023)
OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
by: Liu, Yuliang, et al.
Published: (2023)
by: Liu, Yuliang, et al.
Published: (2023)
MIHBench: Benchmarking and Mitigating Multi-Image Hallucinations in Multimodal Large Language Models
by: Li, Jiale, et al.
Published: (2025)
by: Li, Jiale, et al.
Published: (2025)
CAI: Caption-Sensitive Attention Intervention for Mitigating Object Hallucination in Large Vision-Language Models
by: Li, Qiming, et al.
Published: (2025)
by: Li, Qiming, et al.
Published: (2025)
DrivingRecon: Large 4D Gaussian Reconstruction Model For Autonomous Driving
by: Lu, Hao, et al.
Published: (2024)
by: Lu, Hao, et al.
Published: (2024)
Woodpecker: Hallucination Correction for Multimodal Large Language Models
by: Yin, Shukang, et al.
Published: (2023)
by: Yin, Shukang, et al.
Published: (2023)
The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio
by: Leng, Sicong, et al.
Published: (2024)
by: Leng, Sicong, et al.
Published: (2024)
Score-Control for Hallucination Reduction in Diffusion Models
by: Bhosale, Mahesh, et al.
Published: (2026)
by: Bhosale, Mahesh, et al.
Published: (2026)
Causal Tracing of Object Representations in Large Vision Language Models: Mechanistic Interpretability and Hallucination Mitigation
by: Li, Qiming, et al.
Published: (2025)
by: Li, Qiming, et al.
Published: (2025)
TOUCH: Text-guided Controllable Generation of Free-Form Hand-Object Interactions
by: Han, Guangyi, et al.
Published: (2025)
by: Han, Guangyi, et al.
Published: (2025)
ClearSight: Visual Signal Enhancement for Object Hallucination Mitigation in Multimodal Large language Models
by: Yin, Hao, et al.
Published: (2025)
by: Yin, Hao, et al.
Published: (2025)
Learning to Decode Against Compositional Hallucination in Video Multimodal Large Language Models
by: Xing, Wenbin, et al.
Published: (2026)
by: Xing, Wenbin, et al.
Published: (2026)
Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models
by: Xu, Yifang, et al.
Published: (2025)
by: Xu, Yifang, et al.
Published: (2025)
Similar Items
-
CaptionQA: Is Your Caption as Useful as the Image Itself?
by: Yang, Shijia, et al.
Published: (2025) -
Law of Vision Representation in MLLMs
by: Yang, Shijia, et al.
Published: (2024) -
Immiscible Diffusion: Accelerating Diffusion Training with Noise Assignment
by: Li, Yiheng, et al.
Published: (2024) -
Improved Immiscible Diffusion: Accelerate Diffusion Training by Reducing Its Miscibility
by: Li, Yiheng, et al.
Published: (2025) -
A Lesson in Splats: Teacher-Guided Diffusion for 3D Gaussian Splats Generation with 2D Supervision
by: Peng, Chensheng, et al.
Published: (2024)