CODE: Contrasting Self-generated Description to Combat Hallucination in Large Multi-modal Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kim, Junho, Kim, Hyunjun, Kim, Yeonju, Ro, Yong Man |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
What if...?: Thinking Counterfactual Keywords Helps to Mitigate Hallucination in Large Multi-modal Models
von: Kim, Junho, et al.
Veröffentlicht: (2024)
von: Kim, Junho, et al.
Veröffentlicht: (2024)
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis
von: Kim, Junho, et al.
Veröffentlicht: (2024)
von: Kim, Junho, et al.
Veröffentlicht: (2024)
ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding
von: Lee, Hosu, et al.
Veröffentlicht: (2025)
von: Lee, Hosu, et al.
Veröffentlicht: (2025)
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing
von: Lee, Hosu, et al.
Veröffentlicht: (2024)
von: Lee, Hosu, et al.
Veröffentlicht: (2024)
Empathetic Response in Audio-Visual Conversations Using Emotion Preference Optimization and MambaCompressor
von: Kim, Yeonju, et al.
Veröffentlicht: (2024)
von: Kim, Yeonju, et al.
Veröffentlicht: (2024)
Causal Unsupervised Semantic Segmentation
von: Kim, Junho, et al.
Veröffentlicht: (2023)
von: Kim, Junho, et al.
Veröffentlicht: (2023)
STRIDE: When to Speak Meets Sequence Denoising for Streaming Video Understanding
von: Kim, Junho, et al.
Veröffentlicht: (2026)
von: Kim, Junho, et al.
Veröffentlicht: (2026)
DIP-R1: Deep Inspection and Perception with RL Looking Through and Understanding Complex Scenes
von: Park, Sungjune, et al.
Veröffentlicht: (2025)
von: Park, Sungjune, et al.
Veröffentlicht: (2025)
AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues
von: Park, Se Jin, et al.
Veröffentlicht: (2024)
von: Park, Se Jin, et al.
Veröffentlicht: (2024)
Robust Pedestrian Detection via Constructing Versatile Pedestrian Knowledge Bank
von: Park, Sungjune, et al.
Veröffentlicht: (2024)
von: Park, Sungjune, et al.
Veröffentlicht: (2024)
Integrating Language-Derived Appearance Elements with Visual Cues in Pedestrian Detection
von: Park, Sungjune, et al.
Veröffentlicht: (2023)
von: Park, Sungjune, et al.
Veröffentlicht: (2023)
Language-guided Learning for Object Detection Tackling Multiple Variations in Aerial Images
von: Park, Sungjune, et al.
Veröffentlicht: (2025)
von: Park, Sungjune, et al.
Veröffentlicht: (2025)
Revealing Multi-View Hallucination in Large Vision-Language Models
von: Park, Wooje, et al.
Veröffentlicht: (2026)
von: Park, Wooje, et al.
Veröffentlicht: (2026)
Visual Hallucinations of Multi-modal Large Language Models
von: Huang, Wen, et al.
Veröffentlicht: (2024)
von: Huang, Wen, et al.
Veröffentlicht: (2024)
Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling
von: Park, Sungjune, et al.
Veröffentlicht: (2025)
von: Park, Sungjune, et al.
Veröffentlicht: (2025)
Pygmalion Effect in Vision: Image-to-Clay Translation for Reflective Geometry Reconstruction
von: Lee, Gayoung, et al.
Veröffentlicht: (2025)
von: Lee, Gayoung, et al.
Veröffentlicht: (2025)
Visual Delta Generator with Large Multi-modal Models for Semi-supervised Composed Image Retrieval
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024)
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024)
Prompt Tuning of Deep Neural Networks for Speaker-adaptive Visual Speech Recognition
von: Kim, Minsu, et al.
Veröffentlicht: (2023)
von: Kim, Minsu, et al.
Veröffentlicht: (2023)
Attention-space Contrastive Guidance for Efficient Hallucination Mitigation in LVLMs
von: Jo, Yujin, et al.
Veröffentlicht: (2026)
von: Jo, Yujin, et al.
Veröffentlicht: (2026)
Visual Speech Recognition for Languages with Limited Labeled Data using Automatic Labels from Whisper
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2023)
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2023)
Programmable-Room: Interactive Textured 3D Room Meshes Generation Empowered by Large Language Models
von: Kim, Jihyun, et al.
Veröffentlicht: (2025)
von: Kim, Jihyun, et al.
Veröffentlicht: (2025)
GCAgent: Long-Video Understanding via Schematic and Narrative Episodic Memory
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
von: Yeo, Jeong Hun, et al.
Veröffentlicht: (2025)
VG3T: Visual Geometry Grounded Gaussian Transformer
von: Kim, Junho, et al.
Veröffentlicht: (2025)
von: Kim, Junho, et al.
Veröffentlicht: (2025)
Simultaneous Long-tailed Recognition and Multi-modal Fusion for Highly Imbalanced Multi-modal Data
von: Yoon, Heegeon, et al.
Veröffentlicht: (2026)
von: Yoon, Heegeon, et al.
Veröffentlicht: (2026)
How Blind and Low-Vision Individuals Prefer Large Vision-Language Model-Generated Scene Descriptions
von: An, Na Min, et al.
Veröffentlicht: (2025)
von: An, Na Min, et al.
Veröffentlicht: (2025)
Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face Conversation
von: Park, Se Jin, et al.
Veröffentlicht: (2024)
von: Park, Se Jin, et al.
Veröffentlicht: (2024)
MGHanD: Multi-modal Guidance for authentic Hand Diffusion
von: Eum, Taehyeon, et al.
Veröffentlicht: (2025)
von: Eum, Taehyeon, et al.
Veröffentlicht: (2025)
Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues
von: Park, Beomchan, et al.
Veröffentlicht: (2026)
von: Park, Beomchan, et al.
Veröffentlicht: (2026)
PruneHal: Reducing Hallucinations in Multi-modal Large Language Models through Adaptive KV Cache Pruning
von: Sun, Fengyuan, et al.
Veröffentlicht: (2025)
von: Sun, Fengyuan, et al.
Veröffentlicht: (2025)
MSCoTDet: Language-driven Multi-modal Fusion for Improved Multispectral Pedestrian Detection
von: Kim, Taeheon, et al.
Veröffentlicht: (2024)
von: Kim, Taeheon, et al.
Veröffentlicht: (2024)
M4CXR: Exploring Multi-task Potentials of Multi-modal Large Language Models for Chest X-ray Interpretation
von: Park, Jonggwon, et al.
Veröffentlicht: (2024)
von: Park, Jonggwon, et al.
Veröffentlicht: (2024)
AV2AV: Direct Audio-Visual Speech to Audio-Visual Speech Translation with Unified Audio-Visual Speech Representation
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2023)
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2023)
DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language Models
von: Zhang, Yudong, et al.
Veröffentlicht: (2024)
von: Zhang, Yudong, et al.
Veröffentlicht: (2024)
Solution for SMART-101 Challenge of CVPR Multi-modal Algorithmic Reasoning Task 2024
von: Ahn, Jinwoo, et al.
Veröffentlicht: (2024)
von: Ahn, Jinwoo, et al.
Veröffentlicht: (2024)
DETACH : Decomposed Spatio-Temporal Alignment for Exocentric Video and Ambient Sensors with Staged Learning
von: Yoon, Junho, et al.
Veröffentlicht: (2025)
von: Yoon, Junho, et al.
Veröffentlicht: (2025)
Mitigating Hallucinations in Large Vision-Language Models via Summary-Guided Decoding
von: Min, Kyungmin, et al.
Veröffentlicht: (2024)
von: Min, Kyungmin, et al.
Veröffentlicht: (2024)
Revisiting Misalignment in Multispectral Pedestrian Detection: A Language-Driven Approach for Cross-modal Alignment Fusion
von: Kim, Taeheon, et al.
Veröffentlicht: (2024)
von: Kim, Taeheon, et al.
Veröffentlicht: (2024)
Preserving Pre-trained Representation Space: On Effectiveness of Prefix-tuning for Large Multi-modal Models
von: Kim, Donghoon, et al.
Veröffentlicht: (2024)
von: Kim, Donghoon, et al.
Veröffentlicht: (2024)
SDCD: Structure-Disrupted Contrastive Decoding for Mitigating Hallucinations in Large Vision-Language Models
von: Xia, Yuxuan, et al.
Veröffentlicht: (2026)
von: Xia, Yuxuan, et al.
Veröffentlicht: (2026)
Mitigating Hallucinations in Video Large Language Models via Spatiotemporal-Semantic Contrastive Decoding
von: Gao, Yuansheng, et al.
Veröffentlicht: (2026)
von: Gao, Yuansheng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
What if...?: Thinking Counterfactual Keywords Helps to Mitigate Hallucination in Large Multi-modal Models
von: Kim, Junho, et al.
Veröffentlicht: (2024) -
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis
von: Kim, Junho, et al.
Veröffentlicht: (2024) -
ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding
von: Lee, Hosu, et al.
Veröffentlicht: (2025) -
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing
von: Lee, Hosu, et al.
Veröffentlicht: (2024) -
Empathetic Response in Audio-Visual Conversations Using Emotion Preference Optimization and MambaCompressor
von: Kim, Yeonju, et al.
Veröffentlicht: (2024)