Understanding Multimodal Hallucination with Parameter-Free Representation Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yueqian, Liang, Jianxin, Wang, Yuxuan, Zhang, Huishuai, Zhao, Dongyan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FREAK: A Fine-grained Hallucination Evaluation Benchmark for Advanced MLLMs
by: Yin, Zhihan, et al.
Published: (2026)
by: Yin, Zhihan, et al.
Published: (2026)
VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interaction Format
by: Wang, Yueqian, et al.
Published: (2024)
by: Wang, Yueqian, et al.
Published: (2024)
Beyond Isolated Facts: Synthesizing Narrative and Grounded Supervision for VideoQA
by: Liang, Jianxin, et al.
Published: (2025)
by: Liang, Jianxin, et al.
Published: (2025)
ReasVQA: Advancing VideoQA with Imperfect Reasoning Process
by: Liang, Jianxin, et al.
Published: (2025)
by: Liang, Jianxin, et al.
Published: (2025)
ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models
by: Wang, Yueqian, et al.
Published: (2025)
by: Wang, Yueqian, et al.
Published: (2025)
VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models
by: Wang, Yuxuan, et al.
Published: (2024)
by: Wang, Yuxuan, et al.
Published: (2024)
Efficient Temporal Extrapolation of Multimodal Large Language Models with Temporal Grounding Bridge
by: Wang, Yuxuan, et al.
Published: (2024)
by: Wang, Yuxuan, et al.
Published: (2024)
HawkEye: Training Video-Text LLMs for Grounding Text in Videos
by: Wang, Yueqian, et al.
Published: (2024)
by: Wang, Yueqian, et al.
Published: (2024)
STAIR: Spatial-Temporal Reasoning with Auditable Intermediate Results for Video Question Answering
by: Wang, Yueqian, et al.
Published: (2024)
by: Wang, Yueqian, et al.
Published: (2024)
MMDuet2: Enhancing Proactive Interaction of Video MLLMs with Multi-Turn Reinforcement Learning
by: Wang, Yueqian, et al.
Published: (2025)
by: Wang, Yueqian, et al.
Published: (2025)
End-to-End Video Question Answering with Frame Scoring Mechanisms and Adaptive Sampling
by: Liang, Jianxin, et al.
Published: (2024)
by: Liang, Jianxin, et al.
Published: (2024)
OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts
by: Wang, Yuxuan, et al.
Published: (2025)
by: Wang, Yuxuan, et al.
Published: (2025)
Progressive Local Alignment for Medical Multimodal Pre-training
by: Yan, Huimin, et al.
Published: (2025)
by: Yan, Huimin, et al.
Published: (2025)
Understanding the Role of Hallucination in Reinforcement Post-Training of Multimodal Reasoning Models
by: Zhang, Gengwei, et al.
Published: (2026)
by: Zhang, Gengwei, et al.
Published: (2026)
MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM
by: Dong, Bowen, et al.
Published: (2025)
by: Dong, Bowen, et al.
Published: (2025)
Multimodal Adaptive Retrieval Augmented Generation through Internal Representation Learning
by: Du, Ruoshuang, et al.
Published: (2026)
by: Du, Ruoshuang, et al.
Published: (2026)
Gramian Multimodal Representation Learning and Alignment
by: Cicchetti, Giordano, et al.
Published: (2024)
by: Cicchetti, Giordano, et al.
Published: (2024)
Towards Robust Multimodal Representation: A Unified Approach with Adaptive Experts and Alignment
by: Moradinasab, Nazanin, et al.
Published: (2025)
by: Moradinasab, Nazanin, et al.
Published: (2025)
Multiple Object Stitching for Unsupervised Representation Learning
by: Shen, Chengchao, et al.
Published: (2025)
by: Shen, Chengchao, et al.
Published: (2025)
MoviePuzzle: Visual Narrative Reasoning through Multimodal Order Learning
by: Wang, Jianghui, et al.
Published: (2023)
by: Wang, Jianghui, et al.
Published: (2023)
Learning to Think in Physics: Breaking Shortcut Learning in Scientific Diffusion via Representation Alignment
by: Jia, Haozhe, et al.
Published: (2026)
by: Jia, Haozhe, et al.
Published: (2026)
Suppressing VLM Hallucinations with Spectral Representation Filtering
by: Ali, Ameen, et al.
Published: (2025)
by: Ali, Ameen, et al.
Published: (2025)
Mitigating Object and Action Hallucinations in Multimodal LLMs via Self-Augmented Contrastive Alignment
by: Chang, Kai-Po, et al.
Published: (2025)
by: Chang, Kai-Po, et al.
Published: (2025)
Query-Conditioned Evidential Keyframe Sampling for MLLM-Based Long-Form Video Understanding
by: Wang, Yiheng, et al.
Published: (2026)
by: Wang, Yiheng, et al.
Published: (2026)
Semantic Residual for Multimodal Unified Discrete Representation
by: Huang, Hai, et al.
Published: (2024)
by: Huang, Hai, et al.
Published: (2024)
Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations
by: Jiang, Nick, et al.
Published: (2024)
by: Jiang, Nick, et al.
Published: (2024)
LongViTU: Instruction Tuning for Long-Form Video Understanding
by: Wu, Rujie, et al.
Published: (2025)
by: Wu, Rujie, et al.
Published: (2025)
Causal Decoding for Hallucination-Resistant Multimodal Large Language Models
by: Tan, Shiwei, et al.
Published: (2026)
by: Tan, Shiwei, et al.
Published: (2026)
GOMA: Toward Structure-Driven Multimodal Alignment from a Graph Signal Smoothing Perspective
by: Wang, Xu, et al.
Published: (2026)
by: Wang, Xu, et al.
Published: (2026)
Prompt Disentanglement via Language Guidance and Representation Alignment for Domain Generalization
by: Cheng, De, et al.
Published: (2025)
by: Cheng, De, et al.
Published: (2025)
Understanding Representation Dynamics of Diffusion Models via Low-Dimensional Modeling
by: Li, Xiao, et al.
Published: (2025)
by: Li, Xiao, et al.
Published: (2025)
PolyhedronNet: Representation Learning for Polyhedra with Surface-attributed Graph
by: Yu, Dazhou, et al.
Published: (2025)
by: Yu, Dazhou, et al.
Published: (2025)
FORESEE: Multimodal and Multi-view Representation Learning for Robust Prediction of Cancer Survival
by: Pan, Liangrui, et al.
Published: (2024)
by: Pan, Liangrui, et al.
Published: (2024)
What's in Common? Multimodal Models Hallucinate When Reasoning Across Scenes
by: Ross, Candace, et al.
Published: (2025)
by: Ross, Candace, et al.
Published: (2025)
From Parameter to Representation: A Closed-Form Approach for Controllable Model Merging
by: Wu, Jialin, et al.
Published: (2025)
by: Wu, Jialin, et al.
Published: (2025)
Semantic Editing with Coupled Stochastic Differential Equations
by: Zhang, Jianxin, et al.
Published: (2025)
by: Zhang, Jianxin, et al.
Published: (2025)
OccamVTS: Distilling Vision Models to 1% Parameters for Time Series Forecasting
by: Lyu, Sisuo, et al.
Published: (2025)
by: Lyu, Sisuo, et al.
Published: (2025)
VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding
by: Li, Zongxia, et al.
Published: (2025)
by: Li, Zongxia, et al.
Published: (2025)
Woodpecker: Hallucination Correction for Multimodal Large Language Models
by: Yin, Shukang, et al.
Published: (2023)
by: Yin, Shukang, et al.
Published: (2023)
Graph-based Unsupervised Disentangled Representation Learning via Multimodal Large Language Models
by: Xie, Baao, et al.
Published: (2024)
by: Xie, Baao, et al.
Published: (2024)
Similar Items
-
FREAK: A Fine-grained Hallucination Evaluation Benchmark for Advanced MLLMs
by: Yin, Zhihan, et al.
Published: (2026) -
VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interaction Format
by: Wang, Yueqian, et al.
Published: (2024) -
Beyond Isolated Facts: Synthesizing Narrative and Grounded Supervision for VideoQA
by: Liang, Jianxin, et al.
Published: (2025) -
ReasVQA: Advancing VideoQA with Imperfect Reasoning Process
by: Liang, Jianxin, et al.
Published: (2025) -
ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models
by: Wang, Yueqian, et al.
Published: (2025)