Combating Multimodal LLM Hallucination via Bottom-Up Holistic Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Shengqiong, Fei, Hao, Pan, Liangming, Wang, William Yang, Yan, Shuicheng, Chua, Tat-Seng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Semantic Equivalence of Tokenization in Multimodal LLM
by: Wu, Shengqiong, et al.
Published: (2024)
by: Wu, Shengqiong, et al.
Published: (2024)
Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing
by: Fei, Hao, et al.
Published: (2024)
by: Fei, Hao, et al.
Published: (2024)
Universal Scene Graph Generation
by: Wu, Shengqiong, et al.
Published: (2025)
by: Wu, Shengqiong, et al.
Published: (2025)
Enhancing Video-Language Representations with Structural Spatio-Temporal Alignment
by: Fei, Hao, et al.
Published: (2024)
by: Fei, Hao, et al.
Published: (2024)
Global Commander and Local Operative: A Dual-Agent Framework for Scene Navigation
by: Jin, Kaiming, et al.
Published: (2026)
by: Jin, Kaiming, et al.
Published: (2026)
Modeling Cross-vision Synergy for Unified Large Vision Model
by: Wu, Shengqiong, et al.
Published: (2026)
by: Wu, Shengqiong, et al.
Published: (2026)
Dysen-VDM: Empowering Dynamics-aware Text-to-Video Diffusion with LLMs
by: Fei, Hao, et al.
Published: (2023)
by: Fei, Hao, et al.
Published: (2023)
Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey
by: Wang, Yaoting, et al.
Published: (2025)
by: Wang, Yaoting, et al.
Published: (2025)
Auto-Encoding Morph-Tokens for Multimodal LLM
by: Pan, Kaihang, et al.
Published: (2024)
by: Pan, Kaihang, et al.
Published: (2024)
Synergizing Understanding and Generation with Interleaved Analyzing-Drafting Thinking
by: Wu, Shengqiong, et al.
Published: (2026)
by: Wu, Shengqiong, et al.
Published: (2026)
A Reason-then-Describe Instruction Interpreter for Controllable Video Generation
by: Wu, Shengqiong, et al.
Published: (2025)
by: Wu, Shengqiong, et al.
Published: (2025)
Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation
by: Wu, Shengqiong, et al.
Published: (2025)
by: Wu, Shengqiong, et al.
Published: (2025)
JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation
by: Liu, Kai, et al.
Published: (2025)
by: Liu, Kai, et al.
Published: (2025)
Learning 4D Panoptic Scene Graph Generation from Rich 2D Visual Scene
by: Wu, Shengqiong, et al.
Published: (2025)
by: Wu, Shengqiong, et al.
Published: (2025)
Understanding Long Videos via LLM-Powered Entity Relation Graphs
by: Chu, Meng, et al.
Published: (2025)
by: Chu, Meng, et al.
Published: (2025)
OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
by: Zhang, Tao, et al.
Published: (2024)
by: Zhang, Tao, et al.
Published: (2024)
Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning
by: Qian, Long, et al.
Published: (2024)
by: Qian, Long, et al.
Published: (2024)
Compose Your Aesthetics: Empowering Text-to-Image Models with the Principles of Art
by: Jin, Zhe, et al.
Published: (2025)
by: Jin, Zhe, et al.
Published: (2025)
JavisDiT++: Unified Modeling and Optimization for Joint Audio-Video Generation
by: Liu, Kai, et al.
Published: (2026)
by: Liu, Kai, et al.
Published: (2026)
MuSLR: Multimodal Symbolic Logical Reasoning
by: Xu, Jundong, et al.
Published: (2025)
by: Xu, Jundong, et al.
Published: (2025)
Thinking with Blueprints: Assisting Vision-Language Models in Spatial Reasoning via Structured Object Representation
by: Ma, Weijian, et al.
Published: (2026)
by: Ma, Weijian, et al.
Published: (2026)
UniFGVC: Universal Training-Free Few-Shot Fine-Grained Vision Classification via Attribute-Aware Multimodal Retrieval
by: Guo, Hongyu, et al.
Published: (2025)
by: Guo, Hongyu, et al.
Published: (2025)
Benchmarking Multimodal CoT Reward Model Stepwise by Visual Program
by: Gao, Minghe, et al.
Published: (2025)
by: Gao, Minghe, et al.
Published: (2025)
Principled Multimodal Representation Learning
by: Liu, Xiaohao, et al.
Published: (2025)
by: Liu, Xiaohao, et al.
Published: (2025)
3D Magic Mirror: Clothing Reconstruction from a Single Image via a Causal Perspective
by: Zheng, Zhedong, et al.
Published: (2022)
by: Zheng, Zhedong, et al.
Published: (2022)
Composed Image Retrieval with Text Feedback via Multi-grained Uncertainty Regularization
by: Chen, Yiyang, et al.
Published: (2022)
by: Chen, Yiyang, et al.
Published: (2022)
Precise Shield: Explaining and Aligning VLLM Safety via Neuron-Level Guidance
by: Shi, Enyi, et al.
Published: (2026)
by: Shi, Enyi, et al.
Published: (2026)
Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model
by: Shen, Fei, et al.
Published: (2025)
by: Shen, Fei, et al.
Published: (2025)
Dynamic Multimodal Fusion via Meta-Learning Towards Micro-Video Recommendation
by: Liu, Han, et al.
Published: (2025)
by: Liu, Han, et al.
Published: (2025)
Zero-1-to-A: Zero-Shot One Image to Animatable Head Avatars Using Video Diffusion
by: Zhou, Zhenglin, et al.
Published: (2025)
by: Zhou, Zhenglin, et al.
Published: (2025)
Reasoning-VLA: A Fast and General Vision-Language-Action Reasoning Model for Autonomous Driving
by: Zhang, Dapeng, et al.
Published: (2025)
by: Zhang, Dapeng, et al.
Published: (2025)
Multiple-environment Self-adaptive Network for Aerial-view Geo-localization
by: Wang, Tingyu, et al.
Published: (2022)
by: Wang, Tingyu, et al.
Published: (2022)
Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization
by: Yang, Wenhao, et al.
Published: (2026)
by: Yang, Wenhao, et al.
Published: (2026)
Boosting Reasoning in Large Multimodal Models via Activation Replay
by: Xing, Yun, et al.
Published: (2025)
by: Xing, Yun, et al.
Published: (2025)
ExpLLM: Towards Chain of Thought for Facial Expression Recognition
by: Lan, Xing, et al.
Published: (2024)
by: Lan, Xing, et al.
Published: (2024)
On Path to Multimodal Generalist: General-Level and General-Bench
by: Fei, Hao, et al.
Published: (2025)
by: Fei, Hao, et al.
Published: (2025)
Disentangling Masked Autoencoders for Unsupervised Domain Generalization
by: Zhang, An, et al.
Published: (2024)
by: Zhang, An, et al.
Published: (2024)
Turing Patterns for Multimedia: Reaction-Diffusion Multi-Modal Fusion for Language-Guided Video Moment Retrieval
by: Fang, Xiang, et al.
Published: (2026)
by: Fang, Xiang, et al.
Published: (2026)
Extending Visual Dynamics for Video-to-Music Generation
by: Liu, Xiaohao, et al.
Published: (2025)
by: Liu, Xiaohao, et al.
Published: (2025)
Reinforcing Video Reasoning with Focused Thinking
by: Dang, Jisheng, et al.
Published: (2025)
by: Dang, Jisheng, et al.
Published: (2025)
Similar Items
-
Towards Semantic Equivalence of Tokenization in Multimodal LLM
by: Wu, Shengqiong, et al.
Published: (2024) -
Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing
by: Fei, Hao, et al.
Published: (2024) -
Universal Scene Graph Generation
by: Wu, Shengqiong, et al.
Published: (2025) -
Enhancing Video-Language Representations with Structural Spatio-Temporal Alignment
by: Fei, Hao, et al.
Published: (2024) -
Global Commander and Local Operative: A Dual-Agent Framework for Scene Navigation
by: Jin, Kaiming, et al.
Published: (2026)