HKD4VLM: A Progressive Hybrid Knowledge Distillation Framework for Robust Multimodal Hallucination and Factuality Detection in VLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Zijian, Wu, Xuecheng, Huang, Danlei, Yan, Siyu, Peng, Chong, Cao, Xuezhi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Contrastive Knowledge Distillation for Robust Multimodal Sentiment Analysis
von: Sang, Zhongyi, et al.
Veröffentlicht: (2024)
von: Sang, Zhongyi, et al.
Veröffentlicht: (2024)
AD-AVSR: Asymmetric Dual-stream Enhancement for Robust Audio-Visual Speech Recognition
von: Xue, Junxiao, et al.
Veröffentlicht: (2025)
von: Xue, Junxiao, et al.
Veröffentlicht: (2025)
FakeSV-VLM: Taming VLM for Detecting Fake Short-Video News via Progressive Mixture-Of-Experts Adapter
von: Wang, Junxi, et al.
Veröffentlicht: (2025)
von: Wang, Junxi, et al.
Veröffentlicht: (2025)
Dark Side of Modalities: Reinforced Multimodal Distillation for Multimodal Knowledge Graph Reasoning
von: Zhao, Yu, et al.
Veröffentlicht: (2025)
von: Zhao, Yu, et al.
Veröffentlicht: (2025)
Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed Inputs
von: Ding, Peng, et al.
Veröffentlicht: (2024)
von: Ding, Peng, et al.
Veröffentlicht: (2024)
State-Anchored Complete-View Distillation for Robust Conversational Multimodal Emotion Recognition
von: Pan, Zhaoyan, et al.
Veröffentlicht: (2026)
von: Pan, Zhaoyan, et al.
Veröffentlicht: (2026)
EmoVLM-KD: Fusing Distilled Expertise with Vision-Language Models for Visual Emotion Analysis
von: Lee, SangEun, et al.
Veröffentlicht: (2025)
von: Lee, SangEun, et al.
Veröffentlicht: (2025)
MMC: Iterative Refinement of VLM Reasoning via MCTS-based Multimodal Critique
von: Liu, Shuhang, et al.
Veröffentlicht: (2025)
von: Liu, Shuhang, et al.
Veröffentlicht: (2025)
Hybrid CNN-Mamba Enhancement Network for Robust Multimodal Sentiment Analysis
von: Li, Xiang, et al.
Veröffentlicht: (2025)
von: Li, Xiang, et al.
Veröffentlicht: (2025)
ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs
von: Cui, Shiyao, et al.
Veröffentlicht: (2025)
von: Cui, Shiyao, et al.
Veröffentlicht: (2025)
Multi-source Knowledge Enhanced Graph Attention Networks for Multimodal Fact Verification
von: Cao, Han, et al.
Veröffentlicht: (2024)
von: Cao, Han, et al.
Veröffentlicht: (2024)
Multi-MLLM Knowledge Distillation for Out-of-Context News Detection
von: Gu, Yimeng, et al.
Veröffentlicht: (2025)
von: Gu, Yimeng, et al.
Veröffentlicht: (2025)
Efficient Speech Watermarking for Speech Synthesis via Progressive Knowledge Distillation
von: Cui, Yang, et al.
Veröffentlicht: (2025)
von: Cui, Yang, et al.
Veröffentlicht: (2025)
Augmenting Intra-Modal Understanding in MLLMs for Robust Multimodal Keyphrase Generation
von: Cao, Jiajun, et al.
Veröffentlicht: (2025)
von: Cao, Jiajun, et al.
Veröffentlicht: (2025)
Distilling Implicit Multimodal Knowledge into Large Language Models for Zero-Resource Dialogue Generation
von: Zhang, Bo, et al.
Veröffentlicht: (2024)
von: Zhang, Bo, et al.
Veröffentlicht: (2024)
Multimodal Fusion via Hypergraph Autoencoder and Contrastive Learning for Emotion Recognition in Conversation
von: Yi, Zijian, et al.
Veröffentlicht: (2024)
von: Yi, Zijian, et al.
Veröffentlicht: (2024)
FeatDistill: A Feature Distillation Enhanced Multi-Expert Ensemble Framework for Robust AI-generated Image Detection
von: Tu, Zhilin, et al.
Veröffentlicht: (2026)
von: Tu, Zhilin, et al.
Veröffentlicht: (2026)
Hyper-modal Imputation Diffusion Embedding with Dual-Distillation for Federated Multimodal Knowledge Graph Completion
von: Zhang, Ying, et al.
Veröffentlicht: (2025)
von: Zhang, Ying, et al.
Veröffentlicht: (2025)
Robust Multi-generation Learned Compression of Point Cloud Attribute
von: Liu, Xiangzuo, et al.
Veröffentlicht: (2025)
von: Liu, Xiangzuo, et al.
Veröffentlicht: (2025)
Towards Multimodal Emotional Support Conversation Systems
von: Chu, Yuqi, et al.
Veröffentlicht: (2024)
von: Chu, Yuqi, et al.
Veröffentlicht: (2024)
Zero-Effort Image-to-Music Generation: An Interpretable RAG-based VLM Approach
von: Zhao, Zijian, et al.
Veröffentlicht: (2025)
von: Zhao, Zijian, et al.
Veröffentlicht: (2025)
Multimodal Classification and Out-of-distribution Detection for Multimodal Intent Understanding
von: Zhang, Hanlei, et al.
Veröffentlicht: (2024)
von: Zhang, Hanlei, et al.
Veröffentlicht: (2024)
Towards Alleviating Text-to-Image Retrieval Hallucination for CLIP in Zero-shot Learning
von: Wang, Hanyao, et al.
Veröffentlicht: (2024)
von: Wang, Hanyao, et al.
Veröffentlicht: (2024)
OmniDPO: A Preference Optimization Framework to Address Omni-Modal Hallucination
von: Chen, Junzhe, et al.
Veröffentlicht: (2025)
von: Chen, Junzhe, et al.
Veröffentlicht: (2025)
Hallucination Localization in Video Captioning
von: Nakada, Shota, et al.
Veröffentlicht: (2025)
von: Nakada, Shota, et al.
Veröffentlicht: (2025)
A Progressive Evaluation Framework for Multicultural Analysis of Story Visualization
von: Kapuriya, Janak, et al.
Veröffentlicht: (2025)
von: Kapuriya, Janak, et al.
Veröffentlicht: (2025)
KEN: Knowledge Augmentation and Emotion Guidance Network for Multimodal Fake News Detection
von: Zhu, Peican, et al.
Veröffentlicht: (2025)
von: Zhu, Peican, et al.
Veröffentlicht: (2025)
TMDC: A Two-Stage Modality Denoising and Complementation Framework for Multimodal Sentiment Analysis with Missing and Noisy Modalities
von: Zhuang, Yan, et al.
Veröffentlicht: (2025)
von: Zhuang, Yan, et al.
Veröffentlicht: (2025)
Divide and Conquer: Multimodal Video Deepfake Detection via Cross-Modal Fusion and Localization
von: Li, Qingcao, et al.
Veröffentlicht: (2026)
von: Li, Qingcao, et al.
Veröffentlicht: (2026)
UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts
von: Cheng, Zhi-Qi, et al.
Veröffentlicht: (2024)
von: Cheng, Zhi-Qi, et al.
Veröffentlicht: (2024)
Exploring the Role of Audio in Multimodal Misinformation Detection
von: Liu, Moyang, et al.
Veröffentlicht: (2024)
von: Liu, Moyang, et al.
Veröffentlicht: (2024)
Multimodal LLM-based Query Paraphrasing for Video Search
von: Wu, Jiaxin, et al.
Veröffentlicht: (2024)
von: Wu, Jiaxin, et al.
Veröffentlicht: (2024)
DreamFoley: Scalable VLMs for High-Fidelity Video-to-Audio Generation
von: Li, Fu, et al.
Veröffentlicht: (2025)
von: Li, Fu, et al.
Veröffentlicht: (2025)
InconVAD: A Two-Stage Dual-Tower Framework for Multimodal Emotion Inconsistency Detection
von: Li, Zongyi, et al.
Veröffentlicht: (2025)
von: Li, Zongyi, et al.
Veröffentlicht: (2025)
Graph-based Interaction Augmentation Network for Robust Multimodal Sentiment Analysis
von: Zhangfeng, Hu, et al.
Veröffentlicht: (2025)
von: Zhangfeng, Hu, et al.
Veröffentlicht: (2025)
Mixture of Disentangled Experts with Missing Modalities for Robust Multimodal Sentiment Analysis
von: Li, Xiang, et al.
Veröffentlicht: (2026)
von: Li, Xiang, et al.
Veröffentlicht: (2026)
Unified Hallucination Detection for Multimodal Large Language Models
von: Chen, Xiang, et al.
Veröffentlicht: (2024)
von: Chen, Xiang, et al.
Veröffentlicht: (2024)
TRUST-VL: An Explainable News Assistant for General Multimodal Misinformation Detection
von: Yan, Zehong, et al.
Veröffentlicht: (2025)
von: Yan, Zehong, et al.
Veröffentlicht: (2025)
Multimodal Multi-turn Conversation Stance Detection: A Challenge Dataset and Effective Model
von: Niu, Fuqiang, et al.
Veröffentlicht: (2024)
von: Niu, Fuqiang, et al.
Veröffentlicht: (2024)
Contribution-Guided Asymmetric Learning for Robust Multimodal Fusion under Imbalance and Noise
von: Xu, Zijing, et al.
Veröffentlicht: (2025)
von: Xu, Zijing, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Contrastive Knowledge Distillation for Robust Multimodal Sentiment Analysis
von: Sang, Zhongyi, et al.
Veröffentlicht: (2024) -
AD-AVSR: Asymmetric Dual-stream Enhancement for Robust Audio-Visual Speech Recognition
von: Xue, Junxiao, et al.
Veröffentlicht: (2025) -
FakeSV-VLM: Taming VLM for Detecting Fake Short-Video News via Progressive Mixture-Of-Experts Adapter
von: Wang, Junxi, et al.
Veröffentlicht: (2025) -
Dark Side of Modalities: Reinforced Multimodal Distillation for Multimodal Knowledge Graph Reasoning
von: Zhao, Yu, et al.
Veröffentlicht: (2025) -
Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed Inputs
von: Ding, Peng, et al.
Veröffentlicht: (2024)