IRIS: Implicit Reward-Guided Internal Sifting for Mitigating Multimodal Hallucination
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Yuanshuai, Yan, Yuping, Han, Jirui, Ming, Fei, Lv, Lingjuan, Jin, Yaochu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
HMACE: Heterogeneous Multi-Agent Collaborative Evolution for Combinatorial Optimization
di: Yan, Yuping, et al.
Pubblicazione: (2026)
di: Yan, Yuping, et al.
Pubblicazione: (2026)
TriCon-SF: A Triple-Shuffle and Contribution-Aware Serial Federated Learning Framework for Heterogeneous Healthcare Data
di: Yan, Yuping, et al.
Pubblicazione: (2025)
di: Yan, Yuping, et al.
Pubblicazione: (2025)
Mitigating Visual Hallucinations via Semantic Curriculum Preference Optimization in MLLMs
di: Li, Yuanshuai, et al.
Pubblicazione: (2025)
di: Li, Yuanshuai, et al.
Pubblicazione: (2025)
OutSafe-Bench: A Benchmark for Multimodal Offensive Content Detection in Large Language Models
di: Yan, Yuping, et al.
Pubblicazione: (2025)
di: Yan, Yuping, et al.
Pubblicazione: (2025)
Adversarial Reward Auditing for Active Detection and Mitigation of Reward Hacking
di: Beigi, Mohammad, et al.
Pubblicazione: (2026)
di: Beigi, Mohammad, et al.
Pubblicazione: (2026)
LacaDM: A Latent Causal Diffusion Model for Multiobjective Reinforcement Learning
di: Yan, Xueming, et al.
Pubblicazione: (2025)
di: Yan, Xueming, et al.
Pubblicazione: (2025)
When Foundation Model Meets Federated Learning: Motivations, Challenges, and Future Directions
di: Zhuang, Weiming, et al.
Pubblicazione: (2023)
di: Zhuang, Weiming, et al.
Pubblicazione: (2023)
On the Implicit Reward Overfitting and the Low-rank Dynamics in RLVR
di: Ye, Hao, et al.
Pubblicazione: (2026)
di: Ye, Hao, et al.
Pubblicazione: (2026)
Model Internals-based Answer Attribution for Trustworthy Retrieval-Augmented Generation
di: Qi, Jirui, et al.
Pubblicazione: (2024)
di: Qi, Jirui, et al.
Pubblicazione: (2024)
When Alignment Fails: Multimodal Adversarial Attacks on Vision-Language-Action Models
di: Yan, Yuping, et al.
Pubblicazione: (2025)
di: Yan, Yuping, et al.
Pubblicazione: (2025)
IR$^3$: Contrastive Inverse Reinforcement Learning for Interpretable Detection and Mitigation of Reward Hacking
di: Beigi, Mohammad, et al.
Pubblicazione: (2026)
di: Beigi, Mohammad, et al.
Pubblicazione: (2026)
Entropy-Guided Data-Efficient Training for Multimodal Reasoning Reward Models
di: Yang, Shidong, et al.
Pubblicazione: (2026)
di: Yang, Shidong, et al.
Pubblicazione: (2026)
Process Reinforcement through Implicit Rewards
di: Cui, Ganqu, et al.
Pubblicazione: (2025)
di: Cui, Ganqu, et al.
Pubblicazione: (2025)
Reward Shaping to Mitigate Reward Hacking in RLHF
di: Fu, Jiayi, et al.
Pubblicazione: (2025)
di: Fu, Jiayi, et al.
Pubblicazione: (2025)
Reward Hacking Mitigation using Verifiable Composite Rewards
di: Tarek, Mirza Farhan Bin, et al.
Pubblicazione: (2025)
di: Tarek, Mirza Farhan Bin, et al.
Pubblicazione: (2025)
Repairing Reward Functions with Feedback to Mitigate Reward Hacking
di: Hatgis-Kessell, Stephane, et al.
Pubblicazione: (2025)
di: Hatgis-Kessell, Stephane, et al.
Pubblicazione: (2025)
Local Data Quantity-Aware Weighted Averaging for Federated Learning with Dishonest Clients
di: Wu, Leming, et al.
Pubblicazione: (2025)
di: Wu, Leming, et al.
Pubblicazione: (2025)
Federated Loss Exploration for Improved Convergence on Non-IID Data
di: Internò, Christian, et al.
Pubblicazione: (2025)
di: Internò, Christian, et al.
Pubblicazione: (2025)
MO-EMT-NAS: Multi-Objective Continuous Transfer of Architectural Knowledge Between Tasks from Different Datasets
di: Liao, Peng, et al.
Pubblicazione: (2024)
di: Liao, Peng, et al.
Pubblicazione: (2024)
Mitigating Lost in Multi-turn Conversation via Curriculum RL with Verifiable Accuracy and Abstention Rewards
di: Li, Ming, et al.
Pubblicazione: (2025)
di: Li, Ming, et al.
Pubblicazione: (2025)
Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap
di: Qi, Xuan, et al.
Pubblicazione: (2025)
di: Qi, Xuan, et al.
Pubblicazione: (2025)
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF
di: Zhao, Kangwen, et al.
Pubblicazione: (2025)
di: Zhao, Kangwen, et al.
Pubblicazione: (2025)
DRPO: Efficient Reasoning via Decoupled Reward Policy Optimization
di: Li, Gang, et al.
Pubblicazione: (2025)
di: Li, Gang, et al.
Pubblicazione: (2025)
Policy Filtration for RLHF to Mitigate Noise in Reward Models
di: Zhang, Chuheng, et al.
Pubblicazione: (2024)
di: Zhang, Chuheng, et al.
Pubblicazione: (2024)
Revis: Sparse Latent Steering to Mitigate Object Hallucination in Large Vision-Language Models
di: Wu, Jialin, et al.
Pubblicazione: (2026)
di: Wu, Jialin, et al.
Pubblicazione: (2026)
IRIS: Intrinsic Reward Image Synthesis
di: Chen, Yihang, et al.
Pubblicazione: (2025)
di: Chen, Yihang, et al.
Pubblicazione: (2025)
Awakening Dormant Experts:Counterfactual Routing to Mitigate MoE Hallucinations
di: Hu, Wentao, et al.
Pubblicazione: (2026)
di: Hu, Wentao, et al.
Pubblicazione: (2026)
ConVis: Contrastive Decoding with Hallucination Visualization for Mitigating Hallucinations in Multimodal Large Language Models
di: Park, Yeji, et al.
Pubblicazione: (2024)
di: Park, Yeji, et al.
Pubblicazione: (2024)
IGN : Implicit Generative Networks
di: Luo, Haozheng, et al.
Pubblicazione: (2022)
di: Luo, Haozheng, et al.
Pubblicazione: (2022)
OmniMER: Auxiliary-Enhanced LLM Adaptation for Indonesian Multimodal Emotion Recognition
di: Yan, Xueming, et al.
Pubblicazione: (2025)
di: Yan, Xueming, et al.
Pubblicazione: (2025)
Mitigating LLM Hallucination via Behaviorally Calibrated Reinforcement Learning
di: Wu, Jiayun, et al.
Pubblicazione: (2025)
di: Wu, Jiayun, et al.
Pubblicazione: (2025)
InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling
di: Miao, Yuchun, et al.
Pubblicazione: (2024)
di: Miao, Yuchun, et al.
Pubblicazione: (2024)
Reinforcement Learning for Scalable Train Timetable Rescheduling with Graph Representation
di: Yue, Peng, et al.
Pubblicazione: (2024)
di: Yue, Peng, et al.
Pubblicazione: (2024)
Mitigating Value Hallucination in Dyna Planning via Multistep Predecessor Models
di: Aminmansour, Farzane, et al.
Pubblicazione: (2020)
di: Aminmansour, Farzane, et al.
Pubblicazione: (2020)
Auto-Rubric: Learning From Implicit Weights to Explicit Rubrics for Reward Modeling
di: Xie, Lipeng, et al.
Pubblicazione: (2025)
di: Xie, Lipeng, et al.
Pubblicazione: (2025)
PREFINE: Preference-Based Implicit Reward and Cost Fine-Tuning for Safety Alignment
di: Verma, Richa, et al.
Pubblicazione: (2026)
di: Verma, Richa, et al.
Pubblicazione: (2026)
ExpVG: Investigating the Design Space of Visual Grounding in Multimodal Large Language Model
di: Kang, Weitai, et al.
Pubblicazione: (2025)
di: Kang, Weitai, et al.
Pubblicazione: (2025)
Mitigating Hallucinated Translations in Large Language Models with Hallucination-focused Preference Optimization
di: Tang, Zilu, et al.
Pubblicazione: (2025)
di: Tang, Zilu, et al.
Pubblicazione: (2025)
Towards Mitigation of Hallucination for LLM-empowered Agents: Progressive Generalization Bound Exploration and Watchdog Monitor
di: Liu, Siyuan, et al.
Pubblicazione: (2025)
di: Liu, Siyuan, et al.
Pubblicazione: (2025)
MuPHI: Learning Implicit Multimodal Harm Reasoning via Semantically Grounded Reward Optimization
di: Saha, Anisha, et al.
Pubblicazione: (2026)
di: Saha, Anisha, et al.
Pubblicazione: (2026)
Documenti analoghi
-
HMACE: Heterogeneous Multi-Agent Collaborative Evolution for Combinatorial Optimization
di: Yan, Yuping, et al.
Pubblicazione: (2026) -
TriCon-SF: A Triple-Shuffle and Contribution-Aware Serial Federated Learning Framework for Heterogeneous Healthcare Data
di: Yan, Yuping, et al.
Pubblicazione: (2025) -
Mitigating Visual Hallucinations via Semantic Curriculum Preference Optimization in MLLMs
di: Li, Yuanshuai, et al.
Pubblicazione: (2025) -
OutSafe-Bench: A Benchmark for Multimodal Offensive Content Detection in Large Language Models
di: Yan, Yuping, et al.
Pubblicazione: (2025) -
Adversarial Reward Auditing for Active Detection and Mitigation of Reward Hacking
di: Beigi, Mohammad, et al.
Pubblicazione: (2026)