On-Policy Self-Alignment with Fine-grained Knowledge Feedback for Hallucination Mitigation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wen, Xueru, Lou, Jie, Lu, Xinyu, Yuqiu, Ji, Guan, Xinyan, Lu, Yaojie, Lin, Hongyu, He, Ben, Han, Xianpei, Zhang, Debing, Sun, Le |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Devil Is in the Details: Tackling Unimodal Spurious Correlations for Generalizable Multimodal Reward Models
von: Li, Zichao, et al.
Veröffentlicht: (2025)
von: Li, Zichao, et al.
Veröffentlicht: (2025)
Critic-CoT: Boosting the reasoning abilities of large language model via Chain-of-thoughts Critic
von: Zheng, Xin, et al.
Veröffentlicht: (2024)
von: Zheng, Xin, et al.
Veröffentlicht: (2024)
Rethinking Reward Model Evaluation: Are We Barking up the Wrong Tree?
von: Wen, Xueru, et al.
Veröffentlicht: (2024)
von: Wen, Xueru, et al.
Veröffentlicht: (2024)
Your Teacher Can't Help You Here: Combating Supervision Fidelity Decay in On-Policy Distillation
von: Liu, Yanjiang, et al.
Veröffentlicht: (2026)
von: Liu, Yanjiang, et al.
Veröffentlicht: (2026)
Coupled Variational Reinforcement Learning for Language Model General Reasoning
von: Wen, Xueru, et al.
Veröffentlicht: (2025)
von: Wen, Xueru, et al.
Veröffentlicht: (2025)
Cheems: A Practical Guidance for Building and Evaluating Chinese Reward Models from Scratch
von: Wen, Xueru, et al.
Veröffentlicht: (2025)
von: Wen, Xueru, et al.
Veröffentlicht: (2025)
Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation
von: Yuan, Qianhao, et al.
Veröffentlicht: (2026)
von: Yuan, Qianhao, et al.
Veröffentlicht: (2026)
Scalable Oversight for Superhuman AI via Recursive Self-Critiquing
von: Wen, Xueru, et al.
Veröffentlicht: (2025)
von: Wen, Xueru, et al.
Veröffentlicht: (2025)
Search, Verify and Feedback: Towards Next Generation Post-training Paradigm of Foundation Models via Verifier Engineering
von: Guan, Xinyan, et al.
Veröffentlicht: (2024)
von: Guan, Xinyan, et al.
Veröffentlicht: (2024)
Learning from Failures: Correction-Oriented Policy Optimization with Verifiable Rewards
von: Ren, Mengjie, et al.
Veröffentlicht: (2026)
von: Ren, Mengjie, et al.
Veröffentlicht: (2026)
REInstruct: Building Instruction Data from Unlabeled Corpus
von: Chen, Shu, et al.
Veröffentlicht: (2024)
von: Chen, Shu, et al.
Veröffentlicht: (2024)
When Models Outthink Their Safety: Unveiling and Mitigating Self-Jailbreak in Large Reasoning Models
von: Mao, Yingzhi, et al.
Veröffentlicht: (2025)
von: Mao, Yingzhi, et al.
Veröffentlicht: (2025)
Transferable Post-training via Inverse Value Learning
von: Lu, Xinyu, et al.
Veröffentlicht: (2024)
von: Lu, Xinyu, et al.
Veröffentlicht: (2024)
MemSearcher: Training LLMs to Reason, Search and Manage Memory via End-to-End Reinforcement Learning
von: Yuan, Qianhao, et al.
Veröffentlicht: (2025)
von: Yuan, Qianhao, et al.
Veröffentlicht: (2025)
SoFA: Shielded On-the-fly Alignment via Priority Rule Following
von: Lu, Xinyu, et al.
Veröffentlicht: (2024)
von: Lu, Xinyu, et al.
Veröffentlicht: (2024)
DeepRAG: Thinking to Retrieve Step by Step for Large Language Models
von: Guan, Xinyan, et al.
Veröffentlicht: (2025)
von: Guan, Xinyan, et al.
Veröffentlicht: (2025)
Meta-Cognitive Analysis: Evaluating Declarative and Procedural Knowledge in Datasets and Large Language Models
von: Li, Zhuoqun, et al.
Veröffentlicht: (2024)
von: Li, Zhuoqun, et al.
Veröffentlicht: (2024)
P^2O: Joint Policy and Prompt Optimization
von: Lu, Xinyu, et al.
Veröffentlicht: (2026)
von: Lu, Xinyu, et al.
Veröffentlicht: (2026)
Rule or Story, Which is a Better Commonsense Expression for Talking with Large Language Models?
von: Bian, Ning, et al.
Veröffentlicht: (2024)
von: Bian, Ning, et al.
Veröffentlicht: (2024)
Tackling Length Inflation Without Trade-offs: Group Relative Reward Rescaling for Reinforcement Learning
von: Li, Zichao, et al.
Veröffentlicht: (2026)
von: Li, Zichao, et al.
Veröffentlicht: (2026)
Decoupling Reasoning and Confidence: Resurrecting Calibration in Reinforcement Learning from Verifiable Rewards
von: Ma, Zhengzhao, et al.
Veröffentlicht: (2026)
von: Ma, Zhengzhao, et al.
Veröffentlicht: (2026)
PPTAgent: Generating and Evaluating Presentations Beyond Text-to-Slides
von: Zheng, Hao, et al.
Veröffentlicht: (2025)
von: Zheng, Hao, et al.
Veröffentlicht: (2025)
ConsistentChat: Building Skeleton-Guided Consistent Multi-Turn Dialogues for Large Language Models from Scratch
von: Chen, Jiawei, et al.
Veröffentlicht: (2025)
von: Chen, Jiawei, et al.
Veröffentlicht: (2025)
ChatGPT is a Knowledgeable but Inexperienced Solver: An Investigation of Commonsense Problem in Large Language Models
von: Bian, Ning, et al.
Veröffentlicht: (2023)
von: Bian, Ning, et al.
Veröffentlicht: (2023)
Towards Scalable Automated Alignment of LLMs: A Survey
von: Cao, Boxi, et al.
Veröffentlicht: (2024)
von: Cao, Boxi, et al.
Veröffentlicht: (2024)
Memorizing is Not Enough: Deep Knowledge Injection Through Reasoning
von: Xu, Ruoxi, et al.
Veröffentlicht: (2025)
von: Xu, Ruoxi, et al.
Veröffentlicht: (2025)
PaperRegister: Boosting Flexible-grained Paper Search via Hierarchical Register Indexing
von: Li, Zhuoqun, et al.
Veröffentlicht: (2025)
von: Li, Zhuoqun, et al.
Veröffentlicht: (2025)
LiteCoder-Terminal: Scaling Long-Horizon Terminal Environments for Learning Language Agents
von: Peng, Xiaoxuan, et al.
Veröffentlicht: (2026)
von: Peng, Xiaoxuan, et al.
Veröffentlicht: (2026)
SAISA: Towards Multimodal Large Language Models with Both Training and Inference Efficiency
von: Yuan, Qianhao, et al.
Veröffentlicht: (2025)
von: Yuan, Qianhao, et al.
Veröffentlicht: (2025)
RefCritic: Training Long Chain-of-Thought Critic Models with Refinement Feedback
von: Tang, Qiaoyu, et al.
Veröffentlicht: (2025)
von: Tang, Qiaoyu, et al.
Veröffentlicht: (2025)
Self-Steering Optimization: Autonomous Preference Optimization for Large Language Models
von: Xiang, Hao, et al.
Veröffentlicht: (2024)
von: Xiang, Hao, et al.
Veröffentlicht: (2024)
Executing Natural Language-Described Algorithms with Large Language Models: An Investigation
von: Zheng, Xin, et al.
Veröffentlicht: (2024)
von: Zheng, Xin, et al.
Veröffentlicht: (2024)
Open Grounded Planning: Challenges and Benchmark Construction
von: Guo, Shiguang, et al.
Veröffentlicht: (2024)
von: Guo, Shiguang, et al.
Veröffentlicht: (2024)
A Unified View of Delta Parameter Editing in Post-Trained Large-Scale Models
von: Tang, Qiaoyu, et al.
Veröffentlicht: (2024)
von: Tang, Qiaoyu, et al.
Veröffentlicht: (2024)
Expanding the Boundaries of Vision Prior Knowledge in Multi-modal Large Language Models
von: Liang, Qiao, et al.
Veröffentlicht: (2025)
von: Liang, Qiao, et al.
Veröffentlicht: (2025)
Few-shot Named Entity Recognition via Superposition Concept Discrimination
von: Chen, Jiawei, et al.
Veröffentlicht: (2024)
von: Chen, Jiawei, et al.
Veröffentlicht: (2024)
Multi-Facet Counterfactual Learning for Content Quality Evaluation
von: Zheng, Jiasheng, et al.
Veröffentlicht: (2024)
von: Zheng, Jiasheng, et al.
Veröffentlicht: (2024)
Beyond Text-Dominance: Understanding Modality Preference of Omni-modal Large Language Models
von: Yan, Xinru, et al.
Veröffentlicht: (2026)
von: Yan, Xinru, et al.
Veröffentlicht: (2026)
CRUXEval-X: A Benchmark for Multilingual Code Reasoning, Understanding and Execution
von: Xu, Ruiyang, et al.
Veröffentlicht: (2024)
von: Xu, Ruiyang, et al.
Veröffentlicht: (2024)
All Languages Matter: Understanding and Mitigating Language Bias in Multilingual RAG
von: Wang, Dan, et al.
Veröffentlicht: (2026)
von: Wang, Dan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
The Devil Is in the Details: Tackling Unimodal Spurious Correlations for Generalizable Multimodal Reward Models
von: Li, Zichao, et al.
Veröffentlicht: (2025) -
Critic-CoT: Boosting the reasoning abilities of large language model via Chain-of-thoughts Critic
von: Zheng, Xin, et al.
Veröffentlicht: (2024) -
Rethinking Reward Model Evaluation: Are We Barking up the Wrong Tree?
von: Wen, Xueru, et al.
Veröffentlicht: (2024) -
Your Teacher Can't Help You Here: Combating Supervision Fidelity Decay in On-Policy Distillation
von: Liu, Yanjiang, et al.
Veröffentlicht: (2026) -
Coupled Variational Reinforcement Learning for Language Model General Reasoning
von: Wen, Xueru, et al.
Veröffentlicht: (2025)