Towards Faithful and Controllable Personalization via Critique-Post-Edit Reinforcement Learning
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhu, Chenghao, Tao, Meiling, Wang, Tiannan, Ding, Dongyi, Jiang, Yuchen Eleanor, Zhou, Wangchunshu |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
PersonaFeedback: A Large-scale Human-annotated Benchmark For Personalization
par: Tao, Meiling, et autres
Publié: (2025)
par: Tao, Meiling, et autres
Publié: (2025)
AI PERSONA: Towards Life-long Personalization of LLMs
par: Wang, Tiannan, et autres
Publié: (2024)
par: Wang, Tiannan, et autres
Publié: (2024)
MiCoTA: Bridging the Learnability Gap with Intermediate CoT and Teacher Assistants
par: Ding, Dongyi, et autres
Publié: (2025)
par: Ding, Dongyi, et autres
Publié: (2025)
Symbolic Learning Enables Self-Evolving Agents
par: Zhou, Wangchunshu, et autres
Publié: (2024)
par: Zhou, Wangchunshu, et autres
Publié: (2024)
Towards Personalized Deep Research: Benchmarks and Evaluations
par: Liang, Yuan, et autres
Publié: (2025)
par: Liang, Yuan, et autres
Publié: (2025)
Edit Once, Update Everywhere: A Simple Framework for Cross-Lingual Knowledge Synchronization in LLMs
par: Wu, Yuchen, et autres
Publié: (2025)
par: Wu, Yuchen, et autres
Publié: (2025)
Critique-RL: Training Language Models for Critiquing through Two-Stage Reinforcement Learning
par: Xi, Zhiheng, et autres
Publié: (2025)
par: Xi, Zhiheng, et autres
Publié: (2025)
From Critique to Clarity: A Pathway to Faithful and Personalized Code Explanations with Large Language Models
par: Xu, Zexing, et autres
Publié: (2024)
par: Xu, Zexing, et autres
Publié: (2024)
Flash-Searcher: Fast and Effective Web Agents via DAG-Based Parallel Execution
par: Qin, Tianrui, et autres
Publié: (2025)
par: Qin, Tianrui, et autres
Publié: (2025)
AutoAct: Automatic Agent Learning from Scratch for QA via Self-Planning
par: Qiao, Shuofei, et autres
Publié: (2024)
par: Qiao, Shuofei, et autres
Publié: (2024)
Teaching Language Models to Critique via Reinforcement Learning
par: Xie, Zhihui, et autres
Publié: (2025)
par: Xie, Zhihui, et autres
Publié: (2025)
OThink-SRR1: Search, Refine and Reasoning with Reinforced Learning for Large Language Models
par: Liang, Haijian, et autres
Publié: (2026)
par: Liang, Haijian, et autres
Publié: (2026)
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL
par: Li, Weizhen, et autres
Publié: (2025)
par: Li, Weizhen, et autres
Publié: (2025)
Beyond Sparse Rewards: Enhancing Reinforcement Learning with Language Model Critique in Text Generation
par: Cao, Meng, et autres
Publié: (2024)
par: Cao, Meng, et autres
Publié: (2024)
Towards Generalizable and Faithful Logic Reasoning over Natural Language via Resolution Refutation
par: Sun, Zhouhao, et autres
Publié: (2024)
par: Sun, Zhouhao, et autres
Publié: (2024)
The Critique of Critique
par: Sun, Shichao, et autres
Publié: (2024)
par: Sun, Shichao, et autres
Publié: (2024)
Teaching Large Language Models to Maintain Contextual Faithfulness via Synthetic Tasks and Reinforcement Learning
par: Si, Shuzheng, et autres
Publié: (2025)
par: Si, Shuzheng, et autres
Publié: (2025)
InftyThink+: Effective and Efficient Infinite-Horizon Reasoning via Reinforcement Learning
par: Yan, Yuchen, et autres
Publié: (2026)
par: Yan, Yuchen, et autres
Publié: (2026)
CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation
par: Ke, Pei, et autres
Publié: (2023)
par: Ke, Pei, et autres
Publié: (2023)
A$^2$FM: An Adaptive Agent Foundation Model for Tool-Aware Hybrid Reasoning
par: Chen, Qianben, et autres
Publié: (2025)
par: Chen, Qianben, et autres
Publié: (2025)
MIMIR: A Streamlined Platform for Personalized Agent Tuning in Domain Expertise
par: Deng, Chunyuan, et autres
Publié: (2024)
par: Deng, Chunyuan, et autres
Publié: (2024)
FaithUn: Toward Faithful Forgetting in Language Models by Investigating the Interconnectedness of Knowledge
par: Yang, Nakyeong, et autres
Publié: (2025)
par: Yang, Nakyeong, et autres
Publié: (2025)
Towards Better Chain-of-Thought: A Reflection on Effectiveness and Faithfulness
par: Li, Jiachun, et autres
Publié: (2024)
par: Li, Jiachun, et autres
Publié: (2024)
PokerGPT: An End-to-End Lightweight Solver for Multi-Player Texas Hold'em via Large Language Model
par: Huang, Chenghao, et autres
Publié: (2024)
par: Huang, Chenghao, et autres
Publié: (2024)
Iterative Critique-Refine Framework for Enhancing LLM Personalization
par: Maram, Durga Prasad, et autres
Publié: (2025)
par: Maram, Durga Prasad, et autres
Publié: (2025)
SFR-RAG: Towards Contextually Faithful LLMs
par: Nguyen, Xuan-Phi, et autres
Publié: (2024)
par: Nguyen, Xuan-Phi, et autres
Publié: (2024)
Beyond Output Critique: Self-Correction via Task Distillation
par: Rahmani, Hossein A., et autres
Publié: (2026)
par: Rahmani, Hossein A., et autres
Publié: (2026)
FaithLM: Towards Faithful Explanations for Large Language Models
par: Chuang, Yu-Neng, et autres
Publié: (2024)
par: Chuang, Yu-Neng, et autres
Publié: (2024)
RoleCraft-GLM: Advancing Personalized Role-Playing in Large Language Models
par: Tao, Meiling, et autres
Publié: (2023)
par: Tao, Meiling, et autres
Publié: (2023)
PositionID: LLMs can Control Lengths, Copy and Paste with Explicit Positional Awareness
par: Wang, Zekun, et autres
Publié: (2024)
par: Wang, Zekun, et autres
Publié: (2024)
CTRL-RAG: Contrastive Likelihood Reward Based Reinforcement Learning for Context-Faithful RAG Models
par: Tan, Zhehao, et autres
Publié: (2026)
par: Tan, Zhehao, et autres
Publié: (2026)
Guiding Large Language Models to Post-Edit Machine Translation with Error Annotations
par: Ki, Dayeon, et autres
Publié: (2024)
par: Ki, Dayeon, et autres
Publié: (2024)
Knowledge-Infused Legal Wisdom: Navigating LLM Consultation through the Lens of Diagnostics and Positive-Unlabeled Reinforcement Learning
par: Wu, Yang, et autres
Publié: (2024)
par: Wu, Yang, et autres
Publié: (2024)
Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders
par: Xiong, Guangzhi, et autres
Publié: (2025)
par: Xiong, Guangzhi, et autres
Publié: (2025)
EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies
par: Hu, Xavier, et autres
Publié: (2026)
par: Hu, Xavier, et autres
Publié: (2026)
RealCritic: Towards Effectiveness-Driven Evaluation of Language Model Critiques
par: Tang, Zhengyang, et autres
Publié: (2025)
par: Tang, Zhengyang, et autres
Publié: (2025)
Efficient Agents: Building Effective Agents While Reducing Cost
par: Wang, Ningning, et autres
Publié: (2025)
par: Wang, Ningning, et autres
Publié: (2025)
MM-CRITIC: A Holistic Evaluation of Large Multimodal Models as Multimodal Critique
par: Zeng, Gailun, et autres
Publié: (2025)
par: Zeng, Gailun, et autres
Publié: (2025)
Verify Before You Commit: Towards Faithful Reasoning in LLM Agents via Self-Auditing
par: Yuan, Wenhao, et autres
Publié: (2026)
par: Yuan, Wenhao, et autres
Publié: (2026)
Faithfulness Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution Guidance
par: Alon, Bar, et autres
Publié: (2026)
par: Alon, Bar, et autres
Publié: (2026)
Documents similaires
-
PersonaFeedback: A Large-scale Human-annotated Benchmark For Personalization
par: Tao, Meiling, et autres
Publié: (2025) -
AI PERSONA: Towards Life-long Personalization of LLMs
par: Wang, Tiannan, et autres
Publié: (2024) -
MiCoTA: Bridging the Learnability Gap with Intermediate CoT and Teacher Assistants
par: Ding, Dongyi, et autres
Publié: (2025) -
Symbolic Learning Enables Self-Evolving Agents
par: Zhou, Wangchunshu, et autres
Publié: (2024) -
Towards Personalized Deep Research: Benchmarks and Evaluations
par: Liang, Yuan, et autres
Publié: (2025)