Learning to summarize user information for personalized reinforcement learning from human feedback
Fuente:
arXiv
Saved in:
| Main Authors: | Nam, Hyunji, Wan, Yanming, Liu, Mickel, Ahnn, Peter, Lian, Jianxun, Jaques, Natasha |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Maximizing Mutual Information Between Prompt and Response Improves LLM Performance With No Additional Data
by: Nam, Hyunji, et al.
Published: (2026)
by: Nam, Hyunji, et al.
Published: (2026)
Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning
by: Poddar, Sriyash, et al.
Published: (2024)
by: Poddar, Sriyash, et al.
Published: (2024)
Infer Human's Intentions Before Following Natural Language Instructions
by: Wan, Yanming, et al.
Published: (2024)
by: Wan, Yanming, et al.
Published: (2024)
Delayed homomorphic reinforcement learning for environments with delayed feedback
by: Lee, Jongsoo, et al.
Published: (2026)
by: Lee, Jongsoo, et al.
Published: (2026)
Generative Modeling for Robust Deep Reinforcement Learning on the Traveling Salesman Problem
by: Li, Michael, et al.
Published: (2025)
by: Li, Michael, et al.
Published: (2025)
Deep reinforcement learning for irrigation scheduling using high-dimensional sensor feedback
by: Saikai, Yuji, et al.
Published: (2023)
by: Saikai, Yuji, et al.
Published: (2023)
Logic-informed reinforcement learning for cross-domain optimization of large-scale cyber-physical systems
by: Wan, Guangxi, et al.
Published: (2025)
by: Wan, Guangxi, et al.
Published: (2025)
SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning
by: Liu, Bo, et al.
Published: (2025)
by: Liu, Bo, et al.
Published: (2025)
Impossibility Theorems for Feature Attribution
by: Bilodeau, Blair, et al.
Published: (2022)
by: Bilodeau, Blair, et al.
Published: (2022)
Using reinforcement learning to probe the role of feedback in skill acquisition
by: Terpin, Antonio, et al.
Published: (2025)
by: Terpin, Antonio, et al.
Published: (2025)
Learning to Cooperate with Humans using Generative Agents
by: Liang, Yancheng, et al.
Published: (2024)
by: Liang, Yancheng, et al.
Published: (2024)
Predicting Long Term Sequential Policy Value Using Softer Surrogates
by: Nam, Hyunji, et al.
Published: (2024)
by: Nam, Hyunji, et al.
Published: (2024)
Curriculum reinforcement learning with measurable task representation learning
by: Wen, Yongyan, et al.
Published: (2026)
by: Wen, Yongyan, et al.
Published: (2026)
Modeling Others' Minds as Code
by: Jha, Kunal, et al.
Published: (2025)
by: Jha, Kunal, et al.
Published: (2025)
RankSum An unsupervised extractive text summarization based on rank fusion
by: Joshi, A., et al.
Published: (2024)
by: Joshi, A., et al.
Published: (2024)
Normalization and effective learning rates in reinforcement learning
by: Lyle, Clare, et al.
Published: (2024)
by: Lyle, Clare, et al.
Published: (2024)
MaxInfoRL: Boosting exploration in reinforcement learning through information gain maximization
by: Sukhija, Bhavya, et al.
Published: (2024)
by: Sukhija, Bhavya, et al.
Published: (2024)
Physics-informed offline reinforcement learning eliminates catastrophic fuel waste in maritime routing
by: Bora, Aniruddha, et al.
Published: (2026)
by: Bora, Aniruddha, et al.
Published: (2026)
QoS prediction in radio vehicular environments via prior user information
by: Ain, Noor Ul, et al.
Published: (2024)
by: Ain, Noor Ul, et al.
Published: (2024)
Traffic expertise meets residual RL: Knowledge-informed model-based residual reinforcement learning for CAV trajectory control
by: Sheng, Zihao, et al.
Published: (2024)
by: Sheng, Zihao, et al.
Published: (2024)
Interpretable experiential learning based on state history and global feedback
by: Kolonin, Anton
Published: (2026)
by: Kolonin, Anton
Published: (2026)
Evaluating & Reducing Deceptive Dialogue From Language Models with Multi-turn RL
by: Abdulhai, Marwa, et al.
Published: (2025)
by: Abdulhai, Marwa, et al.
Published: (2025)
Bridging the phenotype-target gap for molecular generation via multi-objective reinforcement learning
by: Guo, Haotian, et al.
Published: (2025)
by: Guo, Haotian, et al.
Published: (2025)
Counterfactual experience augmented off-policy reinforcement learning
by: Lee, Sunbowen, et al.
Published: (2025)
by: Lee, Sunbowen, et al.
Published: (2025)
Causal prompting model-based offline reinforcement learning
by: Yu, Xuehui, et al.
Published: (2024)
by: Yu, Xuehui, et al.
Published: (2024)
Deep reinforcement learning with time-scale invariant memory
by: Kabir, Md Rysul, et al.
Published: (2024)
by: Kabir, Md Rysul, et al.
Published: (2024)
Offline reinforcement learning for job-shop scheduling problems
by: Echeverria, Imanol, et al.
Published: (2024)
by: Echeverria, Imanol, et al.
Published: (2024)
Bellman operator convergence enhancements in reinforcement learning algorithms
by: Kadurha, David Krame, et al.
Published: (2025)
by: Kadurha, David Krame, et al.
Published: (2025)
Cross-environment Cooperation Enables Zero-shot Multi-agent Coordination
by: Jha, Kunal, et al.
Published: (2025)
by: Jha, Kunal, et al.
Published: (2025)
Leveraging weights signals -- Predicting and improving generalizability in reinforcement learning
by: Moulin, Olivier, et al.
Published: (2025)
by: Moulin, Olivier, et al.
Published: (2025)
Dynamic feature selection in medical predictive monitoring by reinforcement learning
by: Chen, Yutong, et al.
Published: (2024)
by: Chen, Yutong, et al.
Published: (2024)
Economic span selection of bridge based on deep reinforcement learning
by: Zhang, Leye, et al.
Published: (2024)
by: Zhang, Leye, et al.
Published: (2024)
Not all tokens are needed(NAT): token efficient reinforcement learning
by: Sang, Hejian, et al.
Published: (2026)
by: Sang, Hejian, et al.
Published: (2026)
Enhancing Personalized Multi-Turn Dialogue with Curiosity Reward
by: Wan, Yanming, et al.
Published: (2025)
by: Wan, Yanming, et al.
Published: (2025)
Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning
by: Su, Xuerui, et al.
Published: (2025)
by: Su, Xuerui, et al.
Published: (2025)
BooookScore: A systematic exploration of book-length summarization in the era of LLMs
by: Chang, Yapei, et al.
Published: (2023)
by: Chang, Yapei, et al.
Published: (2023)
A dataset and benchmark for hospital course summarization with adapted large language models
by: Aali, Asad, et al.
Published: (2024)
by: Aali, Asad, et al.
Published: (2024)
Learning optimal treatment strategies for intraoperative hypotension using deep reinforcement learning
by: Adiyeke, Esra, et al.
Published: (2025)
by: Adiyeke, Esra, et al.
Published: (2025)
An efficient deep reinforcement learning environment for flexible job-shop scheduling
by: Wu, Xinquan, et al.
Published: (2025)
by: Wu, Xinquan, et al.
Published: (2025)
On the consistency of hyper-parameter selection in value-based deep reinforcement learning
by: Obando-Ceron, Johan, et al.
Published: (2024)
by: Obando-Ceron, Johan, et al.
Published: (2024)
Similar Items
-
Maximizing Mutual Information Between Prompt and Response Improves LLM Performance With No Additional Data
by: Nam, Hyunji, et al.
Published: (2026) -
Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning
by: Poddar, Sriyash, et al.
Published: (2024) -
Infer Human's Intentions Before Following Natural Language Instructions
by: Wan, Yanming, et al.
Published: (2024) -
Delayed homomorphic reinforcement learning for environments with delayed feedback
by: Lee, Jongsoo, et al.
Published: (2026) -
Generative Modeling for Robust Deep Reinforcement Learning on the Traveling Salesman Problem
by: Li, Michael, et al.
Published: (2025)