KL-Regularised Q-Learning: A Token-level Action-Value perspective on Online RLHF
Fuente:
arXiv
Saved in:
| Main Authors: | Brown, Jason R, Wells, Lennie, Young, Edward James, Bacallado, Sergio |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning from Preferences and Mixed Demonstrations in General Settings
by: Brown, Jason R, et al.
Published: (2025)
by: Brown, Jason R, et al.
Published: (2025)
An Improved Adaptive PID Optimizer with Enhanced Convergence and Stability for Deep Learning
by: Saini, Saurabh, et al.
Published: (2026)
by: Saini, Saurabh, et al.
Published: (2026)
CellARC: Measuring Intelligence with Cellular Automata
by: Lžičař, Miroslav
Published: (2025)
by: Lžičař, Miroslav
Published: (2025)
Emotion-Inspired Learning Signals (EILS): A Homeostatic Framework for Adaptive Autonomous Agents
by: Tiwari, Dhruv
Published: (2025)
by: Tiwari, Dhruv
Published: (2025)
A Comparison Between Decision Transformers and Traditional Offline Reinforcement Learning Algorithms
by: Caunhye, Ali Murtaza, et al.
Published: (2025)
by: Caunhye, Ali Murtaza, et al.
Published: (2025)
Adaptive Negative Scheduling for Graph Contrastive Learning
by: Ali, Adnan, et al.
Published: (2026)
by: Ali, Adnan, et al.
Published: (2026)
The Impact of Data Characteristics on GNN Evaluation for Detecting Fake News
by: Karn, Isha, et al.
Published: (2025)
by: Karn, Isha, et al.
Published: (2025)
Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
by: Pather, Kaviraj, et al.
Published: (2025)
by: Pather, Kaviraj, et al.
Published: (2025)
Streaming Continual Learning for Unified Adaptive Intelligence in Dynamic Environments
by: Giannini, Federico, et al.
Published: (2026)
by: Giannini, Federico, et al.
Published: (2026)
When Does Global Attention Help? A Unified Empirical Study on Atomistic Graph Learning
by: Chowdhury, Arindam, et al.
Published: (2025)
by: Chowdhury, Arindam, et al.
Published: (2025)
AI and Machine Learning Approaches for Predicting Nanoparticles Toxicity The Critical Role of Physiochemical Properties
by: Yousaf, Iqra
Published: (2024)
by: Yousaf, Iqra
Published: (2024)
From Imitation to Interaction: Mastering Game of Schnapsen with Shallow Reinforcement Learning
by: Klačan, Ján, et al.
Published: (2026)
by: Klačan, Ján, et al.
Published: (2026)
On the Fundamental Limitations of Decentralized Learnable Reward Shaping in Cooperative Multi-Agent Reinforcement Learning
by: Akella, Aditya
Published: (2025)
by: Akella, Aditya
Published: (2025)
Normalisation and Initialisation Strategies for Graph Neural Networks in Blockchain Anomaly Detection
by: Duy, Dang Sy, et al.
Published: (2026)
by: Duy, Dang Sy, et al.
Published: (2026)
Function-Valued Causal Influence in Nonlinear Time Series
by: Kuskova, Valentina V., et al.
Published: (2026)
by: Kuskova, Valentina V., et al.
Published: (2026)
Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models
by: Fu, Tianyu, et al.
Published: (2025)
by: Fu, Tianyu, et al.
Published: (2025)
GoldenStart: Q-Guided Priors and Entropy Control for Distilling Flow Policies
by: Zhang, He, et al.
Published: (2026)
by: Zhang, He, et al.
Published: (2026)
Evolving machine learning workflows through interactive AutoML
by: Barbudo, Rafael, et al.
Published: (2024)
by: Barbudo, Rafael, et al.
Published: (2024)
UWM-JEPA: Predictive World Models That Imagine in Belief Space
by: Radha, Santosh Kumar, et al.
Published: (2026)
by: Radha, Santosh Kumar, et al.
Published: (2026)
Federated Distributional Reinforcement Learning with Distributional Critic Regularization
by: Millard, David, et al.
Published: (2026)
by: Millard, David, et al.
Published: (2026)
Intersymbolic AI: Interlinking Symbolic AI and Subsymbolic AI
by: Platzer, André
Published: (2024)
by: Platzer, André
Published: (2024)
Multimodal Generative AI for Story Point Estimation in Software Development
by: Islam, Mohammad Rubyet, et al.
Published: (2025)
by: Islam, Mohammad Rubyet, et al.
Published: (2025)
AI Agents: Evolution, Architecture, and Real-World Applications
by: Krishnan, Naveen
Published: (2025)
by: Krishnan, Naveen
Published: (2025)
Adaptive Minds: Empowering Agents with LoRA-as-Tools
by: Shekar, Pavan C, et al.
Published: (2025)
by: Shekar, Pavan C, et al.
Published: (2025)
XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoML
by: Estevanell-Valladares, Ernesto L., et al.
Published: (2025)
by: Estevanell-Valladares, Ernesto L., et al.
Published: (2025)
Context Engineering for Multi-Agent LLM Code Assistants Using Elicit, NotebookLM, ChatGPT, and Claude Code
by: Haseeb, Muhammad
Published: (2025)
by: Haseeb, Muhammad
Published: (2025)
Self-Supervised On-Policy Reinforcement Learning via Contrastive Proximal Policy Optimisation
by: Osman, Asim, et al.
Published: (2026)
by: Osman, Asim, et al.
Published: (2026)
Improving LLM Agent Planning with In-Context Learning via Atomic Fact Augmentation and Lookahead Search
by: Holt, Samuel, et al.
Published: (2025)
by: Holt, Samuel, et al.
Published: (2025)
SMOSE: Sparse Mixture of Shallow Experts for Interpretable Reinforcement Learning in Continuous Control Tasks
by: Vincze, Mátyás, et al.
Published: (2024)
by: Vincze, Mátyás, et al.
Published: (2024)
CogniLoad: A Synthetic Natural Language Reasoning Benchmark With Tunable Length, Intrinsic Difficulty, and Distractor Density
by: Kaiser, Daniel, et al.
Published: (2025)
by: Kaiser, Daniel, et al.
Published: (2025)
The two clocks and the innovation window: When and how generative models learn rules
by: Wang, Binxu, et al.
Published: (2026)
by: Wang, Binxu, et al.
Published: (2026)
iLTM: Integrated Large Tabular Model
by: Bonet, David, et al.
Published: (2025)
by: Bonet, David, et al.
Published: (2025)
Learning to Select Goals in Automated Planning with Deep-Q Learning
by: Núñez-Molina, Carlos, et al.
Published: (2024)
by: Núñez-Molina, Carlos, et al.
Published: (2024)
Improving Efficiency of Sampling-based Motion Planning via Message-Passing Monte Carlo
by: Chahine, Makram, et al.
Published: (2024)
by: Chahine, Makram, et al.
Published: (2024)
Agentic UAVs: LLM-Driven Autonomy with Integrated Tool-Calling and Cognitive Reasoning
by: Koubaa, Anis, et al.
Published: (2025)
by: Koubaa, Anis, et al.
Published: (2025)
SuS: Strategy-aware Surprise for Intrinsic Exploration
by: Kashirskiy, Mark, et al.
Published: (2026)
by: Kashirskiy, Mark, et al.
Published: (2026)
FDQN: A Flexible Deep Q-Network Framework for Game Automation
by: Gujavarthy, Prabhath Reddy
Published: (2024)
by: Gujavarthy, Prabhath Reddy
Published: (2024)
TRIZ Agents: A Multi-Agent LLM Approach for TRIZ-Based Innovation
by: Szczepanik, Kamil, et al.
Published: (2025)
by: Szczepanik, Kamil, et al.
Published: (2025)
Collaborative LLM Agents for C4 Software Architecture Design Automation
by: Szczepanik, Kamil, et al.
Published: (2025)
by: Szczepanik, Kamil, et al.
Published: (2025)
Law-Strength Frontiers and a No-Free-Lunch Result for Law-Seeking Reinforcement Learning on Volatility Law Manifolds
by: Zhang, Jian'an
Published: (2025)
by: Zhang, Jian'an
Published: (2025)
Similar Items
-
Learning from Preferences and Mixed Demonstrations in General Settings
by: Brown, Jason R, et al.
Published: (2025) -
An Improved Adaptive PID Optimizer with Enhanced Convergence and Stability for Deep Learning
by: Saini, Saurabh, et al.
Published: (2026) -
CellARC: Measuring Intelligence with Cellular Automata
by: Lžičař, Miroslav
Published: (2025) -
Emotion-Inspired Learning Signals (EILS): A Homeostatic Framework for Adaptive Autonomous Agents
by: Tiwari, Dhruv
Published: (2025) -
A Comparison Between Decision Transformers and Traditional Offline Reinforcement Learning Algorithms
by: Caunhye, Ali Murtaza, et al.
Published: (2025)