RL with Learnable Textual Feedback: A Bilevel Approach
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Singh, Utsav, Sredharan, Sidhaarth, Chakraborty, Souradip, Bedi, Amrit Singh |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Direct Preference Optimization for Primitive-Enabled Hierarchical RL: A Bilevel Approach
par: Singh, Utsav, et autres
Publié: (2024)
par: Singh, Utsav, et autres
Publié: (2024)
REBEL: Reward Regularization-Based Approach for Robotic Reinforcement Learning from Human Feedback
par: Chakraborty, Souradip, et autres
Publié: (2023)
par: Chakraborty, Souradip, et autres
Publié: (2023)
On The Sample Complexity Bounds In Bilevel Reinforcement Learning
par: Gaur, Mudit, et autres
Publié: (2025)
par: Gaur, Mudit, et autres
Publié: (2025)
DIPPER: Direct Preference Optimization to Accelerate Primitive-Enabled Hierarchical Reinforcement Learning
par: Singh, Utsav, et autres
Publié: (2024)
par: Singh, Utsav, et autres
Publié: (2024)
Why Pass@k Optimization Can Degrade Pass@1: Prompt Interference in LLM Post-training
par: Barakat, Anas, et autres
Publié: (2026)
par: Barakat, Anas, et autres
Publié: (2026)
MIRA: Towards Mitigating Reward Hacking in Inference-Time Alignment of T2I Diffusion Models
par: Zhai, Kevin, et autres
Publié: (2025)
par: Zhai, Kevin, et autres
Publié: (2025)
PARL: A Unified Framework for Policy Alignment in Reinforcement Learning from Human Feedback
par: Chakraborty, Souradip, et autres
Publié: (2023)
par: Chakraborty, Souradip, et autres
Publié: (2023)
On the Global Optimality of Policy Gradient Methods in General Utility Reinforcement Learning
par: Barakat, Anas, et autres
Publié: (2024)
par: Barakat, Anas, et autres
Publié: (2024)
Align-Pro: A Principled Approach to Prompt Optimization for LLM Alignment
par: Trivedi, Prashant, et autres
Publié: (2025)
par: Trivedi, Prashant, et autres
Publié: (2025)
Code Comprehension then Auditing for Unsupervised LLM Evaluation
par: Patel, Bhrij, et autres
Publié: (2024)
par: Patel, Bhrij, et autres
Publié: (2024)
VARP: Reinforcement Learning from Vision-Language Model Feedback with Agent Regularized Preferences
par: Singh, Anukriti, et autres
Publié: (2025)
par: Singh, Anukriti, et autres
Publié: (2025)
Uncertainty-Aware Answer Selection for Improved Reasoning in Multi-LLM Systems
par: Agrawal, Aakriti, et autres
Publié: (2025)
par: Agrawal, Aakriti, et autres
Publié: (2025)
PROPS: Progressively Private Self-alignment of Large Language Models
par: Teku, Noel, et autres
Publié: (2025)
par: Teku, Noel, et autres
Publié: (2025)
PIPER: Primitive-Informed Preference-based Hierarchical Reinforcement Learning via Hindsight Relabeling
par: Singh, Utsav, et autres
Publié: (2024)
par: Singh, Utsav, et autres
Publié: (2024)
Transfer Q Star: Principled Decoding for LLM Alignment
par: Chakraborty, Souradip, et autres
Publié: (2024)
par: Chakraborty, Souradip, et autres
Publié: (2024)
MaxMin-RLHF: Alignment with Diverse Human Preferences
par: Chakraborty, Souradip, et autres
Publié: (2024)
par: Chakraborty, Souradip, et autres
Publié: (2024)
Achieving Zero Constraint Violation for Constrained Reinforcement Learning via Conservative Natural Policy Gradient Primal-Dual Algorithm
par: Bai, Qinbo, et autres
Publié: (2022)
par: Bai, Qinbo, et autres
Publié: (2022)
On The Global Convergence Of Online RLHF With Neural Parametrization
par: Gaur, Mudit, et autres
Publié: (2024)
par: Gaur, Mudit, et autres
Publié: (2024)
Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment
par: Ghosal, Soumya Suvra, et autres
Publié: (2024)
par: Ghosal, Soumya Suvra, et autres
Publié: (2024)
Test-Time Scaling in Diffusion LLMs via Hidden Semi-Autoregressive Experts
par: Lee, Jihoon, et autres
Publié: (2025)
par: Lee, Jihoon, et autres
Publié: (2025)
Closing the Gap: Achieving Global Convergence (Last Iterate) of Actor-Critic under Markovian Sampling with Neural Network Parametrization
par: Gaur, Mudit, et autres
Publié: (2024)
par: Gaur, Mudit, et autres
Publié: (2024)
SAIL: Self-Improving Efficient Online Alignment of Large Language Models
par: Ding, Mucong, et autres
Publié: (2024)
par: Ding, Mucong, et autres
Publié: (2024)
Rethinking Adversarial Policies: A Generalized Attack Formulation and Provable Defense in RL
par: Liu, Xiangyu, et autres
Publié: (2023)
par: Liu, Xiangyu, et autres
Publié: (2023)
Draft-Conditioned Constrained Decoding for Structured Generation in LLMs
par: Reddy, Avinash, et autres
Publié: (2026)
par: Reddy, Avinash, et autres
Publié: (2026)
TRAM: Test-Time Risk Adaptation with Mixture of Agents
par: Chehade, Mohamad Fares El Hajj, et autres
Publié: (2024)
par: Chehade, Mohamad Fares El Hajj, et autres
Publié: (2024)
Multi-LLM QA with Embodied Exploration
par: Patel, Bhrij, et autres
Publié: (2024)
par: Patel, Bhrij, et autres
Publié: (2024)
Improved Sample Complexity For Diffusion Model Training Without Empirical Risk Minimizer Access
par: Gaur, Mudit, et autres
Publié: (2025)
par: Gaur, Mudit, et autres
Publié: (2025)
Beyond Text: Utilizing Vocal Cues to Improve Decision Making in LLMs for Robot Navigation Tasks
par: Sun, Xingpeng, et autres
Publié: (2024)
par: Sun, Xingpeng, et autres
Publié: (2024)
PEAR: Primitive Enabled Adaptive Relabeling for Boosting Hierarchical Reinforcement Learning
par: Singh, Utsav, et autres
Publié: (2023)
par: Singh, Utsav, et autres
Publié: (2023)
CRISP: Curriculum Inducing Primitive Informed Subgoal Prediction for Hierarchical Reinforcement Learning
par: Singh, Utsav, et autres
Publié: (2023)
par: Singh, Utsav, et autres
Publié: (2023)
SELF-PERCEPT: Introspection Improves Large Language Models' Detection of Multi-Person Mental Manipulation in Conversations
par: Khanna, Danush, et autres
Publié: (2025)
par: Khanna, Danush, et autres
Publié: (2025)
FACT or Fiction: Can Truthful Mechanisms Eliminate Federated Free Riding?
par: Bornstein, Marco, et autres
Publié: (2024)
par: Bornstein, Marco, et autres
Publié: (2024)
SafeR-CLIP: Mitigating NSFW Content in Vision-Language Models While Preserving Pre-Trained Knowledge
par: Yousaf, Adeel, et autres
Publié: (2025)
par: Yousaf, Adeel, et autres
Publié: (2025)
Safety Recovery in Reasoning Models Is Only a Few Early Steering Steps Away
par: Ghosal, Soumya Suvra, et autres
Publié: (2026)
par: Ghosal, Soumya Suvra, et autres
Publié: (2026)
Generative Modeling with Continuous Flows: Sample Complexity of Flow Matching
par: Gaur, Mudit, et autres
Publié: (2025)
par: Gaur, Mudit, et autres
Publié: (2025)
BalancedDPO: Adaptive Multi-Metric Alignment
par: Tamboli, Dipesh, et autres
Publié: (2025)
par: Tamboli, Dipesh, et autres
Publié: (2025)
LIAR: Leveraging Inference Time Alignment (Best-of-N) to Jailbreak LLMs in Seconds
par: Beetham, James, et autres
Publié: (2024)
par: Beetham, James, et autres
Publié: (2024)
Towards Global Optimality for Practical Average Reward Reinforcement Learning without Mixing Time Oracles
par: Patel, Bhrij, et autres
Publié: (2024)
par: Patel, Bhrij, et autres
Publié: (2024)
Free and Customizable Code Documentation with LLMs: A Fine-Tuning Approach
par: Chakrabarty, Sayak, et autres
Publié: (2024)
par: Chakrabarty, Sayak, et autres
Publié: (2024)
LeARN: Learnable and Adaptive Representations for Nonlinear Dynamics in System Identification
par: Singh, Arunabh, et autres
Publié: (2024)
par: Singh, Arunabh, et autres
Publié: (2024)
Documents similaires
-
Direct Preference Optimization for Primitive-Enabled Hierarchical RL: A Bilevel Approach
par: Singh, Utsav, et autres
Publié: (2024) -
REBEL: Reward Regularization-Based Approach for Robotic Reinforcement Learning from Human Feedback
par: Chakraborty, Souradip, et autres
Publié: (2023) -
On The Sample Complexity Bounds In Bilevel Reinforcement Learning
par: Gaur, Mudit, et autres
Publié: (2025) -
DIPPER: Direct Preference Optimization to Accelerate Primitive-Enabled Hierarchical Reinforcement Learning
par: Singh, Utsav, et autres
Publié: (2024) -
Why Pass@k Optimization Can Degrade Pass@1: Prompt Interference in LLM Post-training
par: Barakat, Anas, et autres
Publié: (2026)