Better than Your Teacher: LLM Agents that learn from Privileged AI Feedback
Fuente:
arXiv
Saved in:
| Main Authors: | Choudhury, Sanjiban, Sodhi, Paloma |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Process Reward Models for LLM Agents: Practical Framework and Directions
by: Choudhury, Sanjiban
Published: (2025)
by: Choudhury, Sanjiban
Published: (2025)
Efficient Imitation under Misspecification
by: Espinosa-Dice, Nicolas, et al.
Published: (2025)
by: Espinosa-Dice, Nicolas, et al.
Published: (2025)
Motion Tracks: A Unified Representation for Human-Robot Transfer in Few-Shot Imitation Learning
by: Ren, Juntao, et al.
Published: (2025)
by: Ren, Juntao, et al.
Published: (2025)
InteRACT: Transformer Models for Human Intent Prediction Conditioned on Robot Actions
by: Kedia, Kushal, et al.
Published: (2023)
by: Kedia, Kushal, et al.
Published: (2023)
Hybrid Inverse Reinforcement Learning
by: Ren, Juntao, et al.
Published: (2024)
by: Ren, Juntao, et al.
Published: (2024)
One-Shot Imitation under Mismatched Execution
by: Kedia, Kushal, et al.
Published: (2024)
by: Kedia, Kushal, et al.
Published: (2024)
Non-Adversarial Inverse Reinforcement Learning via Successor Feature Matching
by: Jain, Arnav Kumar, et al.
Published: (2024)
by: Jain, Arnav Kumar, et al.
Published: (2024)
Imitation Learning via Focused Satisficing
by: Shah, Rushit N., et al.
Published: (2025)
by: Shah, Rushit N., et al.
Published: (2025)
Rethinking Data Curation in LLM Training: Online Reweighting Offers Better Generalization than Offline Methods
by: Zhao, Wanru, et al.
Published: (2026)
by: Zhao, Wanru, et al.
Published: (2026)
How to Train Your LLM Web Agent: A Statistical Diagnosis
by: Vattikonda, Dheeraj, et al.
Published: (2025)
by: Vattikonda, Dheeraj, et al.
Published: (2025)
Is Data Shapley Not Better than Random in Data Selection? Ask NASH
by: Tian, Xiao, et al.
Published: (2026)
by: Tian, Xiao, et al.
Published: (2026)
AVSD: Adaptive-View Self-Distillation by Balancing Consensus and Teacher-Specific Privileged Signals
by: Nguyen, Duy, et al.
Published: (2026)
by: Nguyen, Duy, et al.
Published: (2026)
Privileged Information Distillation for Language Models
by: Penaloza, Emiliano, et al.
Published: (2026)
by: Penaloza, Emiliano, et al.
Published: (2026)
Spend Less, Reason Better: Budget-Aware Value Tree Search for LLM Agents
by: Li, Yushu, et al.
Published: (2026)
by: Li, Yushu, et al.
Published: (2026)
The Devil is in the Condition Numbers: Why is GLU Better than non-GLU Structure?
by: Lyu, Xingyu, et al.
Published: (2026)
by: Lyu, Xingyu, et al.
Published: (2026)
Two Heads Are Better than One: Simulating Large Transformers with Small Ones
by: Yu, Hantao, et al.
Published: (2025)
by: Yu, Hantao, et al.
Published: (2025)
Multi-Turn Code Generation Through Single-Step Rewards
by: Jain, Arnav Kumar, et al.
Published: (2025)
by: Jain, Arnav Kumar, et al.
Published: (2025)
X-Sim: Cross-Embodiment Learning via Real-to-Sim-to-Real
by: Dan, Prithwish, et al.
Published: (2025)
by: Dan, Prithwish, et al.
Published: (2025)
Efficient Bias Mitigation Without Privileged Information
by: Zarlenga, Mateo Espinosa, et al.
Published: (2024)
by: Zarlenga, Mateo Espinosa, et al.
Published: (2024)
An Integrated Approach to AI-Generated Content in e-health
by: Ahmed, Tasnim, et al.
Published: (2025)
by: Ahmed, Tasnim, et al.
Published: (2025)
SOM Directions are Better than One: Multi-Directional Refusal Suppression in Language Models
by: Piras, Giorgio, et al.
Published: (2025)
by: Piras, Giorgio, et al.
Published: (2025)
Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agents
by: Shao, Shuai, et al.
Published: (2025)
by: Shao, Shuai, et al.
Published: (2025)
DataEnvGym: Data Generation Agents in Teacher Environments with Student Feedback
by: Khan, Zaid, et al.
Published: (2024)
by: Khan, Zaid, et al.
Published: (2024)
Privileged Sensing Scaffolds Reinforcement Learning
by: Hu, Edward S., et al.
Published: (2024)
by: Hu, Edward S., et al.
Published: (2024)
Distilling Realizable Students from Unrealizable Teachers
by: Kim, Yujin, et al.
Published: (2025)
by: Kim, Yujin, et al.
Published: (2025)
Distilling Privileged Information for Dubins Traveling Salesman Problems with Neighborhoods
by: Shin, Min Kyu, et al.
Published: (2024)
by: Shin, Min Kyu, et al.
Published: (2024)
Toward Privileged Foundation Models:LUPI for Accelerated and Improved Learning
by: Ding, Xueying, et al.
Published: (2026)
by: Ding, Xueying, et al.
Published: (2026)
This Looks Better than That: Better Interpretable Models with ProtoPNeXt
by: Willard, Frank, et al.
Published: (2024)
by: Willard, Frank, et al.
Published: (2024)
CASE: Efficient Curricular Data Pre-training for Building Assistive Psychology Expert Models
by: Harne, Sarthak, et al.
Published: (2024)
by: Harne, Sarthak, et al.
Published: (2024)
Workspace Optimization: How to Train Your Agent
by: Sarafian, Elad, et al.
Published: (2026)
by: Sarafian, Elad, et al.
Published: (2026)
When Your AIs Deceive You: Challenges of Partial Observability in Reinforcement Learning from Human Feedback
by: Lang, Leon, et al.
Published: (2024)
by: Lang, Leon, et al.
Published: (2024)
HDPO: Hybrid Distillation Policy Optimization via Privileged Self-Distillation
by: Ding, Ken
Published: (2026)
by: Ding, Ken
Published: (2026)
Chemical Reaction Networks Learn Better than Spiking Neural Networks
by: Jaffard, Sophie, et al.
Published: (2026)
by: Jaffard, Sophie, et al.
Published: (2026)
Robotouille: An Asynchronous Planning Benchmark for LLM Agents
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2025)
by: Gonzalez-Pumariega, Gonzalo, et al.
Published: (2025)
From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review
by: Ferrag, Mohamed Amine, et al.
Published: (2025)
by: Ferrag, Mohamed Amine, et al.
Published: (2025)
DUET: Distilled LLM Unlearning from an Efficiently Contextualized Teacher
by: Zhong, Yisheng, et al.
Published: (2026)
by: Zhong, Yisheng, et al.
Published: (2026)
Repeat After Me: Transformers are Better than State Space Models at Copying
by: Jelassi, Samy, et al.
Published: (2024)
by: Jelassi, Samy, et al.
Published: (2024)
Accelerating Inverse Reinforcement Learning with Expert Bootstrapping
by: Wu, David, et al.
Published: (2024)
by: Wu, David, et al.
Published: (2024)
Aligning LLMs with Domain Invariant Reward Models
by: Wu, David, et al.
Published: (2025)
by: Wu, David, et al.
Published: (2025)
MallowsPO: Fine-Tune Your LLM with Preference Dispersions
by: Chen, Haoxian, et al.
Published: (2024)
by: Chen, Haoxian, et al.
Published: (2024)
Similar Items
-
Process Reward Models for LLM Agents: Practical Framework and Directions
by: Choudhury, Sanjiban
Published: (2025) -
Efficient Imitation under Misspecification
by: Espinosa-Dice, Nicolas, et al.
Published: (2025) -
Motion Tracks: A Unified Representation for Human-Robot Transfer in Few-Shot Imitation Learning
by: Ren, Juntao, et al.
Published: (2025) -
InteRACT: Transformer Models for Human Intent Prediction Conditioned on Robot Actions
by: Kedia, Kushal, et al.
Published: (2023) -
Hybrid Inverse Reinforcement Learning
by: Ren, Juntao, et al.
Published: (2024)