Imitation Learning for Multi-turn LM Agents via On-policy Expert Corrections
Fuente:
arXiv
Guardado en:
| Autores principales: | Lauffer, Niklas, Deng, Xiang, Kundurthy, Srivatsa, Kenstler, Brad, Da, Jeff |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Learning Symbolic Task Decompositions for Multi-Agent Teams
por: Shah, Ameesh, et al.
Publicado: (2025)
por: Shah, Ameesh, et al.
Publicado: (2025)
Provably Correct Automata Embeddings for Optimal Automata-Conditioned Reinforcement Learning
por: Yalcinkaya, Beyazit, et al.
Publicado: (2025)
por: Yalcinkaya, Beyazit, et al.
Publicado: (2025)
Learning with Expert Abstractions for Efficient Multi-Task Continuous Control
por: Jewett, Jeff, et al.
Publicado: (2025)
por: Jewett, Jeff, et al.
Publicado: (2025)
Causal Imitation Learning under Expert-Observable and Expert-Unobservable Confounding
por: Shao, Daqian, et al.
Publicado: (2025)
por: Shao, Daqian, et al.
Publicado: (2025)
Thought Cloning: Learning to Think while Acting by Imitating Human Thinking
por: Hu, Shengran, et al.
Publicado: (2023)
por: Hu, Shengran, et al.
Publicado: (2023)
SpreadsheetArena: Decomposing Preference in LLM Generation of Spreadsheet Workbooks
por: Kundurthy, Srivatsa, et al.
Publicado: (2026)
por: Kundurthy, Srivatsa, et al.
Publicado: (2026)
Using Non-Expert Data to Robustify Imitation Learning via Offline Reinforcement Learning
por: Huang, Kevin, et al.
Publicado: (2025)
por: Huang, Kevin, et al.
Publicado: (2025)
IDIL: Imitation Learning of Intent-Driven Expert Behavior
por: Seo, Sangwon, et al.
Publicado: (2024)
por: Seo, Sangwon, et al.
Publicado: (2024)
Sample-Efficient Expert Query Control in Active Imitation Learning via Conformal Prediction
por: Firouzkouhi, Arad, et al.
Publicado: (2025)
por: Firouzkouhi, Arad, et al.
Publicado: (2025)
Reason in Chains, Learn in Trees: Self-Rectification and Grafting for Multi-turn Agent Policy Optimization
por: Li, Yu, et al.
Publicado: (2026)
por: Li, Yu, et al.
Publicado: (2026)
Learning Strategy Representation for Imitation Learning in Multi-Agent Games
por: Lei, Shiqi, et al.
Publicado: (2024)
por: Lei, Shiqi, et al.
Publicado: (2024)
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration
por: Zhao, Heyang, et al.
Publicado: (2025)
por: Zhao, Heyang, et al.
Publicado: (2025)
Coupled Distributional Random Expert Distillation for World Model Online Imitation Learning
por: Li, Shangzhe, et al.
Publicado: (2025)
por: Li, Shangzhe, et al.
Publicado: (2025)
Compositional Automata Embeddings for Goal-Conditioned Reinforcement Learning
por: Yalcinkaya, Beyazit, et al.
Publicado: (2024)
por: Yalcinkaya, Beyazit, et al.
Publicado: (2024)
MAPF-GPT: Imitation Learning for Multi-Agent Pathfinding at Scale
por: Andreychuk, Anton, et al.
Publicado: (2024)
por: Andreychuk, Anton, et al.
Publicado: (2024)
SENSOR: Imitate Third-Person Expert's Behaviors via Active Sensoring
por: Huang, Kaichen, et al.
Publicado: (2024)
por: Huang, Kaichen, et al.
Publicado: (2024)
Scaling Laws for Imitation Learning in Single-Agent Games
por: Tuyls, Jens, et al.
Publicado: (2023)
por: Tuyls, Jens, et al.
Publicado: (2023)
TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents
por: Wang, Jiaqi, et al.
Publicado: (2026)
por: Wang, Jiaqi, et al.
Publicado: (2026)
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations
por: Hoang, Huy, et al.
Publicado: (2025)
por: Hoang, Huy, et al.
Publicado: (2025)
Imitation Learning via Focused Satisficing
por: Shah, Rushit N., et al.
Publicado: (2025)
por: Shah, Rushit N., et al.
Publicado: (2025)
Boolean Satisfiability via Imitation Learning
por: Zhang, Zewei, et al.
Publicado: (2025)
por: Zhang, Zewei, et al.
Publicado: (2025)
Adversarial Imitation Learning via Boosting
por: Chang, Jonathan D., et al.
Publicado: (2024)
por: Chang, Jonathan D., et al.
Publicado: (2024)
Identifying the Risks of LM Agents with an LM-Emulated Sandbox
por: Ruan, Yangjun, et al.
Publicado: (2023)
por: Ruan, Yangjun, et al.
Publicado: (2023)
AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
por: Ma, Chang, et al.
Publicado: (2024)
por: Ma, Chang, et al.
Publicado: (2024)
Learning Representations in Video Game Agents with Supervised Contrastive Imitation Learning
por: Celemin, Carlos, et al.
Publicado: (2025)
por: Celemin, Carlos, et al.
Publicado: (2025)
Sample-Efficient Online Learning in LM Agents via Hindsight Trajectory Rewriting
por: Hu, Michael Y., et al.
Publicado: (2025)
por: Hu, Michael Y., et al.
Publicado: (2025)
Reinforcement Learning via Implicit Imitation Guidance
por: Dong, Perry, et al.
Publicado: (2025)
por: Dong, Perry, et al.
Publicado: (2025)
Unveiling the Role of Expert Guidance: A Comparative Analysis of User-centered Imitation Learning and Traditional Reinforcement Learning
por: Gomaa, Amr, et al.
Publicado: (2024)
por: Gomaa, Amr, et al.
Publicado: (2024)
DecompGAIL: Learning Realistic Traffic Behaviors with Decomposed Multi-Agent Generative Adversarial Imitation Learning
por: Guo, Ke, et al.
Publicado: (2025)
por: Guo, Ke, et al.
Publicado: (2025)
Offline Imitation of Badminton Player Behavior via Experiential Contexts and Brownian Motion
por: Wang, Kuang-Da, et al.
Publicado: (2024)
por: Wang, Kuang-Da, et al.
Publicado: (2024)
SpectR: Dynamically Composing LM Experts with Spectral Routing
por: Fleshman, William, et al.
Publicado: (2025)
por: Fleshman, William, et al.
Publicado: (2025)
Imitation Learning Datasets: A Toolkit For Creating Datasets, Training Agents and Benchmarking
por: Gavenski, Nathan, et al.
Publicado: (2024)
por: Gavenski, Nathan, et al.
Publicado: (2024)
Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation
por: Li, Pengxiang, et al.
Publicado: (2025)
por: Li, Pengxiang, et al.
Publicado: (2025)
Learning Soft Driving Constraints from Vectorized Scene Embeddings while Imitating Expert Trajectories
por: Mobarakeh, Niloufar Saeidi, et al.
Publicado: (2024)
por: Mobarakeh, Niloufar Saeidi, et al.
Publicado: (2024)
Error-Feedback Model for Output Correction in Bilateral Control-Based Imitation Learning
por: Sato, Hiroshi, et al.
Publicado: (2024)
por: Sato, Hiroshi, et al.
Publicado: (2024)
Expert-Free Online Transfer Learning in Multi-Agent Reinforcement Learning
por: Castagna, Alberto
Publicado: (2025)
por: Castagna, Alberto
Publicado: (2025)
Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents
por: Wang, Hao, et al.
Publicado: (2026)
por: Wang, Hao, et al.
Publicado: (2026)
Imitation Learning from Suboptimal Demonstrations via Meta-Learning An Action Ranker
por: Fan, Jiangdong, et al.
Publicado: (2024)
por: Fan, Jiangdong, et al.
Publicado: (2024)
Quantifying Generalisation in Imitation Learning
por: Gavenski, Nathan, et al.
Publicado: (2025)
por: Gavenski, Nathan, et al.
Publicado: (2025)
Imitation Bootstrapped Reinforcement Learning
por: Hu, Hengyuan, et al.
Publicado: (2023)
por: Hu, Hengyuan, et al.
Publicado: (2023)
Ejemplares similares
-
Learning Symbolic Task Decompositions for Multi-Agent Teams
por: Shah, Ameesh, et al.
Publicado: (2025) -
Provably Correct Automata Embeddings for Optimal Automata-Conditioned Reinforcement Learning
por: Yalcinkaya, Beyazit, et al.
Publicado: (2025) -
Learning with Expert Abstractions for Efficient Multi-Task Continuous Control
por: Jewett, Jeff, et al.
Publicado: (2025) -
Causal Imitation Learning under Expert-Observable and Expert-Unobservable Confounding
por: Shao, Daqian, et al.
Publicado: (2025) -
Thought Cloning: Learning to Think while Acting by Imitating Human Thinking
por: Hu, Shengran, et al.
Publicado: (2023)