Imitating Language via Scalable Inverse Reinforcement Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Wulfmeier, Markus, Bloesch, Michael, Vieillard, Nino, Ahuja, Arun, Bornschein, Jorg, Huang, Sandy, Sokolov, Artem, Barnes, Matt, Desjardins, Guillaume, Bewley, Alex, Bechtle, Sarah Maria Elisabeth, Springenberg, Jost Tobias, Momchev, Nikola, Bachem, Olivier, Geist, Matthieu, Riedmiller, Martin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Value from Observations: Towards Large-Scale Imitation Learning via Self-Improvement
por: Bloesch, Michael, et al.
Publicado: (2025)
por: Bloesch, Michael, et al.
Publicado: (2025)
On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
por: Agarwal, Rishabh, et al.
Publicado: (2023)
por: Agarwal, Rishabh, et al.
Publicado: (2023)
Offline Actor-Critic Reinforcement Learning Scales to Large Models
por: Springenberg, Jost Tobias, et al.
Publicado: (2024)
por: Springenberg, Jost Tobias, et al.
Publicado: (2024)
Game On: Towards Language Models as RL Experimenters
por: Zhang, Jingwei, et al.
Publicado: (2024)
por: Zhang, Jingwei, et al.
Publicado: (2024)
Learning from negative feedback, or positive feedback or both
por: Abdolmaleki, Abbas, et al.
Publicado: (2024)
por: Abdolmaleki, Abbas, et al.
Publicado: (2024)
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)
por: Qin, Chongli, et al.
Publicado: (2025)
por: Qin, Chongli, et al.
Publicado: (2025)
WARM: On the Benefits of Weight Averaged Reward Models
por: Ramé, Alexandre, et al.
Publicado: (2024)
por: Ramé, Alexandre, et al.
Publicado: (2024)
Massively Scalable Inverse Reinforcement Learning in Google Maps
por: Barnes, Matt, et al.
Publicado: (2023)
por: Barnes, Matt, et al.
Publicado: (2023)
LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities
por: Schmied, Thomas, et al.
Publicado: (2025)
por: Schmied, Thomas, et al.
Publicado: (2025)
Inshore-offshore sedimentation differences resulting from resuspension in the Eastern Basin of Lake Erie
por: Bloesch, J
Publicado: (1978)
por: Bloesch, J
Publicado: (1978)
BOND: Aligning LLMs with Best-of-N Distillation
por: Sessa, Pier Giuseppe, et al.
Publicado: (2024)
por: Sessa, Pier Giuseppe, et al.
Publicado: (2024)
Private sector investment in Marine Protected Areas-Experiences of the Chumbe Island Coral Park in Zanzibar/Tanzania
por: Riedmiller, S.
Publicado: (2003)
por: Riedmiller, S.
Publicado: (2003)
Private Sector Management of Marine Protected Areas: The Chumbe Island Case
por: Riedmiller, S.
Publicado: (2000)
por: Riedmiller, S.
Publicado: (2000)
Learning Equilibria from Data: Provably Efficient Multi-Agent Imitation Learning
por: Freihaut, Till, et al.
Publicado: (2025)
por: Freihaut, Till, et al.
Publicado: (2025)
Nash Learning from Human Feedback
por: Munos, Rémi, et al.
Publicado: (2023)
por: Munos, Rémi, et al.
Publicado: (2023)
WARP: On the Benefits of Weight Averaged Rewarded Policies
por: Ramé, Alexandre, et al.
Publicado: (2024)
por: Ramé, Alexandre, et al.
Publicado: (2024)
Las izquierdas en Guatemala / Dirk Bornschein ; ilustrador, José Manuel Chacón
por: Bornschein, Dirk
por: Bornschein, Dirk
Real-World Fluid Directed Rigid Body Control via Deep Reinforcement Learning
por: Bhardwaj, Mohak, et al.
Publicado: (2024)
por: Bhardwaj, Mohak, et al.
Publicado: (2024)
Risk-seeking conservative policy iteration with agent-state based policies for Dec-POMDPs with guaranteed convergence
por: Sinha, Amit, et al.
Publicado: (2026)
por: Sinha, Amit, et al.
Publicado: (2026)
Solving robust MDPs as a sequence of static RL problems
por: Zouitine, Adil, et al.
Publicado: (2024)
por: Zouitine, Adil, et al.
Publicado: (2024)
Convergence of regularized agent-state-based Q-learning in POMDPs
por: Sinha, Amit, et al.
Publicado: (2025)
por: Sinha, Amit, et al.
Publicado: (2025)
Periodic agent-state based Q-learning for POMDPs
por: Sinha, Amit, et al.
Publicado: (2024)
por: Sinha, Amit, et al.
Publicado: (2024)
Optimal Connectivity from Idle Qubit residual coupling Cross-Talks in a Cavity Mediated Entangling Gate
por: Mammola, Andrea, et al.
Publicado: (2025)
por: Mammola, Andrea, et al.
Publicado: (2025)
Loss Functions and Operators Generated by f-Divergences
por: Roulet, Vincent, et al.
Publicado: (2025)
por: Roulet, Vincent, et al.
Publicado: (2025)
An Embodied Companion for Visual Storytelling
por: Tresset, Patrick, et al.
Publicado: (2026)
por: Tresset, Patrick, et al.
Publicado: (2026)
Commentary on Epidemiology of mental health comorbidity in patients with atopic dermatitis: An analysis of global trends from 1998 to 2022
por: Anthony Bewley
Publicado: (2024)
por: Anthony Bewley
Publicado: (2024)
Towards Minimax Optimality of Model-based Robust Reinforcement Learning
por: Clavier, Pierre, et al.
Publicado: (2023)
por: Clavier, Pierre, et al.
Publicado: (2023)
DRoP: Distributionally Robust Data Pruning
por: Vysogorets, Artem, et al.
Publicado: (2024)
por: Vysogorets, Artem, et al.
Publicado: (2024)
Moderate rank jumps on rational elliptic surfaces via construction of conics
por: Desjardins, Julie
Publicado: (2025)
por: Desjardins, Julie
Publicado: (2025)
Breve tratado de la emoción / Denise Desjardins ; traducción de Borja Folch
por: Desjardins, Denise
por: Desjardins, Denise
Improving Implementation of the Psychosocial Standards of Care Together
por: Leandra Desjardins
Publicado: (2025)
por: Leandra Desjardins
Publicado: (2025)
3D Audio-Visual Segmentation
por: Sokolov, Artem, et al.
Publicado: (2024)
por: Sokolov, Artem, et al.
Publicado: (2024)
Training and Evaluation of Guideline-Based Medical Reasoning in LLMs
por: Staniek, Michael, et al.
Publicado: (2025)
por: Staniek, Michael, et al.
Publicado: (2025)
Cooperative Face Liveness Detection from Optical Flow
por: Sokolov, Artem, et al.
Publicado: (2025)
por: Sokolov, Artem, et al.
Publicado: (2025)
Closing the Gap between TD Learning and Supervised Learning -- A Generalisation Point of View
por: Ghugare, Raj, et al.
Publicado: (2024)
por: Ghugare, Raj, et al.
Publicado: (2024)
Library Schools React to Continuing Education Demands.
por: Bewley, Lois M.
Publicado: (1980)
por: Bewley, Lois M.
Publicado: (1980)
On Teacher Hacking in Language Model Distillation
por: Tiapkin, Daniil, et al.
Publicado: (2025)
por: Tiapkin, Daniil, et al.
Publicado: (2025)
NFQ2.0: The CartPole Benchmark Revisited
por: Lange, Sascha, et al.
Publicado: (2025)
por: Lange, Sascha, et al.
Publicado: (2025)
RL Token: Bootstrapping Online RL with Vision-Language-Action Models
por: Xu, Charles, et al.
Publicado: (2026)
por: Xu, Charles, et al.
Publicado: (2026)
Learning Robot Soccer from Egocentric Vision with Deep Reinforcement Learning
por: Tirumala, Dhruva, et al.
Publicado: (2024)
por: Tirumala, Dhruva, et al.
Publicado: (2024)
Ejemplares similares
-
Value from Observations: Towards Large-Scale Imitation Learning via Self-Improvement
por: Bloesch, Michael, et al.
Publicado: (2025) -
On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
por: Agarwal, Rishabh, et al.
Publicado: (2023) -
Offline Actor-Critic Reinforcement Learning Scales to Large Models
por: Springenberg, Jost Tobias, et al.
Publicado: (2024) -
Game On: Towards Language Models as RL Experimenters
por: Zhang, Jingwei, et al.
Publicado: (2024) -
Learning from negative feedback, or positive feedback or both
por: Abdolmaleki, Abbas, et al.
Publicado: (2024)