Policy Learning with a Language Bottleneck
Fuente:
arXiv
Saved in:
| Main Authors: | Srivastava, Megha, Colas, Cedric, Sadigh, Dorsa, Andreas, Jacob |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning
by: Sarkar, Bidipta, et al.
Published: (2025)
by: Sarkar, Bidipta, et al.
Published: (2025)
Language and Experience: A Computational Model of Social Learning in Complex Tasks
by: Colas, Cédric, et al.
Published: (2025)
by: Colas, Cédric, et al.
Published: (2025)
Vocal Sandbox: Continual Learning and Adaptation for Situated Human-Robot Collaboration
by: Grannen, Jennifer, et al.
Published: (2024)
by: Grannen, Jennifer, et al.
Published: (2024)
ProVox: Personalization and Proactive Planning for Situated Human-Robot Collaboration
by: Grannen, Jennifer, et al.
Published: (2025)
by: Grannen, Jennifer, et al.
Published: (2025)
Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models
by: Karamcheti, Siddharth, et al.
Published: (2024)
by: Karamcheti, Siddharth, et al.
Published: (2024)
EXPO: Stable Reinforcement Learning with Expressive Policies
by: Dong, Perry, et al.
Published: (2025)
by: Dong, Perry, et al.
Published: (2025)
Chain of Code: Reasoning with a Language Model-Augmented Code Emulator
by: Li, Chengshu, et al.
Published: (2023)
by: Li, Chengshu, et al.
Published: (2023)
Imitation Bootstrapped Reinforcement Learning
by: Hu, Hengyuan, et al.
Published: (2023)
by: Hu, Hengyuan, et al.
Published: (2023)
Language Bottleneck Models for Qualitative Knowledge State Modeling
by: Berthon, Antonin, et al.
Published: (2025)
by: Berthon, Antonin, et al.
Published: (2025)
Batch Active Learning of Reward Functions from Human Preferences
by: Bıyık, Erdem, et al.
Published: (2024)
by: Bıyık, Erdem, et al.
Published: (2024)
Lexicon-Level Contrastive Visual-Grounding Improves Language Modeling
by: Zhuang, Chengxu, et al.
Published: (2024)
by: Zhuang, Chengxu, et al.
Published: (2024)
Heterogeneity in Formal Linguistic Competence of Language Models: Is Data the Real Bottleneck?
by: Renduchintala, H S V N S Kowndinya, et al.
Published: (2026)
by: Renduchintala, H S V N S Kowndinya, et al.
Published: (2026)
LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM Safety
by: Yang, Junxiao, et al.
Published: (2026)
by: Yang, Junxiao, et al.
Published: (2026)
ABBEL: LLM Agents Acting through Belief Bottlenecks Expressed in Language
by: Lidayan, Aly, et al.
Published: (2025)
by: Lidayan, Aly, et al.
Published: (2025)
Training Language Models to Explain Their Own Computations
by: Li, Belinda Z., et al.
Published: (2025)
by: Li, Belinda Z., et al.
Published: (2025)
(How) Do Language Models Track State?
by: Li, Belinda Z., et al.
Published: (2025)
by: Li, Belinda Z., et al.
Published: (2025)
Fine-Tuning Language Models with Reward Learning on Policy
by: Lang, Hao, et al.
Published: (2024)
by: Lang, Hao, et al.
Published: (2024)
RT-Affordance: Affordances are Versatile Intermediate Representations for Robot Manipulation
by: Nasiriany, Soroush, et al.
Published: (2024)
by: Nasiriany, Soroush, et al.
Published: (2024)
Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning
by: Zhao, Shiwan, et al.
Published: (2026)
by: Zhao, Shiwan, et al.
Published: (2026)
The Bicameral Model: Bidirectional Hidden-State Coupling Between Parallel Language Models
by: Flamant, Cedric, et al.
Published: (2026)
by: Flamant, Cedric, et al.
Published: (2026)
Toward In-Context Teaching: Adapting Examples to Students' Misconceptions
by: Ross, Alexis, et al.
Published: (2024)
by: Ross, Alexis, et al.
Published: (2024)
Algorithmic Capabilities of Random Transformers
by: Zhong, Ziqian, et al.
Published: (2024)
by: Zhong, Ziqian, et al.
Published: (2024)
Towards Reasoning Ability of Small Language Models
by: Srivastava, Gaurav, et al.
Published: (2025)
by: Srivastava, Gaurav, et al.
Published: (2025)
UniGuard: Towards Universal Safety Guardrails for Jailbreak Attacks on Multimodal Large Language Models
by: Oh, Sejoon, et al.
Published: (2024)
by: Oh, Sejoon, et al.
Published: (2024)
On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
by: Agarwal, Rishabh, et al.
Published: (2023)
by: Agarwal, Rishabh, et al.
Published: (2023)
Adaptive Social Learning via Mode Policy Optimization for Language Agents
by: Wang, Minzheng, et al.
Published: (2025)
by: Wang, Minzheng, et al.
Published: (2025)
Reaching Beyond the Mode: RL for Distributional Reasoning in Language Models
by: Puri, Isha, et al.
Published: (2026)
by: Puri, Isha, et al.
Published: (2026)
Learning How Hard to Think: Input-Adaptive Allocation of LM Computation
by: Damani, Mehul, et al.
Published: (2024)
by: Damani, Mehul, et al.
Published: (2024)
Sample-Efficient Online Learning in LM Agents via Hindsight Trajectory Rewriting
by: Hu, Michael Y., et al.
Published: (2025)
by: Hu, Michael Y., et al.
Published: (2025)
DEBATE, TRAIN, EVOLVE: Self Evolution of Language Model Reasoning
by: Srivastava, Gaurav, et al.
Published: (2025)
by: Srivastava, Gaurav, et al.
Published: (2025)
ELF: Embedded Language Flows
by: Hu, Keya, et al.
Published: (2026)
by: Hu, Keya, et al.
Published: (2026)
A Hitchhiker's Guide to Scaling Law Estimation
by: Choshen, Leshem, et al.
Published: (2024)
by: Choshen, Leshem, et al.
Published: (2024)
How do Language Models Bind Entities in Context?
by: Feng, Jiahai, et al.
Published: (2023)
by: Feng, Jiahai, et al.
Published: (2023)
Self-Improving Language Models for Evolutionary Program Synthesis: A Case Study on ARC-AGI
by: Pourcel, Julien, et al.
Published: (2025)
by: Pourcel, Julien, et al.
Published: (2025)
Enhancing In-Hospital Mortality Prediction Using Multi-Representational Learning with LLM-Generated Expert Summaries
by: Battula, Harshavardhan, et al.
Published: (2024)
by: Battula, Harshavardhan, et al.
Published: (2024)
FASTER: Value-Guided Sampling for Fast RL
by: Dong, Perry, et al.
Published: (2026)
by: Dong, Perry, et al.
Published: (2026)
Diffusion Models are Secretly Exchangeable: Parallelizing DDPMs via Autospeculation
by: Hu, Hengyuan, et al.
Published: (2025)
by: Hu, Hengyuan, et al.
Published: (2025)
What's the Move? Hybrid Imitation Learning via Salient Points
by: Sundaresan, Priya, et al.
Published: (2024)
by: Sundaresan, Priya, et al.
Published: (2024)
Value Augmented Sampling for Language Model Alignment and Personalization
by: Han, Seungwook, et al.
Published: (2024)
by: Han, Seungwook, et al.
Published: (2024)
A Taxonomy for Evaluating Generalist Robot Manipulation Policies
by: Gao, Jensen, et al.
Published: (2025)
by: Gao, Jensen, et al.
Published: (2025)
Similar Items
-
Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning
by: Sarkar, Bidipta, et al.
Published: (2025) -
Language and Experience: A Computational Model of Social Learning in Complex Tasks
by: Colas, Cédric, et al.
Published: (2025) -
Vocal Sandbox: Continual Learning and Adaptation for Situated Human-Robot Collaboration
by: Grannen, Jennifer, et al.
Published: (2024) -
ProVox: Personalization and Proactive Planning for Situated Human-Robot Collaboration
by: Grannen, Jennifer, et al.
Published: (2025) -
Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models
by: Karamcheti, Siddharth, et al.
Published: (2024)