Meta-RL Induces Exploration in Language Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Yulun, Jiang, Liangze, Teney, Damien, Moor, Michael, Brbic, Maria |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OOD-Chameleon: Is Algorithm Selection for OOD Generalization Learnable?
by: Jiang, Liangze, et al.
Published: (2024)
by: Jiang, Liangze, et al.
Published: (2024)
MARBLE: A Hard Benchmark for Multimodal Spatial Reasoning and Planning
by: Jiang, Yulun, et al.
Published: (2025)
by: Jiang, Yulun, et al.
Published: (2025)
Do We Always Need the Simplicity Bias? Looking for Optimal Inductive Biases in the Wild
by: Teney, Damien, et al.
Published: (2025)
by: Teney, Damien, et al.
Published: (2025)
Procedural Pretraining: Warming Up Language Models with Abstract Data
by: Jiang, Liangze, et al.
Published: (2026)
by: Jiang, Liangze, et al.
Published: (2026)
Risk-Sensitive RL for Alleviating Exploration Dilemmas in Large Language Models
by: Jiang, Yuhua, et al.
Published: (2025)
by: Jiang, Yuhua, et al.
Published: (2025)
Let Go of Your Labels with Unsupervised Transfer
by: Gadetsky, Artyom, et al.
Published: (2024)
by: Gadetsky, Artyom, et al.
Published: (2024)
Transformers Pretrained on Procedural Data Contain Modular Structures for Algorithmic Reasoning
by: Shinnick, Zachary, et al.
Published: (2025)
by: Shinnick, Zachary, et al.
Published: (2025)
How does the optimizer implicitly bias the model merging loss landscape?
by: Zhang, Chenxiang, et al.
Published: (2025)
by: Zhang, Chenxiang, et al.
Published: (2025)
AgentRxiv: Towards Collaborative Autonomous Research
by: Schmidgall, Samuel, et al.
Published: (2025)
by: Schmidgall, Samuel, et al.
Published: (2025)
Neural Redshift: Random Networks are not Random Functions
by: Teney, Damien, et al.
Published: (2024)
by: Teney, Damien, et al.
Published: (2024)
Weak-to-Strong Generalization under Distribution Shifts
by: Jeon, Myeongho, et al.
Published: (2025)
by: Jeon, Myeongho, et al.
Published: (2025)
Scalable Ensemble Diversification for OOD Generalization and Detection
by: Rubinstein, Alexander, et al.
Published: (2024)
by: Rubinstein, Alexander, et al.
Published: (2024)
Zero-shot Meta-learning for Tabular Prediction Tasks with Adversarially Pre-trained Transformer
by: Wu, Yulun, et al.
Published: (2025)
by: Wu, Yulun, et al.
Published: (2025)
RL$^3$: Boosting Meta Reinforcement Learning via RL inside RL$^2$
by: Bhatia, Abhinav, et al.
Published: (2023)
by: Bhatia, Abhinav, et al.
Published: (2023)
Mitigating Shortcut Learning with Diffusion Counterfactuals and Diverse Ensembles
by: Scimeca, Luca, et al.
Published: (2023)
by: Scimeca, Luca, et al.
Published: (2023)
Meta-World+: An Improved, Standardized, RL Benchmark
by: McLean, Reginald, et al.
Published: (2025)
by: McLean, Reginald, et al.
Published: (2025)
Capacity-Aware Planning and Scheduling in Budget-Constrained Multi-Agent MDPs: A Meta-RL Approach
by: Vora, Manav, et al.
Published: (2024)
by: Vora, Manav, et al.
Published: (2024)
Settling Decentralized Multi-Agent Coordinated Exploration by Novelty Sharing
by: Jiang, Haobin, et al.
Published: (2024)
by: Jiang, Haobin, et al.
Published: (2024)
Toward Efficient Exploration by Large Language Model Agents
by: Arumugam, Dilip, et al.
Published: (2025)
by: Arumugam, Dilip, et al.
Published: (2025)
Sample-efficient and Scalable Exploration in Continuous-Time RL
by: Iten, Klemens, et al.
Published: (2025)
by: Iten, Klemens, et al.
Published: (2025)
With Limited Data for Multimodal Alignment, Let the STRUCTURE Guide You
by: Gröger, Fabian, et al.
Published: (2025)
by: Gröger, Fabian, et al.
Published: (2025)
Explaining RL Decisions with Trajectories
by: Deshmukh, Shripad Vilasrao, et al.
Published: (2023)
by: Deshmukh, Shripad Vilasrao, et al.
Published: (2023)
SLEA-RL: Step-Level Experience Augmented Reinforcement Learning for Multi-Turn Agentic Training
by: Wang, Prince Zizhuang, et al.
Published: (2026)
by: Wang, Prince Zizhuang, et al.
Published: (2026)
Soft-Masked Diffusion Language Models
by: Hersche, Michael, et al.
Published: (2025)
by: Hersche, Michael, et al.
Published: (2025)
LPPG-RL: Lexicographically Projected Policy Gradient Reinforcement Learning with Subproblem Exploration
by: Qiu, Ruiyu, et al.
Published: (2025)
by: Qiu, Ruiyu, et al.
Published: (2025)
Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback
by: Yang, Yongjin, et al.
Published: (2025)
by: Yang, Yongjin, et al.
Published: (2025)
A Single Goal is All You Need: Skills and Exploration Emerge from Contrastive RL without Rewards, Demonstrations, or Subgoals
by: Liu, Grace, et al.
Published: (2024)
by: Liu, Grace, et al.
Published: (2024)
Revisiting the Platonic Representation Hypothesis: An Aristotelian View
by: Gröger, Fabian, et al.
Published: (2026)
by: Gröger, Fabian, et al.
Published: (2026)
Learning Off-policy with Model-based Intrinsic Motivation For Active Online Exploration
by: Wang, Yibo, et al.
Published: (2024)
by: Wang, Yibo, et al.
Published: (2024)
Are Expressive Models Truly Necessary for Offline RL?
by: Wang, Guan, et al.
Published: (2024)
by: Wang, Guan, et al.
Published: (2024)
CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation
by: Dai, Weinan, et al.
Published: (2026)
by: Dai, Weinan, et al.
Published: (2026)
AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench
by: Toledo, Edan, et al.
Published: (2025)
by: Toledo, Edan, et al.
Published: (2025)
SortedRL: Accelerating RL Training for LLMs through Online Length-Aware Scheduling
by: Zhang, Yiqi, et al.
Published: (2026)
by: Zhang, Yiqi, et al.
Published: (2026)
ArenaRL: Scaling RL for Open-Ended Agents via Tournament-based Relative Ranking
by: Zhang, Qiang, et al.
Published: (2026)
by: Zhang, Qiang, et al.
Published: (2026)
ProgAgent:A Continual RL Agent with Progress-Aware Rewards
by: Tan, Jinzhou, et al.
Published: (2026)
by: Tan, Jinzhou, et al.
Published: (2026)
Verification-Guided Falsification for Safe RL via Explainable Abstraction and Risk-Aware Exploration
by: Le, Tuan, et al.
Published: (2025)
by: Le, Tuan, et al.
Published: (2025)
RF-Agent: Automated Reward Function Design via Language Agent Tree Search
by: Gao, Ning, et al.
Published: (2026)
by: Gao, Ning, et al.
Published: (2026)
Inefficiencies of Meta Agents for Agent Design
by: El, Batu, et al.
Published: (2025)
by: El, Batu, et al.
Published: (2025)
MALLM-GAN: Multi-Agent Large Language Model as Generative Adversarial Network for Synthesizing Tabular Data
by: Ling, Yaobin, et al.
Published: (2024)
by: Ling, Yaobin, et al.
Published: (2024)
Words as Beacons: Guiding RL Agents with High-Level Language Prompts
by: Ruiz-Gonzalez, Unai, et al.
Published: (2024)
by: Ruiz-Gonzalez, Unai, et al.
Published: (2024)
Similar Items
-
OOD-Chameleon: Is Algorithm Selection for OOD Generalization Learnable?
by: Jiang, Liangze, et al.
Published: (2024) -
MARBLE: A Hard Benchmark for Multimodal Spatial Reasoning and Planning
by: Jiang, Yulun, et al.
Published: (2025) -
Do We Always Need the Simplicity Bias? Looking for Optimal Inductive Biases in the Wild
by: Teney, Damien, et al.
Published: (2025) -
Procedural Pretraining: Warming Up Language Models with Abstract Data
by: Jiang, Liangze, et al.
Published: (2026) -
Risk-Sensitive RL for Alleviating Exploration Dilemmas in Large Language Models
by: Jiang, Yuhua, et al.
Published: (2025)