Assessing the Zero-Shot Capabilities of LLMs for Action Evaluation in RL
Fuente:
arXiv
Guardado en:
| Autores principales: | Pignatelli, Eduardo, Ferret, Johan, Rockäschel, Tim, Grefenstette, Edward, Paglieri, Davide, Coward, Samuel, Toni, Laura |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A Survey of Temporal Credit Assignment in Deep Reinforcement Learning
por: Pignatelli, Eduardo, et al.
Publicado: (2023)
por: Pignatelli, Eduardo, et al.
Publicado: (2023)
The impact of intrinsic rewards on exploration in Reinforcement Learning
por: Kayal, Aya, et al.
Publicado: (2025)
por: Kayal, Aya, et al.
Publicado: (2025)
Outliers and Calibration Sets have Diminishing Effect on Quantization of Modern LLMs
por: Paglieri, Davide, et al.
Publicado: (2024)
por: Paglieri, Davide, et al.
Publicado: (2024)
Preference-Based Alignment of Discrete Diffusion Models
por: Borso, Umberto, et al.
Publicado: (2025)
por: Borso, Umberto, et al.
Publicado: (2025)
Interaction Dynamics as a Reward Signal for LLMs
por: Gooding, Sian, et al.
Publicado: (2025)
por: Gooding, Sian, et al.
Publicado: (2025)
NAVIX: Scaling MiniGrid Environments with JAX
por: Pignatelli, Eduardo, et al.
Publicado: (2024)
por: Pignatelli, Eduardo, et al.
Publicado: (2024)
minimax: Efficient Baselines for Autocurricula in JAX
por: Jiang, Minqi, et al.
Publicado: (2023)
por: Jiang, Minqi, et al.
Publicado: (2023)
BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
por: Paglieri, Davide, et al.
Publicado: (2024)
por: Paglieri, Davide, et al.
Publicado: (2024)
Zero-Shot LLMs in Human-in-the-Loop RL: Replacing Human Feedback for Reward Shaping
por: Nazir, Mohammad Saif, et al.
Publicado: (2025)
por: Nazir, Mohammad Saif, et al.
Publicado: (2025)
Multi-Agent Diagnostics for Robustness via Illuminated Diversity
por: Samvelyan, Mikayel, et al.
Publicado: (2024)
por: Samvelyan, Mikayel, et al.
Publicado: (2024)
Learning When to Plan: Efficiently Allocating Test-Time Compute for LLM Agents
por: Paglieri, Davide, et al.
Publicado: (2025)
por: Paglieri, Davide, et al.
Publicado: (2025)
Graph-R1: Incentivizing the Zero-Shot Graph Learning Capability in LLMs via Explicit Reasoning
por: Wu, Yicong, et al.
Publicado: (2025)
por: Wu, Yicong, et al.
Publicado: (2025)
Zero-Shot Instruction Following in RL via Structured LTL Representations
por: Giuri, Mattia, et al.
Publicado: (2025)
por: Giuri, Mattia, et al.
Publicado: (2025)
Zero-Shot Instruction Following in RL via Structured LTL Representations
por: Jackermeier, Mathias, et al.
Publicado: (2026)
por: Jackermeier, Mathias, et al.
Publicado: (2026)
JaxUED: A simple and useable UED library in Jax
por: Coward, Samuel, et al.
Publicado: (2024)
por: Coward, Samuel, et al.
Publicado: (2024)
Grounding LTL Tasks in Sub-Symbolic RL Environments for Zero-Shot Generalization
por: Pannacci, Matteo, et al.
Publicado: (2026)
por: Pannacci, Matteo, et al.
Publicado: (2026)
CoBA-RL: Capability-Oriented Budget Allocation for Reinforcement Learning in LLMs
por: Yao, Zhiyuan, et al.
Publicado: (2026)
por: Yao, Zhiyuan, et al.
Publicado: (2026)
Is PRM Necessary? Problem-Solving RL Implicitly Induces PRM Capability in LLMs
por: Feng, Zhangying, et al.
Publicado: (2025)
por: Feng, Zhangying, et al.
Publicado: (2025)
Know Your Neighborhood: General and Zero-Shot Capable Binary Function Search Powered by Call Graphlets
por: Collyer, Joshua, et al.
Publicado: (2024)
por: Collyer, Joshua, et al.
Publicado: (2024)
Drag-and-Drop LLMs: Zero-Shot Prompt-to-Weights
por: Liang, Zhiyuan, et al.
Publicado: (2025)
por: Liang, Zhiyuan, et al.
Publicado: (2025)
DéjàQ: Open-Ended Evolution of Diverse, Learnable and Verifiable Problems
por: Röpke, Willem, et al.
Publicado: (2026)
por: Röpke, Willem, et al.
Publicado: (2026)
A Comedy of Estimators: On KL Regularization in RL Training of LLMs
por: Shah, Vedant, et al.
Publicado: (2025)
por: Shah, Vedant, et al.
Publicado: (2025)
Zero-Shot Robustification of Zero-Shot Models
por: Adila, Dyah, et al.
Publicado: (2023)
por: Adila, Dyah, et al.
Publicado: (2023)
Mixture of Experts in a Mixture of RL settings
por: Willi, Timon, et al.
Publicado: (2024)
por: Willi, Timon, et al.
Publicado: (2024)
Tool Zero: Training Tool-Augmented LLMs via Pure RL from Scratch
por: Zeng, Yirong, et al.
Publicado: (2025)
por: Zeng, Yirong, et al.
Publicado: (2025)
Zero-Shot Conditioning of Score-Based Diffusion Models by Neuro-Symbolic Constraints
por: Scassola, Davide, et al.
Publicado: (2023)
por: Scassola, Davide, et al.
Publicado: (2023)
Univariate to Multivariate: LLMs as Zero-Shot Predictors for Time-Series Forecasting
por: Madarasingha, Chamara, et al.
Publicado: (2025)
por: Madarasingha, Chamara, et al.
Publicado: (2025)
Zero-Shot Action Generalization with Limited Observations
por: Alchihabi, Abdullah, et al.
Publicado: (2025)
por: Alchihabi, Abdullah, et al.
Publicado: (2025)
Evaluating LLMs Capabilities Towards Understanding Social Dynamics
por: Tahir, Anique, et al.
Publicado: (2024)
por: Tahir, Anique, et al.
Publicado: (2024)
LLMs are not Zero-Shot Reasoners for Biomedical Information Extraction
por: Nagar, Aishik, et al.
Publicado: (2024)
por: Nagar, Aishik, et al.
Publicado: (2024)
Scaling In-Context Online Learning Capability of LLMs via Cross-Episode Meta-RL
por: Lin, Xiaofeng, et al.
Publicado: (2026)
por: Lin, Xiaofeng, et al.
Publicado: (2026)
MultiCast: Zero-Shot Multivariate Time Series Forecasting Using LLMs
por: Chatzigeorgakidis, Georgios, et al.
Publicado: (2024)
por: Chatzigeorgakidis, Georgios, et al.
Publicado: (2024)
On Evaluating LLMs' Capabilities as Functional Approximators: A Bayesian Perspective
por: Siddiqui, Shoaib Ahmed, et al.
Publicado: (2024)
por: Siddiqui, Shoaib Ahmed, et al.
Publicado: (2024)
Model Merging Improves Zero-Shot Generalization in Bioacoustic Foundation Models
por: Marincione, Davide, et al.
Publicado: (2025)
por: Marincione, Davide, et al.
Publicado: (2025)
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization
por: Dong, Yihong, et al.
Publicado: (2025)
por: Dong, Yihong, et al.
Publicado: (2025)
Don't flatten, tokenize! Unlocking the key to SoftMoE's efficacy in deep RL
por: Sokar, Ghada, et al.
Publicado: (2024)
por: Sokar, Ghada, et al.
Publicado: (2024)
A Subgoal-driven Framework for Improving Long-Horizon LLM Agents
por: Wang, Taiyi, et al.
Publicado: (2026)
por: Wang, Taiyi, et al.
Publicado: (2026)
Zero-Shot Generalization of Vision-Based RL Without Data Augmentation
por: Batra, Sumeet, et al.
Publicado: (2024)
por: Batra, Sumeet, et al.
Publicado: (2024)
EVIL: Evolving Interpretable Algorithms for Zero-Shot Inference on Event Sequences and Time Series with LLMs
por: Berghaus, David
Publicado: (2026)
por: Berghaus, David
Publicado: (2026)
On Zero-Shot Reinforcement Learning
por: Jeen, Scott
Publicado: (2025)
por: Jeen, Scott
Publicado: (2025)
Ejemplares similares
-
A Survey of Temporal Credit Assignment in Deep Reinforcement Learning
por: Pignatelli, Eduardo, et al.
Publicado: (2023) -
The impact of intrinsic rewards on exploration in Reinforcement Learning
por: Kayal, Aya, et al.
Publicado: (2025) -
Outliers and Calibration Sets have Diminishing Effect on Quantization of Modern LLMs
por: Paglieri, Davide, et al.
Publicado: (2024) -
Preference-Based Alignment of Discrete Diffusion Models
por: Borso, Umberto, et al.
Publicado: (2025) -
Interaction Dynamics as a Reward Signal for LLMs
por: Gooding, Sian, et al.
Publicado: (2025)