BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
Fuente:
arXiv
Guardado en:
| Autores principales: | Paglieri, Davide, Cupiał, Bartłomiej, Coward, Samuel, Piterbarg, Ulyana, Wolczyk, Maciej, Khan, Akbir, Pignatelli, Eduardo, Kuciński, Łukasz, Pinto, Lerrel, Fergus, Rob, Foerster, Jakob Nicolaus, Parker-Holder, Jack, Rocktäschel, Tim |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Learning When to Plan: Efficiently Allocating Test-Time Compute for LLM Agents
por: Paglieri, Davide, et al.
Publicado: (2025)
por: Paglieri, Davide, et al.
Publicado: (2025)
diff History for Neural Language Agents
por: Piterbarg, Ulyana, et al.
Publicado: (2023)
por: Piterbarg, Ulyana, et al.
Publicado: (2023)
Training Language Models on Synthetic Edit Sequences Improves Code Synthesis
por: Piterbarg, Ulyana, et al.
Publicado: (2024)
por: Piterbarg, Ulyana, et al.
Publicado: (2024)
Outliers and Calibration Sets have Diminishing Effect on Quantization of Modern LLMs
por: Paglieri, Davide, et al.
Publicado: (2024)
por: Paglieri, Davide, et al.
Publicado: (2024)
Multi-Agent Diagnostics for Robustness via Illuminated Diversity
por: Samvelyan, Mikayel, et al.
Publicado: (2024)
por: Samvelyan, Mikayel, et al.
Publicado: (2024)
Fine-tuning Reinforcement Learning Models is Secretly a Forgetting Mitigation Problem
por: Wołczyk, Maciej, et al.
Publicado: (2024)
por: Wołczyk, Maciej, et al.
Publicado: (2024)
Preference-Based Alignment of Discrete Diffusion Models
por: Borso, Umberto, et al.
Publicado: (2025)
por: Borso, Umberto, et al.
Publicado: (2025)
Assessing the Zero-Shot Capabilities of LLMs for Action Evaluation in RL
por: Pignatelli, Eduardo, et al.
Publicado: (2024)
por: Pignatelli, Eduardo, et al.
Publicado: (2024)
Programming by Backprop: An Instruction is Worth 100 Examples When Finetuning LLMs
por: Cook, Jonathan, et al.
Publicado: (2025)
por: Cook, Jonathan, et al.
Publicado: (2025)
DéjàQ: Open-Ended Evolution of Diverse, Learnable and Verifiable Problems
por: Röpke, Willem, et al.
Publicado: (2026)
por: Röpke, Willem, et al.
Publicado: (2026)
Imagined Autocurricula
por: Güzel, Ahmet H., et al.
Publicado: (2025)
por: Güzel, Ahmet H., et al.
Publicado: (2025)
Scaling Opponent Shaping to High Dimensional Games
por: Khan, Akbir, et al.
Publicado: (2023)
por: Khan, Akbir, et al.
Publicado: (2023)
JaxUED: A simple and useable UED library in Jax
por: Coward, Samuel, et al.
Publicado: (2024)
por: Coward, Samuel, et al.
Publicado: (2024)
Learning Multi-Agent Coordination via Sheaf-ADMM
por: Seely, Jeffrey, et al.
Publicado: (2026)
por: Seely, Jeffrey, et al.
Publicado: (2026)
GUIDE: Guidance-based Incremental Learning with Diffusion Models
por: Cywiński, Bartosz, et al.
Publicado: (2024)
por: Cywiński, Bartosz, et al.
Publicado: (2024)
TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and Generation
por: Cook, Jonathan, et al.
Publicado: (2024)
por: Cook, Jonathan, et al.
Publicado: (2024)
There and Back Again: On the relation between Noise and Image Inversions in Diffusion Models
por: Staniszewski, Łukasz, et al.
Publicado: (2024)
por: Staniszewski, Łukasz, et al.
Publicado: (2024)
Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts
por: Samvelyan, Mikayel, et al.
Publicado: (2024)
por: Samvelyan, Mikayel, et al.
Publicado: (2024)
Open-Endedness is Essential for Artificial Superhuman Intelligence
por: Hughes, Edward, et al.
Publicado: (2024)
por: Hughes, Edward, et al.
Publicado: (2024)
Kendall-Cancer-Lab/PAX3-FOXO1_invivo_6hpf: code for publication
por: Jack Kucinski, et al.
Publicado: (2025)
por: Jack Kucinski, et al.
Publicado: (2025)
Factorio Learning Environment
por: Hopkins, Jack, et al.
Publicado: (2025)
por: Hopkins, Jack, et al.
Publicado: (2025)
Seeing Through Their Eyes: Evaluating Visual Perspective Taking in Vision Language Models
por: Góral, Gracjan, et al.
Publicado: (2024)
por: Góral, Gracjan, et al.
Publicado: (2024)
JaxLife: An Open-Ended Agentic Simulator
por: Lu, Chris, et al.
Publicado: (2024)
por: Lu, Chris, et al.
Publicado: (2024)
State Soup: In-Context Skill Learning, Retrieval and Mixing
por: Pióro, Maciej, et al.
Publicado: (2024)
por: Pióro, Maciej, et al.
Publicado: (2024)
An evaluation of the BALROG and RoboBA algorithms for determining the position of Fermi/GBM GRBs
por: López, K. Océlotl. C., et al.
Publicado: (2024)
por: López, K. Océlotl. C., et al.
Publicado: (2024)
Synthetic Data is Sufficient for Zero-Shot Visual Generalization from Offline Data
por: Güzel, Ahmet H., et al.
Publicado: (2025)
por: Güzel, Ahmet H., et al.
Publicado: (2025)
Low regularity potentials in heterogeneous Cahn--Hilliard functionals
por: Cristoferi, Riccardo, et al.
Publicado: (2026)
por: Cristoferi, Riccardo, et al.
Publicado: (2026)
Stochastic Video Generation with a Learned Prior
por: Denton, Remi, et al.
Publicado: (2018)
por: Denton, Remi, et al.
Publicado: (2018)
Catalytic Role Of Noise And Necessity Of Inductive Biases In The Emergence Of Compositional Communication
por: Kuciński, Łukasz, et al.
Publicado: (2021)
por: Kuciński, Łukasz, et al.
Publicado: (2021)
Craftax: A Lightning-Fast Benchmark for Open-Ended Reinforcement Learning
por: Matthews, Michael, et al.
Publicado: (2024)
por: Matthews, Michael, et al.
Publicado: (2024)
Refining Minimax Regret for Unsupervised Environment Design
por: Beukman, Michael, et al.
Publicado: (2024)
por: Beukman, Michael, et al.
Publicado: (2024)
Jornalismo, saúde e cidadania
por: Bernardo Kucinski
Publicado: (2000)
por: Bernardo Kucinski
Publicado: (2000)
Synthesis of Alkynylsilanes: A Review of the State of the Art
por: Krzysztof Kuciński
Publicado: (2024)
por: Krzysztof Kuciński
Publicado: (2024)
Simple formulas of π in terms of ϕ
por: Pignatelli, Angelo
Publicado: (2024)
por: Pignatelli, Angelo
Publicado: (2024)
On canonical threefolds near the Noether line
por: Pignatelli, Roberto
Publicado: (2024)
por: Pignatelli, Roberto
Publicado: (2024)
Entre o idealizado e o exequível: memórias da primeira mulher antropóloga portuguesa. Entrevista com Maria Beatriz Rocha-Trindade
por: Marina Pignatelli
Publicado: (2014)
por: Marina Pignatelli
Publicado: (2014)
Una aproximación empírica al análisis de las percepciones del consumidor sobre el envase
por: Paola Pignatelli
Publicado: (2020)
por: Paola Pignatelli
Publicado: (2020)
Antropologia em Portugal nos últimos 50 anos: introdução
por: Marina Pignatelli
Publicado: (2014)
por: Marina Pignatelli
Publicado: (2014)
RapidDock: Unlocking Proteome-scale Molecular Docking
por: Powalski, Rafał, et al.
Publicado: (2024)
por: Powalski, Rafał, et al.
Publicado: (2024)
AI & Human Co-Improvement for Safer Co-Superintelligence
por: Weston, Jason, et al.
Publicado: (2025)
por: Weston, Jason, et al.
Publicado: (2025)
Ejemplares similares
-
Learning When to Plan: Efficiently Allocating Test-Time Compute for LLM Agents
por: Paglieri, Davide, et al.
Publicado: (2025) -
diff History for Neural Language Agents
por: Piterbarg, Ulyana, et al.
Publicado: (2023) -
Training Language Models on Synthetic Edit Sequences Improves Code Synthesis
por: Piterbarg, Ulyana, et al.
Publicado: (2024) -
Outliers and Calibration Sets have Diminishing Effect on Quantization of Modern LLMs
por: Paglieri, Davide, et al.
Publicado: (2024) -
Multi-Agent Diagnostics for Robustness via Illuminated Diversity
por: Samvelyan, Mikayel, et al.
Publicado: (2024)