Memory Gym: Towards Endless Tasks to Benchmark Memory Capabilities of Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Pleines, Marco, Pallasch, Matthias, Zimmer, Frank, Preuss, Mike |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Pokemon Red via Reinforcement Learning
by: Pleines, Marco, et al.
Published: (2025)
by: Pleines, Marco, et al.
Published: (2025)
Endless Terminals: Scaling RL Environments for Terminal Agents
by: Gandhi, Kanishk, et al.
Published: (2026)
by: Gandhi, Kanishk, et al.
Published: (2026)
CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale
by: Wang, Zhun, et al.
Published: (2025)
by: Wang, Zhun, et al.
Published: (2025)
Hybrid-Gym: Training Coding Agents to Generalize Across Tasks
by: Xie, Yiqing, et al.
Published: (2026)
by: Xie, Yiqing, et al.
Published: (2026)
AExGym: Benchmarks and Environments for Adaptive Experimentation
by: Wang, Jimmy, et al.
Published: (2024)
by: Wang, Jimmy, et al.
Published: (2024)
CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents
by: Wang, Bowen, et al.
Published: (2026)
by: Wang, Bowen, et al.
Published: (2026)
Autoformalizing Memory Specifications with Agents
by: Ernst, Jan Ole, et al.
Published: (2026)
by: Ernst, Jan Ole, et al.
Published: (2026)
Memory, Benchmark & Robots: A Benchmark for Solving Complex Tasks with Reinforcement Learning
by: Cherepanov, Egor, et al.
Published: (2025)
by: Cherepanov, Egor, et al.
Published: (2025)
SeekerGym: A Benchmark for Reliable Information Seeking
by: Kim, Remy, et al.
Published: (2026)
by: Kim, Remy, et al.
Published: (2026)
WebGym: Scaling Training Environments for Visual Web Agents with Realistic Tasks
by: Bai, Hao, et al.
Published: (2026)
by: Bai, Hao, et al.
Published: (2026)
Mastering Memory Tasks with World Models
by: Samsami, Mohammad Reza, et al.
Published: (2024)
by: Samsami, Mohammad Reza, et al.
Published: (2024)
A Benchmark for Procedural Memory Retrieval in Language Agents
by: Kohar, Ishant, et al.
Published: (2025)
by: Kohar, Ishant, et al.
Published: (2025)
Memory Retention Is Not Enough to Master Memory Tasks in Reinforcement Learning
by: Shchendrigin, Oleg, et al.
Published: (2026)
by: Shchendrigin, Oleg, et al.
Published: (2026)
Chain-of-Memory: Lightweight Memory Construction with Dynamic Evolution for LLM Agents
by: Xu, Xiucheng, et al.
Published: (2026)
by: Xu, Xiucheng, et al.
Published: (2026)
Stable Hadamard Memory: Revitalizing Memory-Augmented Agents for Reinforcement Learning
by: Le, Hung, et al.
Published: (2024)
by: Le, Hung, et al.
Published: (2024)
InnoGym: Benchmarking the Innovation Potential of AI Agents
by: Zhang, Jintian, et al.
Published: (2025)
by: Zhang, Jintian, et al.
Published: (2025)
MoralityGym: A Benchmark for Evaluating Hierarchical Moral Alignment in Sequential Decision-Making Agents
by: Rosen, Simon, et al.
Published: (2026)
by: Rosen, Simon, et al.
Published: (2026)
TimeSeriesGym: A Scalable Benchmark for (Time Series) Machine Learning Engineering Agents
by: Cai, Yifu, et al.
Published: (2025)
by: Cai, Yifu, et al.
Published: (2025)
PARMESAN: Parameter-Free Memory Search and Transduction for Dense Prediction Tasks
by: Winter, Philip Matthias, et al.
Published: (2024)
by: Winter, Philip Matthias, et al.
Published: (2024)
The BrowserGym Ecosystem for Web Agent Research
by: De Chezelles, Thibault Le Sellier, et al.
Published: (2024)
by: De Chezelles, Thibault Le Sellier, et al.
Published: (2024)
Memori: A Persistent Memory Layer for Efficient, Context-Aware LLM Agents
by: Borro, Luiz C., et al.
Published: (2026)
by: Borro, Luiz C., et al.
Published: (2026)
OceanGym: A Benchmark Environment for Underwater Embodied Agents
by: Xue, Yida, et al.
Published: (2025)
by: Xue, Yida, et al.
Published: (2025)
BlenderGym: Benchmarking Foundational Model Systems for Graphics Editing
by: Gu, Yunqi, et al.
Published: (2025)
by: Gu, Yunqi, et al.
Published: (2025)
TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking
by: Cheng, Yu, et al.
Published: (2026)
by: Cheng, Yu, et al.
Published: (2026)
Toward Conversational Agents with Context and Time Sensitive Long-term Memory
by: Alonso, Nick, et al.
Published: (2024)
by: Alonso, Nick, et al.
Published: (2024)
Towards Compressive and Scalable Recurrent Memory
by: Song, Yunchong, et al.
Published: (2026)
by: Song, Yunchong, et al.
Published: (2026)
Gym-Anything: Turn any Software into an Agent Environment
by: Aggarwal, Pranjal, et al.
Published: (2026)
by: Aggarwal, Pranjal, et al.
Published: (2026)
Task Vector Quantization for Memory-Efficient Model Merging
by: Kim, Youngeun, et al.
Published: (2025)
by: Kim, Youngeun, et al.
Published: (2025)
CrystalGym: A New Benchmark for Materials Discovery Using Reinforcement Learning
by: Govindarajan, Prashant, et al.
Published: (2025)
by: Govindarajan, Prashant, et al.
Published: (2025)
PersonaGym: Evaluating Persona Agents and LLMs
by: Samuel, Vinay, et al.
Published: (2024)
by: Samuel, Vinay, et al.
Published: (2024)
Gym4ReaL: A Suite for Benchmarking Real-World Reinforcement Learning
by: Salaorni, Davide, et al.
Published: (2025)
by: Salaorni, Davide, et al.
Published: (2025)
Forgetful but Faithful: A Cognitive Memory Architecture and Benchmark for Privacy-Aware Generative Agents
by: Alqithami, Saad
Published: (2025)
by: Alqithami, Saad
Published: (2025)
Illuminating the Diversity-Fitness Trade-Off in Black-Box Optimization
by: Santoni, Maria Laura, et al.
Published: (2024)
by: Santoni, Maria Laura, et al.
Published: (2024)
PeersimGym: An Environment for Solving the Task Offloading Problem with Reinforcement Learning
by: Metelo, Frederico, et al.
Published: (2024)
by: Metelo, Frederico, et al.
Published: (2024)
MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems
by: Ai, Qingyao, et al.
Published: (2025)
by: Ai, Qingyao, et al.
Published: (2025)
Reservoir Computing for Fast, Simplified Reinforcement Learning on Memory Tasks
by: McKee, Kevin
Published: (2024)
by: McKee, Kevin
Published: (2024)
G-Memory: Tracing Hierarchical Memory for Multi-Agent Systems
by: Zhang, Guibin, et al.
Published: (2025)
by: Zhang, Guibin, et al.
Published: (2025)
Sliding Puzzles Gym: A Scalable Benchmark for State Representation in Visual Reinforcement Learning
by: de Oliveira, Bryan L. M., et al.
Published: (2024)
by: de Oliveira, Bryan L. M., et al.
Published: (2024)
BoxingGym: Benchmarking Progress in Automated Experimental Design and Model Discovery
by: Gandhi, Kanishk, et al.
Published: (2025)
by: Gandhi, Kanishk, et al.
Published: (2025)
NextMem: Towards Latent Factual Memory for LLM-based Agents
by: Zhang, Zeyu, et al.
Published: (2026)
by: Zhang, Zeyu, et al.
Published: (2026)
Similar Items
-
Pokemon Red via Reinforcement Learning
by: Pleines, Marco, et al.
Published: (2025) -
Endless Terminals: Scaling RL Environments for Terminal Agents
by: Gandhi, Kanishk, et al.
Published: (2026) -
CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale
by: Wang, Zhun, et al.
Published: (2025) -
Hybrid-Gym: Training Coding Agents to Generalize Across Tasks
by: Xie, Yiqing, et al.
Published: (2026) -
AExGym: Benchmarks and Environments for Adaptive Experimentation
by: Wang, Jimmy, et al.
Published: (2024)