Saved in:
| Main Authors: | Moussa, Hanane Nour, Li, Yifei, Li, Zhuoyang, Yang, Yankai, Tang, Cheng, Zhang, Tianshu, Ahmed, Nesreen K., Payani, Ali, Chen, Ziru, Sun, Huan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2604.27977 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AutoSDT: Scaling Data-Driven Discovery Tasks Toward Open Co-Scientists
by: Li, Yifei, et al.
Published: (2025)
by: Li, Yifei, et al.
Published: (2025)
Toxicity and Tolerability of the OEPA/COPDAC Regimen in Children and Adolescents With Hodgkin Lymphoma: A Real‐World Experience
by: Hala Bekhit, et al.
Published: (2025)
by: Hala Bekhit, et al.
Published: (2025)
When is Tree Search Useful for LLM Planning? It Depends on the Discriminator
by: Chen, Ziru, et al.
Published: (2024)
by: Chen, Ziru, et al.
Published: (2024)
WorldGym: World Model as An Environment for Policy Evaluation
by: Quevedo, Julian, et al.
Published: (2025)
by: Quevedo, Julian, et al.
Published: (2025)
TableLlama: Towards Open Large Generalist Models for Tables
by: Zhang, Tianshu, et al.
Published: (2023)
by: Zhang, Tianshu, et al.
Published: (2023)
CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents
by: Wang, Bowen, et al.
Published: (2026)
by: Wang, Bowen, et al.
Published: (2026)
SWE-Universe: Scale Real-World Verifiable Environments to Millions
by: Chen, Mouxiang, et al.
Published: (2026)
by: Chen, Mouxiang, et al.
Published: (2026)
UserBench: An Interactive Gym Environment for User-Centric Agents
by: Qian, Cheng, et al.
Published: (2025)
by: Qian, Cheng, et al.
Published: (2025)
Training Software Engineering Agents and Verifiers with SWE-Gym
by: Pan, Jiayi, et al.
Published: (2024)
by: Pan, Jiayi, et al.
Published: (2024)
Deliberate Reasoning in Language Models as Structure-Aware Planning with an Accurate World Model
by: Xiong, Siheng, et al.
Published: (2024)
by: Xiong, Siheng, et al.
Published: (2024)
SearchGym: Bootstrapping Real-World Search Agents via Cost-Effective and High-Fidelity Environment Simulation
by: Zhang, Xichen, et al.
Published: (2026)
by: Zhang, Xichen, et al.
Published: (2026)
Effect of Activation Mode on the Use of Black Slag in Cementitious Materials for Ecological and Sustainable Construction
by: Hanane Zadri, et al.
Published: (2026)
by: Hanane Zadri, et al.
Published: (2026)
ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery
by: Chen, Ziru, et al.
Published: (2024)
by: Chen, Ziru, et al.
Published: (2024)
R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents
by: Jain, Naman, et al.
Published: (2025)
by: Jain, Naman, et al.
Published: (2025)
PC-Gym: Benchmark Environments For Process Control Problems
by: Bloor, Maximilian, et al.
Published: (2024)
by: Bloor, Maximilian, et al.
Published: (2024)
Multilingual Reasoning Gym: Multilingual Scaling of Procedural Reasoning Environments
by: Dobler, Konstantin, et al.
Published: (2026)
by: Dobler, Konstantin, et al.
Published: (2026)
Data-Efficient Contrastive Language-Image Pretraining: Prioritizing Data Quality over Quantity
by: Joshi, Siddharth, et al.
Published: (2024)
by: Joshi, Siddharth, et al.
Published: (2024)
Severe Domain Shift in Skeleton-Based Action Recognition:A Study of Uncertainty Failure in Real-World Gym Environments
by: Khanal, Aaditya, et al.
Published: (2026)
by: Khanal, Aaditya, et al.
Published: (2026)
The Station: An Open-World Environment for AI-Driven Discovery
by: Chung, Stephen, et al.
Published: (2025)
by: Chung, Stephen, et al.
Published: (2025)
The Station: An Open-World Environment for AI-Driven Discovery
by: Chung, Stephen, et al.
Published: (2025)
by: Chung, Stephen, et al.
Published: (2025)
Graffiti and the Discursive Construction of Fitness in Gyms
by: Raymond Echitchi
Published: (2022)
by: Raymond Echitchi
Published: (2022)
BoxingGym: Benchmarking Progress in Automated Experimental Design and Model Discovery
by: Gandhi, Kanishk, et al.
Published: (2025)
by: Gandhi, Kanishk, et al.
Published: (2025)
Agent-ToM: Learning to Monitor Autonomous LLM Agents via Theory-of-Mind Reasoning
by: Ahmed, Nesreen K., et al.
Published: (2026)
by: Ahmed, Nesreen K., et al.
Published: (2026)
DataEnvGym: Data Generation Agents in Teacher Environments with Student Feedback
by: Khan, Zaid, et al.
Published: (2024)
by: Khan, Zaid, et al.
Published: (2024)
CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale
by: Wang, Zhun, et al.
Published: (2025)
by: Wang, Zhun, et al.
Published: (2025)
ResearchGym: Evaluating Language Model Agents on Real-World AI Research
by: Garikaparthi, Aniketh, et al.
Published: (2026)
by: Garikaparthi, Aniketh, et al.
Published: (2026)
EO-Gym: A Multimodal, Interactive Environment for Earth Observation Agents
by: Ma, Sai, et al.
Published: (2026)
by: Ma, Sai, et al.
Published: (2026)
ScholarEval: Research Idea Evaluation Grounded in Literature
by: Moussa, Hanane Nour, et al.
Published: (2025)
by: Moussa, Hanane Nour, et al.
Published: (2025)
Prompt Mining for Language-based Human Mobility Forecasting
by: Xue, Hao, et al.
Published: (2024)
by: Xue, Hao, et al.
Published: (2024)
Online Dynamic Goal Recognition in Gym Environments
by: Matan, Shamir, et al.
Published: (2025)
by: Matan, Shamir, et al.
Published: (2025)
pyRDDLGym: From RDDL to Gym Environments
by: Taitler, Ayal, et al.
Published: (2022)
by: Taitler, Ayal, et al.
Published: (2022)
AExGym: Benchmarks and Environments for Adaptive Experimentation
by: Wang, Jimmy, et al.
Published: (2024)
by: Wang, Jimmy, et al.
Published: (2024)
HistoGym: A Reinforcement Learning Environment for Histopathological Image Analysis
by: Liu, Zhi-Bo, et al.
Published: (2024)
by: Liu, Zhi-Bo, et al.
Published: (2024)
Structure Guided Prompt: Instructing Large Language Model in Multi-Step Reasoning by Exploring Graph Structure of the Text
by: Cheng, Kewei, et al.
Published: (2024)
by: Cheng, Kewei, et al.
Published: (2024)
SciNav: A General Agent Framework for Scientific Coding Tasks
by: Zhang, Tianshu, et al.
Published: (2026)
by: Zhang, Tianshu, et al.
Published: (2026)
Customizing Emotional Support: How Do Individuals Construct and Interact With LLM-Powered Chatbots
by: Zheng, Xi, et al.
Published: (2025)
by: Zheng, Xi, et al.
Published: (2025)
Gym4ReaL: A Suite for Benchmarking Real-World Reinforcement Learning
by: Salaorni, Davide, et al.
Published: (2025)
by: Salaorni, Davide, et al.
Published: (2025)
Design Challenges for Robots in Industrial Applications
by: Mufid, Nesreen
Published: (2024)
by: Mufid, Nesreen
Published: (2024)
Generalization Error Bounds for Learning under Censored Feedback
by: Yang, Yifan, et al.
Published: (2024)
by: Yang, Yifan, et al.
Published: (2024)
Friends in Unexpected Places: Enhancing Local Fairness in Federated Learning through Clustering
by: Yang, Yifan, et al.
Published: (2024)
by: Yang, Yifan, et al.
Published: (2024)
Similar Items
-
AutoSDT: Scaling Data-Driven Discovery Tasks Toward Open Co-Scientists
by: Li, Yifei, et al.
Published: (2025) -
Toxicity and Tolerability of the OEPA/COPDAC Regimen in Children and Adolescents With Hodgkin Lymphoma: A Real‐World Experience
by: Hala Bekhit, et al.
Published: (2025) -
When is Tree Search Useful for LLM Planning? It Depends on the Discriminator
by: Chen, Ziru, et al.
Published: (2024) -
WorldGym: World Model as An Environment for Policy Evaluation
by: Quevedo, Julian, et al.
Published: (2025) -
TableLlama: Towards Open Large Generalist Models for Tables
by: Zhang, Tianshu, et al.
Published: (2023)