CUBE: A Standard for Unifying Agent Benchmarks
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866915868154265600 |
|---|---|
| author | Lacoste, Alexandre Gontier, Nicolas Shliazhko, Oleh Jaiswal, Aman Sareen, Kusha Nanisetty, Shailesh Cabezas, Joan Del Verme, Manuel Younis, Omar G. Baratta, Simone Avalle, Matteo Kerboua, Imene Lù, Xing Han Bandel, Elron Shmueli-Scheuer, Michal Yehudai, Asaf Choshen, Leshem Lebensold, Jonathan Hughes, Sean Caccia, Massimo Drouin, Alexandre Reddy, Siva Yu, Tao Su, Yu Neubig, Graham Song, Dawn |
| author_facet | Lacoste, Alexandre Gontier, Nicolas Shliazhko, Oleh Jaiswal, Aman Sareen, Kusha Nanisetty, Shailesh Cabezas, Joan Del Verme, Manuel Younis, Omar G. Baratta, Simone Avalle, Matteo Kerboua, Imene Lù, Xing Han Bandel, Elron Shmueli-Scheuer, Michal Yehudai, Asaf Choshen, Leshem Lebensold, Jonathan Hughes, Sean Caccia, Massimo Drouin, Alexandre Reddy, Siva Yu, Tao Su, Yu Neubig, Graham Song, Dawn |
| contents | The proliferation of agent benchmarks has created critical fragmentation that threatens research productivity. Each new benchmark requires substantial custom integration, creating an "integration tax" that limits comprehensive evaluation. We propose CUBE (Common Unified Benchmark Environments), a universal protocol standard built on MCP and Gym that allows benchmarks to be wrapped once and used everywhere. By separating task, benchmark, package, and registry concerns into distinct API layers, CUBE enables any compliant platform to access any compliant benchmark for evaluation, RL training, or data generation without custom integration. We call on the community to contribute to the development of this standard before platform-specific implementations deepen fragmentation as benchmark production accelerates through 2026. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_15798 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | CUBE: A Standard for Unifying Agent Benchmarks Lacoste, Alexandre Gontier, Nicolas Shliazhko, Oleh Jaiswal, Aman Sareen, Kusha Nanisetty, Shailesh Cabezas, Joan Del Verme, Manuel Younis, Omar G. Baratta, Simone Avalle, Matteo Kerboua, Imene Lù, Xing Han Bandel, Elron Shmueli-Scheuer, Michal Yehudai, Asaf Choshen, Leshem Lebensold, Jonathan Hughes, Sean Caccia, Massimo Drouin, Alexandre Reddy, Siva Yu, Tao Su, Yu Neubig, Graham Song, Dawn Artificial Intelligence The proliferation of agent benchmarks has created critical fragmentation that threatens research productivity. Each new benchmark requires substantial custom integration, creating an "integration tax" that limits comprehensive evaluation. We propose CUBE (Common Unified Benchmark Environments), a universal protocol standard built on MCP and Gym that allows benchmarks to be wrapped once and used everywhere. By separating task, benchmark, package, and registry concerns into distinct API layers, CUBE enables any compliant platform to access any compliant benchmark for evaluation, RL training, or data generation without custom integration. We call on the community to contribute to the development of this standard before platform-specific implementations deepen fragmentation as benchmark production accelerates through 2026. |
| title | CUBE: A Standard for Unifying Agent Benchmarks |
| topic | Artificial Intelligence |
| url | https://arxiv.org/abs/2603.15798 |