CUBE: A Standard for Unifying Agent Benchmarks

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Lacoste, Alexandre, Gontier, Nicolas, Shliazhko, Oleh, Jaiswal, Aman, Sareen, Kusha, Nanisetty, Shailesh, Cabezas, Joan, Del Verme, Manuel, Younis, Omar G., Baratta, Simone, Avalle, Matteo, Kerboua, Imene, Lù, Xing Han, Bandel, Elron, Shmueli-Scheuer, Michal, Yehudai, Asaf, Choshen, Leshem, Lebensold, Jonathan, Hughes, Sean, Caccia, Massimo, Drouin, Alexandre, Reddy, Siva, Yu, Tao, Su, Yu, Neubig, Graham, Song, Dawn
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915868154265600
author Lacoste, Alexandre
Gontier, Nicolas
Shliazhko, Oleh
Jaiswal, Aman
Sareen, Kusha
Nanisetty, Shailesh
Cabezas, Joan
Del Verme, Manuel
Younis, Omar G.
Baratta, Simone
Avalle, Matteo
Kerboua, Imene
Lù, Xing Han
Bandel, Elron
Shmueli-Scheuer, Michal
Yehudai, Asaf
Choshen, Leshem
Lebensold, Jonathan
Hughes, Sean
Caccia, Massimo
Drouin, Alexandre
Reddy, Siva
Yu, Tao
Su, Yu
Neubig, Graham
Song, Dawn
author_facet Lacoste, Alexandre
Gontier, Nicolas
Shliazhko, Oleh
Jaiswal, Aman
Sareen, Kusha
Nanisetty, Shailesh
Cabezas, Joan
Del Verme, Manuel
Younis, Omar G.
Baratta, Simone
Avalle, Matteo
Kerboua, Imene
Lù, Xing Han
Bandel, Elron
Shmueli-Scheuer, Michal
Yehudai, Asaf
Choshen, Leshem
Lebensold, Jonathan
Hughes, Sean
Caccia, Massimo
Drouin, Alexandre
Reddy, Siva
Yu, Tao
Su, Yu
Neubig, Graham
Song, Dawn
contents The proliferation of agent benchmarks has created critical fragmentation that threatens research productivity. Each new benchmark requires substantial custom integration, creating an "integration tax" that limits comprehensive evaluation. We propose CUBE (Common Unified Benchmark Environments), a universal protocol standard built on MCP and Gym that allows benchmarks to be wrapped once and used everywhere. By separating task, benchmark, package, and registry concerns into distinct API layers, CUBE enables any compliant platform to access any compliant benchmark for evaluation, RL training, or data generation without custom integration. We call on the community to contribute to the development of this standard before platform-specific implementations deepen fragmentation as benchmark production accelerates through 2026.
format Preprint
id arxiv_https___arxiv_org_abs_2603_15798
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CUBE: A Standard for Unifying Agent Benchmarks
Lacoste, Alexandre
Gontier, Nicolas
Shliazhko, Oleh
Jaiswal, Aman
Sareen, Kusha
Nanisetty, Shailesh
Cabezas, Joan
Del Verme, Manuel
Younis, Omar G.
Baratta, Simone
Avalle, Matteo
Kerboua, Imene
Lù, Xing Han
Bandel, Elron
Shmueli-Scheuer, Michal
Yehudai, Asaf
Choshen, Leshem
Lebensold, Jonathan
Hughes, Sean
Caccia, Massimo
Drouin, Alexandre
Reddy, Siva
Yu, Tao
Su, Yu
Neubig, Graham
Song, Dawn
Artificial Intelligence
The proliferation of agent benchmarks has created critical fragmentation that threatens research productivity. Each new benchmark requires substantial custom integration, creating an "integration tax" that limits comprehensive evaluation. We propose CUBE (Common Unified Benchmark Environments), a universal protocol standard built on MCP and Gym that allows benchmarks to be wrapped once and used everywhere. By separating task, benchmark, package, and registry concerns into distinct API layers, CUBE enables any compliant platform to access any compliant benchmark for evaluation, RL training, or data generation without custom integration. We call on the community to contribute to the development of this standard before platform-specific implementations deepen fragmentation as benchmark production accelerates through 2026.
title CUBE: A Standard for Unifying Agent Benchmarks
topic Artificial Intelligence
url https://arxiv.org/abs/2603.15798