Results and Retrospective Analysis of the CODS 2025 AssetOpsBench Challenge

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Patel, Dhaval, Shyalika, Chathurangi, Yarrabothula, Suryanarayana Reddy, Yue, Ling, Lin, Shuxin, Zhou, Nianjun, Rayfield, James
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913104535748608
author Patel, Dhaval
Shyalika, Chathurangi
Yarrabothula, Suryanarayana Reddy
Yue, Ling
Lin, Shuxin
Zhou, Nianjun
Rayfield, James
author_facet Patel, Dhaval
Shyalika, Chathurangi
Yarrabothula, Suryanarayana Reddy
Yue, Ling
Lin, Shuxin
Zhou, Nianjun
Rayfield, James
contents Competition retrospectives are useful when they explain what a leaderboard measured, how hidden evaluation changed conclusions, and which design patterns were rewarded. We revisit the CODS 2025 \assetopslive{} challenge, a privacy-aware Codabench competition on industrial multi-agent orchestration built on \assetops{}. We combine final rank sheets, a 300-submission server log, 149-team registrations, best-submission exports, the organizer winners report, the companion \assetopslive{} system paper, and verified planning-track source trees. Five results stand out. First, the public planning leaderboard saturates at 72.73\%, and richer prompts do not improve that peak. Second, hidden evaluation changes the story: public and private scores correlate moderately in planning ($r{=}0.69$) but negatively in execution ($r{=}{-}0.13$), with several 45.45\% public execution systems reaching 63.64\% on the hidden set. Third, the \tmatch{} term is numerically almost inert in the official composite -- combined on a 0--1 scale with 0--100 percentage scores, it contributes at most 0.05 points per track, and rescaling would swap the top two teams. Fourth, the competition is operationally account-based but substantively team-based: 149 registered teams reduce to 24 with non-zero public scores and 11 fully ranked, while 52.3\% of deduplicated registrations list multiple usernames. Fifth, successful execution methods mostly improve guardrails -- response selection, contamination cleanup, fallback, and context control -- rather than novel agent architectures. These findings identify which behaviors the evaluation rewarded, and motivate scale-aware composites, skill-level diagnostics, and versioned artifact release.
format Preprint
id arxiv_https___arxiv_org_abs_2605_08518
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Results and Retrospective Analysis of the CODS 2025 AssetOpsBench Challenge
Patel, Dhaval
Shyalika, Chathurangi
Yarrabothula, Suryanarayana Reddy
Yue, Ling
Lin, Shuxin
Zhou, Nianjun
Rayfield, James
Artificial Intelligence
Competition retrospectives are useful when they explain what a leaderboard measured, how hidden evaluation changed conclusions, and which design patterns were rewarded. We revisit the CODS 2025 \assetopslive{} challenge, a privacy-aware Codabench competition on industrial multi-agent orchestration built on \assetops{}. We combine final rank sheets, a 300-submission server log, 149-team registrations, best-submission exports, the organizer winners report, the companion \assetopslive{} system paper, and verified planning-track source trees. Five results stand out. First, the public planning leaderboard saturates at 72.73\%, and richer prompts do not improve that peak. Second, hidden evaluation changes the story: public and private scores correlate moderately in planning ($r{=}0.69$) but negatively in execution ($r{=}{-}0.13$), with several 45.45\% public execution systems reaching 63.64\% on the hidden set. Third, the \tmatch{} term is numerically almost inert in the official composite -- combined on a 0--1 scale with 0--100 percentage scores, it contributes at most 0.05 points per track, and rescaling would swap the top two teams. Fourth, the competition is operationally account-based but substantively team-based: 149 registered teams reduce to 24 with non-zero public scores and 11 fully ranked, while 52.3\% of deduplicated registrations list multiple usernames. Fifth, successful execution methods mostly improve guardrails -- response selection, contamination cleanup, fallback, and context control -- rather than novel agent architectures. These findings identify which behaviors the evaluation rewarded, and motivate scale-aware composites, skill-level diagnostics, and versioned artifact release.
title Results and Retrospective Analysis of the CODS 2025 AssetOpsBench Challenge
topic Artificial Intelligence
url https://arxiv.org/abs/2605.08518