MOCHA: Multi-Objective Chebyshev Annealing for Agent Skill Optimization

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Tanjim, Md Mehrab, Subramanian, Jayakumar, Chen, Xiang, Kveton, Branislav, Mukherjee, Subhojyoti, Zhang, Anlan, Kim, Sungchul, Sarkhel, Somdeb, Choudhury, Sunav
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911696953540608
author Tanjim, Md Mehrab
Subramanian, Jayakumar
Chen, Xiang
Kveton, Branislav
Mukherjee, Subhojyoti
Zhang, Anlan
Kim, Sungchul
Sarkhel, Somdeb
Choudhury, Sunav
author_facet Tanjim, Md Mehrab
Subramanian, Jayakumar
Chen, Xiang
Kveton, Branislav
Mukherjee, Subhojyoti
Zhang, Anlan
Kim, Sungchul
Sarkhel, Somdeb
Choudhury, Sunav
contents LLM agents organize behavior through skills - structured natural-language specifications governing how an agent reasons, retrieves, and responds. Unlike monolithic prompts, skills are multi-field artifacts subject to hard platform constraints: description fields are truncated for routing, instruction bodies are compacted via progressive disclosure, and co-resident skills compete for limited context windows. These constraints make skill optimization inherently multi-objective: a skill must simultaneously maximize task performance and satisfy platform limits. Yet existing prompt optimizers either ignore these trade-offs or collapse them into a weighted sum, missing Pareto-optimal variants in non-convex objective regions. We introduce MOCHA (Multi-Objective Chebyshev Annealing), which replaces single-objective selection with Chebyshev scalarization - covering the full Pareto front, including non-convex regions - combined with exponential annealing that transitions from exploration to exploitation. In our experiments across six diverse agent skills - where all methods share the same multi-objective mutation operator and baselines receive identical per-objective textual feedback - existing optimizers fail to improve the seed skill on 4 of 6 tasks: 1000 rollouts yield zero progress. MOCHA breaks through on every task, achieving 7.5% relative improvement in mean correctness over the strongest baseline (up to 14.9% on FEVER and 10.4% on TheoremQA) while discovering twice as many more Pareto-optimal skill variants.
format Preprint
id arxiv_https___arxiv_org_abs_2605_19330
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MOCHA: Multi-Objective Chebyshev Annealing for Agent Skill Optimization
Tanjim, Md Mehrab
Subramanian, Jayakumar
Chen, Xiang
Kveton, Branislav
Mukherjee, Subhojyoti
Zhang, Anlan
Kim, Sungchul
Sarkhel, Somdeb
Choudhury, Sunav
Artificial Intelligence
Machine Learning
Software Engineering
I.2.7; I.2.6; I.2.4; I.2.8
LLM agents organize behavior through skills - structured natural-language specifications governing how an agent reasons, retrieves, and responds. Unlike monolithic prompts, skills are multi-field artifacts subject to hard platform constraints: description fields are truncated for routing, instruction bodies are compacted via progressive disclosure, and co-resident skills compete for limited context windows. These constraints make skill optimization inherently multi-objective: a skill must simultaneously maximize task performance and satisfy platform limits. Yet existing prompt optimizers either ignore these trade-offs or collapse them into a weighted sum, missing Pareto-optimal variants in non-convex objective regions. We introduce MOCHA (Multi-Objective Chebyshev Annealing), which replaces single-objective selection with Chebyshev scalarization - covering the full Pareto front, including non-convex regions - combined with exponential annealing that transitions from exploration to exploitation. In our experiments across six diverse agent skills - where all methods share the same multi-objective mutation operator and baselines receive identical per-objective textual feedback - existing optimizers fail to improve the seed skill on 4 of 6 tasks: 1000 rollouts yield zero progress. MOCHA breaks through on every task, achieving 7.5% relative improvement in mean correctness over the strongest baseline (up to 14.9% on FEVER and 10.4% on TheoremQA) while discovering twice as many more Pareto-optimal skill variants.
title MOCHA: Multi-Objective Chebyshev Annealing for Agent Skill Optimization
topic Artificial Intelligence
Machine Learning
Software Engineering
I.2.7; I.2.6; I.2.4; I.2.8
url https://arxiv.org/abs/2605.19330