RoboPhD: Evolving Diverse Complex Agents Under Tight Evaluation Budgets
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Borthwick, Andrew, Ash, Stephen, Galczak, Anthony |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RoboPhD: Self-Improving Text-to-SQL Through Autonomous Agent Evolution
von: Borthwick, Andrew, et al.
Veröffentlicht: (2026)
von: Borthwick, Andrew, et al.
Veröffentlicht: (2026)
ORBIT: Scalable and Verifiable Data Generation for Search Agents on a Tight Budget
von: Thakur, Nandan, et al.
Veröffentlicht: (2026)
von: Thakur, Nandan, et al.
Veröffentlicht: (2026)
AgentGym: Evolving Large Language Model-based Agents across Diverse Environments
von: Xi, Zhiheng, et al.
Veröffentlicht: (2024)
von: Xi, Zhiheng, et al.
Veröffentlicht: (2024)
ContextBudget: Budget-Aware Context Management for Long-Horizon Search Agents
von: Wu, Yong, et al.
Veröffentlicht: (2026)
von: Wu, Yong, et al.
Veröffentlicht: (2026)
Phase Transition for Budgeted Multi-Agent Synergy
von: Liu, Bang, et al.
Veröffentlicht: (2026)
von: Liu, Bang, et al.
Veröffentlicht: (2026)
SEA-Eval: A Benchmark for Evaluating Self-Evolving Agents Beyond Episodic Assessment
von: Jiang, Sihang, et al.
Veröffentlicht: (2026)
von: Jiang, Sihang, et al.
Veröffentlicht: (2026)
RoboLayout: Differentiable 3D Scene Generation for Embodied Agents
von: Shamsaddinlou, Ali
Veröffentlicht: (2026)
von: Shamsaddinlou, Ali
Veröffentlicht: (2026)
When Agents Evolve, Institutions Follow
von: Fei, Chao, et al.
Veröffentlicht: (2026)
von: Fei, Chao, et al.
Veröffentlicht: (2026)
Inference-Time Budget Control for LLM Search Agents
von: Fang, Zhengru, et al.
Veröffentlicht: (2026)
von: Fang, Zhengru, et al.
Veröffentlicht: (2026)
PhD: A ChatGPT-Prompted Visual hallucination Evaluation Dataset
von: Liu, Jiazhen, et al.
Veröffentlicht: (2024)
von: Liu, Jiazhen, et al.
Veröffentlicht: (2024)
Towards Goal-Oriented Agents for Evolving Problems Observed via Conversation
von: Free, Michael, et al.
Veröffentlicht: (2024)
von: Free, Michael, et al.
Veröffentlicht: (2024)
Alita-G: Self-Evolving Generative Agent for Agent Generation
von: Qiu, Jiahao, et al.
Veröffentlicht: (2025)
von: Qiu, Jiahao, et al.
Veröffentlicht: (2025)
Agents of Change: Self-Evolving LLM Agents for Strategic Planning
von: Belle, Nikolas, et al.
Veröffentlicht: (2025)
von: Belle, Nikolas, et al.
Veröffentlicht: (2025)
RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation
von: Jiang, Feng, et al.
Veröffentlicht: (2026)
von: Jiang, Feng, et al.
Veröffentlicht: (2026)
Beyond Perfect APIs: A Comprehensive Evaluation of LLM Agents Under Real-World API Complexity
von: Kim, Doyoung, et al.
Veröffentlicht: (2026)
von: Kim, Doyoung, et al.
Veröffentlicht: (2026)
Efficient Agent Evaluation via Diversity-Guided User Simulation
von: Nakash, Itay, et al.
Veröffentlicht: (2026)
von: Nakash, Itay, et al.
Veröffentlicht: (2026)
EXG: Self-Evolving Agents with Experience Graphs
von: Jin, Yuxin, et al.
Veröffentlicht: (2026)
von: Jin, Yuxin, et al.
Veröffentlicht: (2026)
Autogenesis: A Self-Evolving Agent Protocol
von: Zhang, Wentao, et al.
Veröffentlicht: (2026)
von: Zhang, Wentao, et al.
Veröffentlicht: (2026)
Real-Time Reasoning Agents in Evolving Environments
von: Wen, Yule, et al.
Veröffentlicht: (2025)
von: Wen, Yule, et al.
Veröffentlicht: (2025)
Self-Evolving Software Agents
von: Robol, Marco, et al.
Veröffentlicht: (2026)
von: Robol, Marco, et al.
Veröffentlicht: (2026)
RoboCertProb: Property Specification for Probabilistic RoboChart Models
von: Ye, Kangfeng, et al.
Veröffentlicht: (2024)
von: Ye, Kangfeng, et al.
Veröffentlicht: (2024)
Evolving-RL: End-to-End Optimization of Experience-Driven Self-Evolving Capability within Agents
von: Fan, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Fan, Zhiyuan, et al.
Veröffentlicht: (2026)
Budget-Aware Tool-Use Enables Effective Agent Scaling
von: Liu, Tengxiao, et al.
Veröffentlicht: (2025)
von: Liu, Tengxiao, et al.
Veröffentlicht: (2025)
AutoAgent: Evolving Cognition and Elastic Memory Orchestration for Adaptive Agents
von: Wang, Xiaoxing, et al.
Veröffentlicht: (2026)
von: Wang, Xiaoxing, et al.
Veröffentlicht: (2026)
AgentDevel: Reframing Self-Evolving LLM Agents as Release Engineering
von: Zhang, Di
Veröffentlicht: (2026)
von: Zhang, Di
Veröffentlicht: (2026)
DSGBench: A Diverse Strategic Game Benchmark for Evaluating LLM-based Agents in Complex Decision-Making Environments
von: Tang, Wenjie, et al.
Veröffentlicht: (2025)
von: Tang, Wenjie, et al.
Veröffentlicht: (2025)
AppAgentX: Evolving GUI Agents as Proficient Smartphone Users
von: Jiang, Wenjia, et al.
Veröffentlicht: (2025)
von: Jiang, Wenjia, et al.
Veröffentlicht: (2025)
EvoTool: Self-Evolving Tool-Use Policy Optimization in LLM Agents via Blame-Aware Mutation and Diversity-Aware Selection
von: Yang, Shuo, et al.
Veröffentlicht: (2026)
von: Yang, Shuo, et al.
Veröffentlicht: (2026)
RoboCurate: Harnessing Diversity with Action-Verified Neural Trajectory for Robot Learning
von: Kim, Seungku, et al.
Veröffentlicht: (2026)
von: Kim, Seungku, et al.
Veröffentlicht: (2026)
CODESKILL: Learning Self-Evolving Skills for Coding Agents
von: Li, Yanzhou, et al.
Veröffentlicht: (2026)
von: Li, Yanzhou, et al.
Veröffentlicht: (2026)
SEDM: Scalable Self-Evolving Distributed Memory for Agents
von: Xu, Haoran, et al.
Veröffentlicht: (2025)
von: Xu, Haoran, et al.
Veröffentlicht: (2025)
MorphAgent: Empowering Agents through Self-Evolving Profiles and Decentralized Collaboration
von: Lu, Siyuan, et al.
Veröffentlicht: (2024)
von: Lu, Siyuan, et al.
Veröffentlicht: (2024)
EVE-Agent: Evidence-Verifiable Self-Evolving Agents
von: Arai, Yamato, et al.
Veröffentlicht: (2026)
von: Arai, Yamato, et al.
Veröffentlicht: (2026)
Evaluating Frontier LLMs on PhD-Level Mathematical Reasoning: A Benchmark on a Textbook in Theoretical Computer Science about Randomized Algorithms
von: Cao, Yang, et al.
Veröffentlicht: (2025)
von: Cao, Yang, et al.
Veröffentlicht: (2025)
Agent Alignment in Evolving Social Norms
von: Li, Shimin, et al.
Veröffentlicht: (2024)
von: Li, Shimin, et al.
Veröffentlicht: (2024)
A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence
von: Gao, Huan-ang, et al.
Veröffentlicht: (2025)
von: Gao, Huan-ang, et al.
Veröffentlicht: (2025)
Observation Denoising in CYRUS Soccer Simulation 2D Team For RoboCup 2024
von: Zare, Nader, et al.
Veröffentlicht: (2024)
von: Zare, Nader, et al.
Veröffentlicht: (2024)
SOP-Bench: Complex Industrial SOPs for Evaluating LLM Agents
von: Nandi, Subhrangshu, et al.
Veröffentlicht: (2025)
von: Nandi, Subhrangshu, et al.
Veröffentlicht: (2025)
EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle
von: Wu, Rong, et al.
Veröffentlicht: (2025)
von: Wu, Rong, et al.
Veröffentlicht: (2025)
SD-E$^2$: Semantic Exploration for Reasoning Under Token Budgets
von: Mishra, Kshitij, et al.
Veröffentlicht: (2026)
von: Mishra, Kshitij, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
RoboPhD: Self-Improving Text-to-SQL Through Autonomous Agent Evolution
von: Borthwick, Andrew, et al.
Veröffentlicht: (2026) -
ORBIT: Scalable and Verifiable Data Generation for Search Agents on a Tight Budget
von: Thakur, Nandan, et al.
Veröffentlicht: (2026) -
AgentGym: Evolving Large Language Model-based Agents across Diverse Environments
von: Xi, Zhiheng, et al.
Veröffentlicht: (2024) -
ContextBudget: Budget-Aware Context Management for Long-Horizon Search Agents
von: Wu, Yong, et al.
Veröffentlicht: (2026) -
Phase Transition for Budgeted Multi-Agent Synergy
von: Liu, Bang, et al.
Veröffentlicht: (2026)