Towards Self-Evolving Benchmarks: Synthesizing Agent Trajectories via Test-Time Exploration under Validate-by-Reproduce Paradigm
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Dadi, Zhou, Tianyi, Liu, Dongrui, Qian, Chen, Ren, Qihan, Shao, Shuai, Fan, Zhiyuan, Fung, Yi R., Wang, Kun, Zhang, Linfeng, Shao, Jing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Code2Math: Can Your Code Agent Effectively Evolve Math Problems Through Exploration?
by: Guo, Dadi, et al.
Published: (2026)
by: Guo, Dadi, et al.
Published: (2026)
Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agents
by: Shao, Shuai, et al.
Published: (2025)
by: Shao, Shuai, et al.
Published: (2025)
ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use Agents
by: Yang, Jingyi, et al.
Published: (2025)
by: Yang, Jingyi, et al.
Published: (2025)
Are Your Agents Upward Deceivers?
by: Guo, Dadi, et al.
Published: (2025)
by: Guo, Dadi, et al.
Published: (2025)
Attributing Emergence in Million-Agent Systems
by: Tang, Ling, et al.
Published: (2026)
by: Tang, Ling, et al.
Published: (2026)
Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability
by: Ren, Qihan, et al.
Published: (2026)
by: Ren, Qihan, et al.
Published: (2026)
The Why Behind the Action: Unveiling Internal Drivers via Agentic Attribution
by: Qian, Chen, et al.
Published: (2026)
by: Qian, Chen, et al.
Published: (2026)
LED-Merging: Mitigating Safety-Utility Conflicts in Model Merging with Location-Election-Disjoint
by: Ma, Qianli, et al.
Published: (2025)
by: Ma, Qianli, et al.
Published: (2025)
Self-Consolidation for Self-Evolving Agents
by: Yu, Hongzhuo, et al.
Published: (2026)
by: Yu, Hongzhuo, et al.
Published: (2026)
APEX: Autonomous Policy Exploration for Self-Evolving LLM Agents
by: Li, Yibo, et al.
Published: (2026)
by: Li, Yibo, et al.
Published: (2026)
EE-MCP: Self-Evolving MCP-GUI Agents via Automated Environment Generation and Experience Learning
by: He, Tiantian, et al.
Published: (2026)
by: He, Tiantian, et al.
Published: (2026)
Benchmarks for Trajectory Safety Evaluation and Diagnosis in OpenClaw and Codex: ATBench-Claw and ATBench-Codex
by: Yang, Zhonghao, et al.
Published: (2026)
by: Yang, Zhonghao, et al.
Published: (2026)
Independent characterization of the elastic and the mixing parts of hydrogel osmotic pressure
by: Shao, Zefan, et al.
Published: (2023)
by: Shao, Zefan, et al.
Published: (2023)
Towards the Dynamics of a DNN Learning Symbolic Interactions
by: Ren, Qihan, et al.
Published: (2024)
by: Ren, Qihan, et al.
Published: (2024)
Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
by: Wei, Tianxin, et al.
Published: (2025)
by: Wei, Tianxin, et al.
Published: (2025)
MorphAgent: Empowering Agents through Self-Evolving Profiles and Decentralized Collaboration
by: Lu, Siyuan, et al.
Published: (2024)
by: Lu, Siyuan, et al.
Published: (2024)
Conditional Advantage Estimation for Reinforcement Learning in Large Reasoning Models
by: Chen, Guanxu, et al.
Published: (2025)
by: Chen, Guanxu, et al.
Published: (2025)
MetaAgent: Toward Self-Evolving Agent via Tool Meta-Learning
by: Qian, Hongjin, et al.
Published: (2025)
by: Qian, Hongjin, et al.
Published: (2025)
AgentSlimming: Towards Efficient and Cost-Aware Multi-Agent Systems
by: Chen, Yulang, et al.
Published: (2026)
by: Chen, Yulang, et al.
Published: (2026)
Entropy-Gradient Inversion: Moving Toward Internal Mechanism of Large Reasoning Models
by: Yang, Junyao, et al.
Published: (2026)
by: Yang, Junyao, et al.
Published: (2026)
A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence
by: Gao, Huan-ang, et al.
Published: (2025)
by: Gao, Huan-ang, et al.
Published: (2025)
Autonomous Self‐Evolving Research on Biomedical Data: The DREAM Paradigm
by: Luojia Deng, et al.
Published: (2025)
by: Luojia Deng, et al.
Published: (2025)
REEF: Representation Encoding Fingerprints for Large Language Models
by: Zhang, Jie, et al.
Published: (2024)
by: Zhang, Jie, et al.
Published: (2024)
EvolveReason: Self-Evolving Reasoning Paradigm for Explainable Deepfake Facial Image Identification
by: Zhou, Binjia, et al.
Published: (2026)
by: Zhou, Binjia, et al.
Published: (2026)
TradeTrap: Are LLM-based Trading Agents Truly Reliable and Faithful?
by: Yan, Lewen, et al.
Published: (2025)
by: Yan, Lewen, et al.
Published: (2025)
Benchmark Test-Time Scaling of General LLM Agents
by: Li, Xiaochuan, et al.
Published: (2026)
by: Li, Xiaochuan, et al.
Published: (2026)
COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation
by: Zhou, Tianyi, et al.
Published: (2026)
by: Zhou, Tianyi, et al.
Published: (2026)
Evaluating the Correctness of Inference Patterns Used by LLMs for Judgment
by: Chen, Lu, et al.
Published: (2024)
by: Chen, Lu, et al.
Published: (2024)
SEARL: Joint Optimization of Policy and Tool Graph Memory for Self-Evolving Agents
by: Feng, Xinshun, et al.
Published: (2026)
by: Feng, Xinshun, et al.
Published: (2026)
Tree-frog-inspired osmocapillary adhesive bonding to diverse substrates
by: Shao, Zefan, et al.
Published: (2025)
by: Shao, Zefan, et al.
Published: (2025)
SEFRQO: A Self-Evolving Fine-Tuned RAG-Based Query Optimizer
by: Liu, Hanwen, et al.
Published: (2025)
by: Liu, Hanwen, et al.
Published: (2025)
Adapting Like Humans: A Metacognitive Agent with Test-time Reasoning
by: Li, Yang, et al.
Published: (2025)
by: Li, Yang, et al.
Published: (2025)
Mathematical Proof as a Litmus Test: Revealing Failure Modes of Advanced Large Reasoning Models
by: Guo, Dadi, et al.
Published: (2025)
by: Guo, Dadi, et al.
Published: (2025)
LatentEvolve: Self-Evolving Test-Time Scaling in Latent Space
by: Zhang, Guibin, et al.
Published: (2025)
by: Zhang, Guibin, et al.
Published: (2025)
Interpreting Emergent Extreme Events in Multi-Agent Systems
by: Tang, Ling, et al.
Published: (2026)
by: Tang, Ling, et al.
Published: (2026)
Loop as a Bridge: Can Looped Transformers Truly Link Representation Space and Natural Language Outputs?
by: Chen, Guanxu, et al.
Published: (2026)
by: Chen, Guanxu, et al.
Published: (2026)
AgentEvolver: Towards Efficient Self-Evolving Agent System
by: Zhai, Yunpeng, et al.
Published: (2025)
by: Zhai, Yunpeng, et al.
Published: (2025)
MANTRA: Synthesizing SMT-Validated Compliance Benchmarks for Tool-Using LLM Agents
by: Anand, Ashwani, et al.
Published: (2026)
by: Anand, Ashwani, et al.
Published: (2026)
RBoard: A Unified Platform for Reproducible and Reusable Recommender System Benchmarks
by: Shao, Xinyang, et al.
Published: (2024)
by: Shao, Xinyang, et al.
Published: (2024)
Similar Items
-
Code2Math: Can Your Code Agent Effectively Evolve Math Problems Through Exploration?
by: Guo, Dadi, et al.
Published: (2026) -
Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agents
by: Shao, Shuai, et al.
Published: (2025) -
ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis
by: Li, Yu, et al.
Published: (2026) -
RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use Agents
by: Yang, Jingyi, et al.
Published: (2025) -
Are Your Agents Upward Deceivers?
by: Guo, Dadi, et al.
Published: (2025)