SEA-Eval: A Benchmark for Evaluating Self-Evolving Agents Beyond Episodic Assessment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jiang, Sihang, Ma, Lipeng, Hong, Zhonghua, Wang, Keyi, Lu, Zhiyu, Wang, Tengfei, Chen, Shisong, Zhang, Jinghao, Pan, Tianjun, Li, Weijia, Liang, Jiaqing, Xiao, Yanghua |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GenericAgent: A Token-Efficient Self-Evolving LLM Agent via Contextual Information Density Maximization (V1.0)
von: Liang, Jiaqing, et al.
Veröffentlicht: (2026)
von: Liang, Jiaqing, et al.
Veröffentlicht: (2026)
RubricEval: A Rubric-Level Meta-Evaluation Benchmark for LLM Judges in Instruction Following
von: Pan, Tianjun, et al.
Veröffentlicht: (2026)
von: Pan, Tianjun, et al.
Veröffentlicht: (2026)
CultureScope: A Dimensional Lens for Probing Cultural Understanding in LLMs
von: Zhang, Jinghao, et al.
Veröffentlicht: (2025)
von: Zhang, Jinghao, et al.
Veröffentlicht: (2025)
MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory
von: Zhang, Shengtao, et al.
Veröffentlicht: (2026)
von: Zhang, Shengtao, et al.
Veröffentlicht: (2026)
Enhancing Quantitative Reasoning Skills of Large Language Models through Dimension Perception
von: Huang, Yuncheng, et al.
Veröffentlicht: (2023)
von: Huang, Yuncheng, et al.
Veröffentlicht: (2023)
Beyond Asymptotics: Practical Insights into Community Detection in Complex Networks
von: Ke, Tianjun, et al.
Veröffentlicht: (2024)
von: Ke, Tianjun, et al.
Veröffentlicht: (2024)
SEIF: Self-Evolving Reinforcement Learning for Instruction Following
von: Ren, Qingyu, et al.
Veröffentlicht: (2026)
von: Ren, Qingyu, et al.
Veröffentlicht: (2026)
CEM: A Data-Efficient Method for Large Language Models to Continue Evolving From Mistakes
von: Zhao, Haokun, et al.
Veröffentlicht: (2024)
von: Zhao, Haokun, et al.
Veröffentlicht: (2024)
Enhancing Confidence Expression in Large Language Models Through Learning from Past Experience
von: Han, Haixia, et al.
Veröffentlicht: (2024)
von: Han, Haixia, et al.
Veröffentlicht: (2024)
When Agents Evolve, Institutions Follow
von: Fei, Chao, et al.
Veröffentlicht: (2026)
von: Fei, Chao, et al.
Veröffentlicht: (2026)
ComLQ: Benchmarking Complex Logical Queries in Information Retrieval
von: Xu, Ganlin, et al.
Veröffentlicht: (2025)
von: Xu, Ganlin, et al.
Veröffentlicht: (2025)
Small Language Model Can Self-correct
von: Han, Haixia, et al.
Veröffentlicht: (2024)
von: Han, Haixia, et al.
Veröffentlicht: (2024)
ESC-Eval: Evaluating Emotion Support Conversations in Large Language Models
von: Zhao, Haiquan, et al.
Veröffentlicht: (2024)
von: Zhao, Haiquan, et al.
Veröffentlicht: (2024)
BookWorld: From Novels to Interactive Agent Societies for Creative Story Generation
von: Ran, Yiting, et al.
Veröffentlicht: (2025)
von: Ran, Yiting, et al.
Veröffentlicht: (2025)
Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows
von: Li, Chenxin, et al.
Veröffentlicht: (2026)
von: Li, Chenxin, et al.
Veröffentlicht: (2026)
SEA-TS: Self-Evolving Agent for Autonomous Code Generation of Time Series Forecasting Algorithms
von: Xu, Longkun, et al.
Veröffentlicht: (2026)
von: Xu, Longkun, et al.
Veröffentlicht: (2026)
A Stitch in Time Saves Nine: Proactive Self-Refinement for Language Models
von: Han, Jinyi, et al.
Veröffentlicht: (2025)
von: Han, Jinyi, et al.
Veröffentlicht: (2025)
Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following
von: Ren, Qingyu, et al.
Veröffentlicht: (2025)
von: Ren, Qingyu, et al.
Veröffentlicht: (2025)
Mind the Generation Process: Fine-Grained Confidence Estimation During LLM Generation
von: Han, Jinyi, et al.
Veröffentlicht: (2025)
von: Han, Jinyi, et al.
Veröffentlicht: (2025)
AgenticEval: Toward Agentic and Self-Evolving Safety Evaluation of Large Language Models
von: Wang, Yixu, et al.
Veröffentlicht: (2025)
von: Wang, Yixu, et al.
Veröffentlicht: (2025)
ConcEPT: Concept-Enhanced Pre-Training for Language Models
von: Wang, Xintao, et al.
Veröffentlicht: (2024)
von: Wang, Xintao, et al.
Veröffentlicht: (2024)
HINT: Helping Ineffective Rollouts Navigate Towards Effectiveness
von: Wang, Xinyi, et al.
Veröffentlicht: (2025)
von: Wang, Xinyi, et al.
Veröffentlicht: (2025)
Selective Expert Guidance for Effective and Diverse Exploration in Reinforcement Learning of LLMs
von: Jiang, Zishang, et al.
Veröffentlicht: (2025)
von: Jiang, Zishang, et al.
Veröffentlicht: (2025)
Embodied Co-Design for Rapidly Evolving Agents: Taxonomy, Frontiers, and Challenges
von: Wang, Yuxing, et al.
Veröffentlicht: (2025)
von: Wang, Yuxing, et al.
Veröffentlicht: (2025)
Do Large Language Models Truly Understand Cross-cultural Differences?
von: Guo, Shiwei, et al.
Veröffentlicht: (2025)
von: Guo, Shiwei, et al.
Veröffentlicht: (2025)
OVEL: Large Language Model as Memory Manager for Online Video Entity Linking
von: Zhao, Haiquan, et al.
Veröffentlicht: (2024)
von: Zhao, Haiquan, et al.
Veröffentlicht: (2024)
DiagEval: Trajectory-Conditioned Diagnosis for Reliable Software Evaluation with GUI Agents
von: Hong, Sirui, et al.
Veröffentlicht: (2026)
von: Hong, Sirui, et al.
Veröffentlicht: (2026)
Experience-Evolving Multi-Turn Tool-Use Agent with Hybrid Episodic-Procedural Memory
von: Li, Sijia, et al.
Veröffentlicht: (2025)
von: Li, Sijia, et al.
Veröffentlicht: (2025)
Agentified Assessment of Logical Reasoning Agents
von: Ni, Zhiyu, et al.
Veröffentlicht: (2026)
von: Ni, Zhiyu, et al.
Veröffentlicht: (2026)
Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation
von: Wang, Siyuan, et al.
Veröffentlicht: (2024)
von: Wang, Siyuan, et al.
Veröffentlicht: (2024)
Light Up the Shadows: Enhance Long-Tailed Entity Grounding with Concept-Guided Vision-Language Models
von: Zhang, Yikai, et al.
Veröffentlicht: (2024)
von: Zhang, Yikai, et al.
Veröffentlicht: (2024)
Chain-of-Knowledge: Integrating Knowledge Reasoning into Large Language Models by Learning from Knowledge Graphs
von: Zhang, Yifei, et al.
Veröffentlicht: (2024)
von: Zhang, Yifei, et al.
Veröffentlicht: (2024)
Laying the Foundation First? Investigating the Generalization from Atomic Skills to Complex Reasoning Tasks
von: Huang, Yuncheng, et al.
Veröffentlicht: (2024)
von: Huang, Yuncheng, et al.
Veröffentlicht: (2024)
Adaptive Ordered Information Extraction with Deep Reinforcement Learning
von: Huang, Wenhao, et al.
Veröffentlicht: (2023)
von: Huang, Wenhao, et al.
Veröffentlicht: (2023)
From Complex to Simple: Enhancing Multi-Constraint Complex Instruction Following Ability of Large Language Models
von: He, Qianyu, et al.
Veröffentlicht: (2024)
von: He, Qianyu, et al.
Veröffentlicht: (2024)
Is There a One-Model-Fits-All Approach to Information Extraction? Revisiting Task Definition Biases
von: Huang, Wenhao, et al.
Veröffentlicht: (2024)
von: Huang, Wenhao, et al.
Veröffentlicht: (2024)
Improving Recall of Large Language Models: A Model Collaboration Approach for Relational Triple Extraction
von: Ding, Zepeng, et al.
Veröffentlicht: (2024)
von: Ding, Zepeng, et al.
Veröffentlicht: (2024)
EDGE: Enhanced Grounded GUI Understanding with Enriched Multi-Granularity Synthetic Data
von: Chen, Xuetian, et al.
Veröffentlicht: (2024)
von: Chen, Xuetian, et al.
Veröffentlicht: (2024)
From Coarse to Fine: Benchmarking and Reward Modeling for Writing-Centric Generation Tasks
von: Ren, Qingyu, et al.
Veröffentlicht: (2026)
von: Ren, Qingyu, et al.
Veröffentlicht: (2026)
Building Self-Evolving Agents via Experience-Driven Lifelong Learning: A Framework and Benchmark
von: Cai, Yuxuan, et al.
Veröffentlicht: (2025)
von: Cai, Yuxuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
GenericAgent: A Token-Efficient Self-Evolving LLM Agent via Contextual Information Density Maximization (V1.0)
von: Liang, Jiaqing, et al.
Veröffentlicht: (2026) -
RubricEval: A Rubric-Level Meta-Evaluation Benchmark for LLM Judges in Instruction Following
von: Pan, Tianjun, et al.
Veröffentlicht: (2026) -
CultureScope: A Dimensional Lens for Probing Cultural Understanding in LLMs
von: Zhang, Jinghao, et al.
Veröffentlicht: (2025) -
MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory
von: Zhang, Shengtao, et al.
Veröffentlicht: (2026) -
Enhancing Quantitative Reasoning Skills of Large Language Models through Dimension Perception
von: Huang, Yuncheng, et al.
Veröffentlicht: (2023)