Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Siyuan, Long, Zhuohan, Fan, Zhihao, Wei, Zhongyu, Huang, Xuanjing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
From LLMs to MLLMs: Exploring the Landscape of Multimodal Jailbreaking
von: Wang, Siyuan, et al.
Veröffentlicht: (2024)
von: Wang, Siyuan, et al.
Veröffentlicht: (2024)
Synergistic Multi-Agent Framework with Trajectory Learning for Knowledge-Intensive Tasks
von: Yue, Shengbin, et al.
Veröffentlicht: (2024)
von: Yue, Shengbin, et al.
Veröffentlicht: (2024)
Strong Reasoning Isn't Enough: Evaluating Evidence Elicitation in Interactive Diagnosis
von: Long, Zhuohan, et al.
Veröffentlicht: (2026)
von: Long, Zhuohan, et al.
Veröffentlicht: (2026)
Unveiling the Truth and Facilitating Change: Towards Agent-based Large-scale Social Movement Simulation
von: Mou, Xinyi, et al.
Veröffentlicht: (2024)
von: Mou, Xinyi, et al.
Veröffentlicht: (2024)
LifeSim: Long-Horizon User Life Simulator for Personalized Assistant Evaluation
von: Duan, Feiyu, et al.
Veröffentlicht: (2026)
von: Duan, Feiyu, et al.
Veröffentlicht: (2026)
AI Hospital: Benchmarking Large Language Models in a Multi-agent Medical Interaction Simulator
von: Fan, Zhihao, et al.
Veröffentlicht: (2024)
von: Fan, Zhihao, et al.
Veröffentlicht: (2024)
Multi-Agent Simulator Drives Language Models for Legal Intensive Interaction
von: Yue, Shengbin, et al.
Veröffentlicht: (2025)
von: Yue, Shengbin, et al.
Veröffentlicht: (2025)
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs
von: Ding, Xuanwen, et al.
Veröffentlicht: (2025)
von: Ding, Xuanwen, et al.
Veröffentlicht: (2025)
How Jailbreak Defenses Work and Ensemble? A Mechanistic Investigation
von: Long, Zhuohang, et al.
Veröffentlicht: (2025)
von: Long, Zhuohang, et al.
Veröffentlicht: (2025)
HAF-RM: A Hybrid Alignment Framework for Reward Model Training
von: Liu, Shujun, et al.
Veröffentlicht: (2024)
von: Liu, Shujun, et al.
Veröffentlicht: (2024)
EcoLANG: Efficient and Effective Agent Communication Language Induction for Social Simulation
von: Mou, Xinyi, et al.
Veröffentlicht: (2025)
von: Mou, Xinyi, et al.
Veröffentlicht: (2025)
ALaRM: Align Language Models via Hierarchical Rewards Modeling
von: Lai, Yuhang, et al.
Veröffentlicht: (2024)
von: Lai, Yuhang, et al.
Veröffentlicht: (2024)
Debatrix: Multi-dimensional Debate Judge with Iterative Chronological Analysis Based on LLM
von: Liang, Jingcong, et al.
Veröffentlicht: (2024)
von: Liang, Jingcong, et al.
Veröffentlicht: (2024)
HyLaT: Efficient Multi-Agent Communication via Hybrid Latent-Text Protocol
von: Mou, Xinyi, et al.
Veröffentlicht: (2026)
von: Mou, Xinyi, et al.
Veröffentlicht: (2026)
Stepwise Informativeness Search for Efficient and Effective LLM Reasoning
von: Wang, Siyuan, et al.
Veröffentlicht: (2025)
von: Wang, Siyuan, et al.
Veröffentlicht: (2025)
EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models
von: Du, Mengfei, et al.
Veröffentlicht: (2024)
von: Du, Mengfei, et al.
Veröffentlicht: (2024)
AgentSense: Benchmarking Social Intelligence of Language Agents through Interactive Scenarios
von: Mou, Xinyi, et al.
Veröffentlicht: (2024)
von: Mou, Xinyi, et al.
Veröffentlicht: (2024)
CURP: Codebook-based Continuous User Representation for Personalized Generation with LLMs
von: Wang, Liang, et al.
Veröffentlicht: (2026)
von: Wang, Liang, et al.
Veröffentlicht: (2026)
Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
von: Wei, Tianxin, et al.
Veröffentlicht: (2025)
von: Wei, Tianxin, et al.
Veröffentlicht: (2025)
Benchmark^2: Systematic Evaluation of LLM Benchmarks
von: Qian, Qi, et al.
Veröffentlicht: (2026)
von: Qian, Qi, et al.
Veröffentlicht: (2026)
SEAD: Self-Evolving Agent for Multi-Turn Service Dialogue
von: Dai, Yuqin, et al.
Veröffentlicht: (2026)
von: Dai, Yuqin, et al.
Veröffentlicht: (2026)
MetaGen: Self-Evolving Roles and Topologies for Multi-Agent LLM Reasoning
von: Wang, Yimeng, et al.
Veröffentlicht: (2026)
von: Wang, Yimeng, et al.
Veröffentlicht: (2026)
AI-Press: A Multi-Agent News Generating and Feedback Simulation System Powered by Large Language Models
von: Liu, Xiawei, et al.
Veröffentlicht: (2024)
von: Liu, Xiawei, et al.
Veröffentlicht: (2024)
SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents
von: Shen, Yujiong, et al.
Veröffentlicht: (2026)
von: Shen, Yujiong, et al.
Veröffentlicht: (2026)
DELAN: Dual-Level Alignment for Vision-and-Language Navigation by Cross-Modal Contrastive Learning
von: Du, Mengfei, et al.
Veröffentlicht: (2024)
von: Du, Mengfei, et al.
Veröffentlicht: (2024)
Functional Consistency of LLM Code Embeddings: A Self-Evolving Data Synthesis Framework for Benchmarking
von: Li, Zhuohao, et al.
Veröffentlicht: (2025)
von: Li, Zhuohao, et al.
Veröffentlicht: (2025)
Ratchet: A Minimal Hygiene Recipe for Self-Evolving LLM Agents
von: Zhang, Xing, et al.
Veröffentlicht: (2026)
von: Zhang, Xing, et al.
Veröffentlicht: (2026)
VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models
von: Li, Zejun, et al.
Veröffentlicht: (2024)
von: Li, Zejun, et al.
Veröffentlicht: (2024)
EvoWiki: Evaluating LLMs on Evolving Knowledge
von: Tang, Wei, et al.
Veröffentlicht: (2024)
von: Tang, Wei, et al.
Veröffentlicht: (2024)
MulDimIF: A Multi-Dimensional Constraint Framework for Evaluating and Improving Instruction Following in Large Language Models
von: Ye, Junjie, et al.
Veröffentlicht: (2025)
von: Ye, Junjie, et al.
Veröffentlicht: (2025)
EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle
von: Wu, Rong, et al.
Veröffentlicht: (2025)
von: Wu, Rong, et al.
Veröffentlicht: (2025)
CryptoBench: A Dynamic Benchmark for Expert-Level Evaluation of LLM Agents in Cryptocurrency
von: Guo, Jiacheng, et al.
Veröffentlicht: (2025)
von: Guo, Jiacheng, et al.
Veröffentlicht: (2025)
SELAUR: Self Evolving LLM Agent via Uncertainty-aware Rewards
von: Zhang, Dengjia, et al.
Veröffentlicht: (2026)
von: Zhang, Dengjia, et al.
Veröffentlicht: (2026)
SEMA-RAG: A Self-Evolving Multi-Agent Retrieval-Augmented Generation Framework for Medical Reasoning
von: Huang, Yongfeng, et al.
Veröffentlicht: (2026)
von: Huang, Yongfeng, et al.
Veröffentlicht: (2026)
EvolveSearch: An Iterative Self-Evolving Search Agent
von: Zhang, Dingchu, et al.
Veröffentlicht: (2025)
von: Zhang, Dingchu, et al.
Veröffentlicht: (2025)
STELLA: Self-Evolving LLM Agent for Biomedical Research
von: Jin, Ruofan, et al.
Veröffentlicht: (2025)
von: Jin, Ruofan, et al.
Veröffentlicht: (2025)
Building Self-Evolving Agents via Experience-Driven Lifelong Learning: A Framework and Benchmark
von: Cai, Yuxuan, et al.
Veröffentlicht: (2025)
von: Cai, Yuxuan, et al.
Veröffentlicht: (2025)
PIORS: Personalized Intelligent Outpatient Reception based on Large Language Model with Multi-Agents Medical Scenario Simulation
von: Bao, Zhijie, et al.
Veröffentlicht: (2024)
von: Bao, Zhijie, et al.
Veröffentlicht: (2024)
WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning
von: Qi, Zehan, et al.
Veröffentlicht: (2024)
von: Qi, Zehan, et al.
Veröffentlicht: (2024)
A Self-Evolving Framework for Efficient Terminal Agents via Observational Context Compression
von: Ren, Jincheng, et al.
Veröffentlicht: (2026)
von: Ren, Jincheng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
From LLMs to MLLMs: Exploring the Landscape of Multimodal Jailbreaking
von: Wang, Siyuan, et al.
Veröffentlicht: (2024) -
Synergistic Multi-Agent Framework with Trajectory Learning for Knowledge-Intensive Tasks
von: Yue, Shengbin, et al.
Veröffentlicht: (2024) -
Strong Reasoning Isn't Enough: Evaluating Evidence Elicitation in Interactive Diagnosis
von: Long, Zhuohan, et al.
Veröffentlicht: (2026) -
Unveiling the Truth and Facilitating Change: Towards Agent-based Large-scale Social Movement Simulation
von: Mou, Xinyi, et al.
Veröffentlicht: (2024) -
LifeSim: Long-Horizon User Life Simulator for Personalized Assistant Evaluation
von: Duan, Feiyu, et al.
Veröffentlicht: (2026)