On the Importance of Task Complexity in Evaluating LLM-Based Multi-Agent Systems
Fuente:
arXiv
Guardado en:
| Autores principales: | Tang, Bohan, Liang, Huidong, Jiang, Keyue, Dong, Xiaowen |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Heterogeneous Graph Structure Learning through the Lens of Data-generating Processes
por: Jiang, Keyue, et al.
Publicado: (2025)
por: Jiang, Keyue, et al.
Publicado: (2025)
Training-Free Message Passing for Learning on Hypergraphs
por: Tang, Bohan, et al.
Publicado: (2024)
por: Tang, Bohan, et al.
Publicado: (2024)
A Markov Random Field model for Hypergraph-based Machine Learning
por: Tang, Bohan, et al.
Publicado: (2023)
por: Tang, Bohan, et al.
Publicado: (2023)
Bures-Wasserstein Flow Matching for Graph Generation
por: Jiang, Keyue, et al.
Publicado: (2025)
por: Jiang, Keyue, et al.
Publicado: (2025)
Towards Quantifying Long-Range Interactions in Graph Machine Learning: a Large Graph Dataset and a Measurement
por: Liang, Huidong, et al.
Publicado: (2025)
por: Liang, Huidong, et al.
Publicado: (2025)
Understanding Agent Scaling in LLM-Based Multi-Agent Systems via Diversity
por: Yang, Yingxuan, et al.
Publicado: (2026)
por: Yang, Yingxuan, et al.
Publicado: (2026)
Decentralized and Lifelong-Adaptive Multi-Agent Collaborative Learning
por: Tang, Shuo, et al.
Publicado: (2024)
por: Tang, Shuo, et al.
Publicado: (2024)
Active Learning for Communication Structure Optimization in LLM-Based Multi-Agent Systems
por: Yang, Huchen, et al.
Publicado: (2026)
por: Yang, Huchen, et al.
Publicado: (2026)
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks
por: Long, Xiang, et al.
Publicado: (2026)
por: Long, Xiang, et al.
Publicado: (2026)
TemporalBench: A Benchmark for Evaluating LLM-Based Agents on Contextual and Event-Informed Time Series Tasks
por: Weng, Muyan, et al.
Publicado: (2026)
por: Weng, Muyan, et al.
Publicado: (2026)
BrowserArena: Evaluating LLM Agents on Real-World Web Navigation Tasks
por: Anupam, Sagnik, et al.
Publicado: (2025)
por: Anupam, Sagnik, et al.
Publicado: (2025)
ARM: Discovering Agentic Reasoning Modules for Generalizable Multi-Agent Systems
por: Yao, Bohan, et al.
Publicado: (2025)
por: Yao, Bohan, et al.
Publicado: (2025)
Molecular Generative Adversarial Network with Multi-Property Optimization
por: Tang, Huidong, et al.
Publicado: (2024)
por: Tang, Huidong, et al.
Publicado: (2024)
GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models
por: Tang, Xiaohang, et al.
Publicado: (2026)
por: Tang, Xiaohang, et al.
Publicado: (2026)
Evaluating Long-Context Reasoning in LLM-Based WebAgents
por: Chung, Andy, et al.
Publicado: (2025)
por: Chung, Andy, et al.
Publicado: (2025)
COMET: Benchmark for Comprehensive Biological Multi-omics Evaluation Tasks and Language Models
por: Ren, Yuchen, et al.
Publicado: (2024)
por: Ren, Yuchen, et al.
Publicado: (2024)
Graph and Simplicial Complex Prediction Gaussian Process via the Hodgelet Representations
por: Alain, Mathieu, et al.
Publicado: (2025)
por: Alain, Mathieu, et al.
Publicado: (2025)
Beyond Distribution Sharpening: The Importance of Task Rewards
por: Mittal, Sarthak, et al.
Publicado: (2026)
por: Mittal, Sarthak, et al.
Publicado: (2026)
MLE-Smith: Scaling MLE Tasks with Automated Multi-Agent Pipeline
por: Qiang, Rushi, et al.
Publicado: (2025)
por: Qiang, Rushi, et al.
Publicado: (2025)
When Planning Fails Despite Correct Execution: On Epistemic Calibration for LLM-Based Multi-Agent Systems
por: Wang, Zehao, et al.
Publicado: (2026)
por: Wang, Zehao, et al.
Publicado: (2026)
Enhancing Player Enjoyment with a Two-Tier DRL and LLM-Based Agent System for Fighting Games
por: Wang, Shouren, et al.
Publicado: (2025)
por: Wang, Shouren, et al.
Publicado: (2025)
Agentic Neural Networks: Self-Evolving Multi-Agent Systems via Textual Backpropagation
por: Ma, Xiaowen, et al.
Publicado: (2025)
por: Ma, Xiaowen, et al.
Publicado: (2025)
Dr. MAS: Stable Reinforcement Learning for Multi-Agent LLM Systems
por: Feng, Lang, et al.
Publicado: (2026)
por: Feng, Lang, et al.
Publicado: (2026)
The Behavior Gap: Evaluating Zero-shot LLM Agents in Complex Task-Oriented Dialogs
por: Baidya, Avinash, et al.
Publicado: (2025)
por: Baidya, Avinash, et al.
Publicado: (2025)
Learning Evolving Latent Strategies for Multi-Agent Language Systems without Model Fine-Tuning
por: Tang, Wenlong
Publicado: (2025)
por: Tang, Wenlong
Publicado: (2025)
Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective
por: Zhang, Yuheng, et al.
Publicado: (2026)
por: Zhang, Yuheng, et al.
Publicado: (2026)
MASTEST: A LLM-Based Multi-Agent System For RESTful API Tests
por: Han, Xiaoke, et al.
Publicado: (2025)
por: Han, Xiaoke, et al.
Publicado: (2025)
AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
por: Ma, Chang, et al.
Publicado: (2024)
por: Ma, Chang, et al.
Publicado: (2024)
Towards General Computer Control with Hierarchical Agents and Multi-Level Action Spaces
por: Dong, Zihan, et al.
Publicado: (2025)
por: Dong, Zihan, et al.
Publicado: (2025)
AdamS: Momentum Itself Can Be A Normalizer for LLM Pretraining and Post-training
por: Zhang, Huishuai, et al.
Publicado: (2025)
por: Zhang, Huishuai, et al.
Publicado: (2025)
Evaluation and Benchmarking of LLM Agents: A Survey
por: Mohammadi, Mahmoud, et al.
Publicado: (2025)
por: Mohammadi, Mahmoud, et al.
Publicado: (2025)
AgenticPay: A Multi-Agent LLM Negotiation System for Buyer-Seller Transactions
por: Liu, Xianyang, et al.
Publicado: (2026)
por: Liu, Xianyang, et al.
Publicado: (2026)
Reliable Self-Harm Risk Screening via Adaptive Multi-Agent LLM Systems
por: Karnam, Meghana, et al.
Publicado: (2026)
por: Karnam, Meghana, et al.
Publicado: (2026)
On the Trainability of Masked Diffusion Language Models via Blockwise Locality
por: Wang, Yuxiang, et al.
Publicado: (2026)
por: Wang, Yuxiang, et al.
Publicado: (2026)
ComplexityNet: Increasing LLM Inference Efficiency by Learning Task Complexity
por: Bae, Henry, et al.
Publicado: (2023)
por: Bae, Henry, et al.
Publicado: (2023)
CoPS: Empowering LLM Agents with Provable Cross-Task Experience Sharing
por: Yang, Chen, et al.
Publicado: (2024)
por: Yang, Chen, et al.
Publicado: (2024)
Topological Structure Learning Should Be A Research Priority for LLM-Based Multi-Agent Systems
por: Yang, Jiaxi, et al.
Publicado: (2025)
por: Yang, Jiaxi, et al.
Publicado: (2025)
Unraveling the Complexity of Memory in RL Agents: an Approach for Classification and Evaluation
por: Cherepanov, Egor, et al.
Publicado: (2024)
por: Cherepanov, Egor, et al.
Publicado: (2024)
One-step Structure Prediction and Screening for Protein-Ligand Complexes using Multi-Task Geometric Deep Learning
por: He, Kelei, et al.
Publicado: (2024)
por: He, Kelei, et al.
Publicado: (2024)
GoAgent: Group-of-Agents Communication Topology Generation for LLM-based Multi-Agent Systems
por: Chen, Hongjiang, et al.
Publicado: (2026)
por: Chen, Hongjiang, et al.
Publicado: (2026)
Ejemplares similares
-
Heterogeneous Graph Structure Learning through the Lens of Data-generating Processes
por: Jiang, Keyue, et al.
Publicado: (2025) -
Training-Free Message Passing for Learning on Hypergraphs
por: Tang, Bohan, et al.
Publicado: (2024) -
A Markov Random Field model for Hypergraph-based Machine Learning
por: Tang, Bohan, et al.
Publicado: (2023) -
Bures-Wasserstein Flow Matching for Graph Generation
por: Jiang, Keyue, et al.
Publicado: (2025) -
Towards Quantifying Long-Range Interactions in Graph Machine Learning: a Large Graph Dataset and a Measurement
por: Liang, Huidong, et al.
Publicado: (2025)