Benchmarking LLMs' Swarm intelligence
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ruan, Kai, Huang, Mowen, Wen, Ji-Rong, Sun, Hao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LingxiDiagBench: A Multi-Agent Framework for Benchmarking LLMs in Chinese Psychiatric Consultation and Diagnosis
von: Xu, Shihao, et al.
Veröffentlicht: (2026)
von: Xu, Shihao, et al.
Veröffentlicht: (2026)
A Novel Hierarchical Multi-Agent System for Payments Using LLMs
von: Chua, Joon Kiat, et al.
Veröffentlicht: (2026)
von: Chua, Joon Kiat, et al.
Veröffentlicht: (2026)
Agents that Matter: Optimizing Multi-Agent LLMs via Removal-Based Attribution
von: Lu, Mingyu, et al.
Veröffentlicht: (2026)
von: Lu, Mingyu, et al.
Veröffentlicht: (2026)
Collab-Overcooked: Benchmarking and Evaluating Large Language Models as Collaborative Agents
von: Sun, Haochen, et al.
Veröffentlicht: (2025)
von: Sun, Haochen, et al.
Veröffentlicht: (2025)
StoryBox: Collaborative Multi-Agent Simulation for Hybrid Bottom-Up Long-Form Story Generation Using Large Language Models
von: Chen, Zehao, et al.
Veröffentlicht: (2025)
von: Chen, Zehao, et al.
Veröffentlicht: (2025)
Scaling Personality Control in LLMs with Big Five Scaler Prompts
von: Cho, Gunhee, et al.
Veröffentlicht: (2025)
von: Cho, Gunhee, et al.
Veröffentlicht: (2025)
Debate-Driven Multi-Agent LLMs for Phishing Email Detection
von: Nguyen, Ngoc Tuong Vy, et al.
Veröffentlicht: (2025)
von: Nguyen, Ngoc Tuong Vy, et al.
Veröffentlicht: (2025)
Specialists or Generalists? Multi-Agent and Single-Agent LLMs for Essay Grading
von: Idowu, Jamiu Adekunle, et al.
Veröffentlicht: (2026)
von: Idowu, Jamiu Adekunle, et al.
Veröffentlicht: (2026)
Enhancing Value Alignment of LLMs with Multi-agent system and Combinatorial Fusion
von: Wu, Yuanhong, et al.
Veröffentlicht: (2026)
von: Wu, Yuanhong, et al.
Veröffentlicht: (2026)
CLUES: Collaborative High-Quality Data Selection for LLMs via Training Dynamics
von: Zhao, Wanru, et al.
Veröffentlicht: (2025)
von: Zhao, Wanru, et al.
Veröffentlicht: (2025)
From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems
von: Chen, Jiayi, et al.
Veröffentlicht: (2025)
von: Chen, Jiayi, et al.
Veröffentlicht: (2025)
MAS-GPT: Training LLMs to Build LLM-based Multi-Agent Systems
von: Ye, Rui, et al.
Veröffentlicht: (2025)
von: Ye, Rui, et al.
Veröffentlicht: (2025)
Agentic Medical Knowledge Graphs Enhance Medical Question Answering: Bridging the Gap Between LLMs and Evolving Medical Knowledge
von: Rezaei, Mohammad Reza, et al.
Veröffentlicht: (2025)
von: Rezaei, Mohammad Reza, et al.
Veröffentlicht: (2025)
Single-Agent LLMs Outperform Multi-Agent Systems on Multi-Hop Reasoning Under Equal Thinking Token Budgets
von: Tran, Dat, et al.
Veröffentlicht: (2026)
von: Tran, Dat, et al.
Veröffentlicht: (2026)
MedPriv-Bench: Benchmarking the Privacy-Utility Trade-off of Large Language Models in Medical Open-End Question Answering
von: Guan, Shaowei, et al.
Veröffentlicht: (2026)
von: Guan, Shaowei, et al.
Veröffentlicht: (2026)
ReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical Reasoning
von: Sun, Yu, et al.
Veröffentlicht: (2025)
von: Sun, Yu, et al.
Veröffentlicht: (2025)
Assemble Your Crew: Automatic Multi-agent Communication Topology Design via Autoregressive Graph Generation
von: Li, Shiyuan, et al.
Veröffentlicht: (2025)
von: Li, Shiyuan, et al.
Veröffentlicht: (2025)
Mitigating Bias in Queer Representation within Large Language Models: A Collaborative Agent Approach
von: Huang, Tianyi, et al.
Veröffentlicht: (2024)
von: Huang, Tianyi, et al.
Veröffentlicht: (2024)
Fortytwo: Swarm Inference with Peer-Ranked Consensus
von: Larin, Vladyslav, et al.
Veröffentlicht: (2025)
von: Larin, Vladyslav, et al.
Veröffentlicht: (2025)
How Brittle is Agent Safety? Rethinking Agent Risk under Intent Concealment and Task Complexity
von: Ma, Zihan, et al.
Veröffentlicht: (2025)
von: Ma, Zihan, et al.
Veröffentlicht: (2025)
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities
von: Li, Haoming, et al.
Veröffentlicht: (2025)
von: Li, Haoming, et al.
Veröffentlicht: (2025)
Multi-User Large Language Model Agents
von: Yang, Shu, et al.
Veröffentlicht: (2026)
von: Yang, Shu, et al.
Veröffentlicht: (2026)
You Only Align Once: Propagating Cooperative Behaviors in Multi-Agent Systems through Seed Agents
von: Hsing, Nicole, et al.
Veröffentlicht: (2026)
von: Hsing, Nicole, et al.
Veröffentlicht: (2026)
EvoRoute: Experience-Driven Self-Routing LLM Agent Systems
von: Zhang, Guibin, et al.
Veröffentlicht: (2026)
von: Zhang, Guibin, et al.
Veröffentlicht: (2026)
LLMs with Personalities in Multi-issue Negotiation Games
von: Noh, Sean, et al.
Veröffentlicht: (2024)
von: Noh, Sean, et al.
Veröffentlicht: (2024)
Lying with Truths: Open-Channel Multi-Agent Collusion for Belief Manipulation via Generative Montage
von: Hu, Jinwei, et al.
Veröffentlicht: (2026)
von: Hu, Jinwei, et al.
Veröffentlicht: (2026)
GEMA-Score: Granular Explainable Multi-Agent Scoring Framework for Radiology Report Evaluation
von: Zhang, Zhenxuan, et al.
Veröffentlicht: (2025)
von: Zhang, Zhenxuan, et al.
Veröffentlicht: (2025)
Cash or Comfort? How LLMs Value Your Inconvenience
von: Cedro, Mateusz, et al.
Veröffentlicht: (2025)
von: Cedro, Mateusz, et al.
Veröffentlicht: (2025)
Integrating Expert Knowledge into Logical Programs via LLMs
von: Górski, Franciszek, et al.
Veröffentlicht: (2025)
von: Górski, Franciszek, et al.
Veröffentlicht: (2025)
Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning
von: Yan, Sikuan, et al.
Veröffentlicht: (2025)
von: Yan, Sikuan, et al.
Veröffentlicht: (2025)
ColorEcosystem: Powering Personalized, Standardized, and Trustworthy Agentic Service in massive-agent Ecosystem
von: Wu, Fangwen, et al.
Veröffentlicht: (2025)
von: Wu, Fangwen, et al.
Veröffentlicht: (2025)
DSBC : Data Science task Benchmarking with Context engineering
von: Kadiyala, Ram Mohan Rao, et al.
Veröffentlicht: (2025)
von: Kadiyala, Ram Mohan Rao, et al.
Veröffentlicht: (2025)
SciReplicate-Bench: Benchmarking LLMs in Agent-driven Algorithmic Reproduction from Research Papers
von: Xiang, Yanzheng, et al.
Veröffentlicht: (2025)
von: Xiang, Yanzheng, et al.
Veröffentlicht: (2025)
Responsible Agentic AI Requires Explicit Provenance
von: Hu, Jinwei, et al.
Veröffentlicht: (2026)
von: Hu, Jinwei, et al.
Veröffentlicht: (2026)
360$^\circ$REA: Towards A Reusable Experience Accumulation with 360° Assessment for Multi-Agent System
von: Gao, Shen, et al.
Veröffentlicht: (2024)
von: Gao, Shen, et al.
Veröffentlicht: (2024)
From Intention To Implementation: Automating Biomedical Research via LLMs
von: Luo, Yi, et al.
Veröffentlicht: (2024)
von: Luo, Yi, et al.
Veröffentlicht: (2024)
X-MAS: Towards Building Multi-Agent Systems with Heterogeneous LLMs
von: Ye, Rui, et al.
Veröffentlicht: (2025)
von: Ye, Rui, et al.
Veröffentlicht: (2025)
Do Mixed-Vendor Multi-Agent LLMs Improve Clinical Diagnosis?
von: Yuan, Grace Chang, et al.
Veröffentlicht: (2026)
von: Yuan, Grace Chang, et al.
Veröffentlicht: (2026)
AgentArch: A Comprehensive Benchmark to Evaluate Agent Architectures in Enterprise
von: Bogavelli, Tara, et al.
Veröffentlicht: (2025)
von: Bogavelli, Tara, et al.
Veröffentlicht: (2025)
ATOD: An Evaluation Framework and Benchmark for Agentic Task-Oriented Dialogue Systems
von: Zhang, Yifei, et al.
Veröffentlicht: (2026)
von: Zhang, Yifei, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
LingxiDiagBench: A Multi-Agent Framework for Benchmarking LLMs in Chinese Psychiatric Consultation and Diagnosis
von: Xu, Shihao, et al.
Veröffentlicht: (2026) -
A Novel Hierarchical Multi-Agent System for Payments Using LLMs
von: Chua, Joon Kiat, et al.
Veröffentlicht: (2026) -
Agents that Matter: Optimizing Multi-Agent LLMs via Removal-Based Attribution
von: Lu, Mingyu, et al.
Veröffentlicht: (2026) -
Collab-Overcooked: Benchmarking and Evaluating Large Language Models as Collaborative Agents
von: Sun, Haochen, et al.
Veröffentlicht: (2025) -
StoryBox: Collaborative Multi-Agent Simulation for Hybrid Bottom-Up Long-Form Story Generation Using Large Language Models
von: Chen, Zehao, et al.
Veröffentlicht: (2025)