Towards a Science of Collective AI: LLM-based Multi-Agent Systems Need a Transition from Blind Trial-and-Error to Rigorous Science
Fuente:
arXiv
Saved in:
| Main Authors: | Fan, Jingru, Liu, Dewen, Dang, Yufan, Li, Huatao, Wang, Yuheng, Liu, Wei, Duan, Feiyu, Ding, Xuanwen, Yao, Shu, Wu, Lin, Shi, Ruijie, Leung, Wai-Shing, Cheng, Yuan, Wei, Zhongyu, Yang, Cheng, Qian, Chen, Liu, Zhiyuan, Sun, Maosong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AgentXRay: White-Boxing Agentic Systems via Workflow Reconstruction
by: Shi, Ruijie, et al.
Published: (2026)
by: Shi, Ruijie, et al.
Published: (2026)
AppCopilot: Toward General, Accurate, Long-Horizon, and Efficient Mobile Agent
by: Fan, Jingru, et al.
Published: (2025)
by: Fan, Jingru, et al.
Published: (2025)
Multi-Agent Collaboration via Evolving Orchestration
by: Dang, Yufan, et al.
Published: (2025)
by: Dang, Yufan, et al.
Published: (2025)
Iterative Experience Refinement of Software-Developing Agents
by: Qian, Chen, et al.
Published: (2024)
by: Qian, Chen, et al.
Published: (2024)
Metashinobius Wei & Liu 2025, gen. nov.
by: Cheng, Yufan, et al.
Published: (2025)
by: Cheng, Yufan, et al.
Published: (2025)
Experiential Co-Learning of Software-Developing Agents
by: Qian, Chen, et al.
Published: (2023)
by: Qian, Chen, et al.
Published: (2023)
LifeSim: Long-Horizon User Life Simulator for Personalized Assistant Evaluation
by: Duan, Feiyu, et al.
Published: (2026)
by: Duan, Feiyu, et al.
Published: (2026)
ProcureGym: A Multi-Agent Markov Game Framework for Modeling National Volume-based Drug Procurement
by: Wang, Jia, et al.
Published: (2026)
by: Wang, Jia, et al.
Published: (2026)
Scaling Large Language Model-based Multi-Agent Collaboration
by: Qian, Chen, et al.
Published: (2024)
by: Qian, Chen, et al.
Published: (2024)
ChatDev: Communicative Agents for Software Development
by: Qian, Chen, et al.
Published: (2023)
by: Qian, Chen, et al.
Published: (2023)
Visual-information-driven model for crowd simulation using temporal convolutional network
by: Liang, Xuanwen, et al.
Published: (2023)
by: Liang, Xuanwen, et al.
Published: (2023)
Simulation of collision avoidance behavior in crowd movement by data-driven approach
by: Liang, Xuanwen, et al.
Published: (2026)
by: Liang, Xuanwen, et al.
Published: (2026)
Cross-Task Experiential Learning on LLM-based Multi-Agent Collaboration
by: Li, Yilong, et al.
Published: (2025)
by: Li, Yilong, et al.
Published: (2025)
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs
by: Ding, Xuanwen, et al.
Published: (2025)
by: Ding, Xuanwen, et al.
Published: (2025)
Science Democratization for Rigor, Relevance, and Resilience
by: Marlen Z. Gonzalez, et al.
Published: (2025)
by: Marlen Z. Gonzalez, et al.
Published: (2025)
Improved visual-information-driven model for crowd simulation and its modular application
by: Liang, Xuanwen, et al.
Published: (2025)
by: Liang, Xuanwen, et al.
Published: (2025)
Seeing No Evil: Blinding Large Vision-Language Models to Safety Instructions via Adversarial Attention Hijacking
by: Li, Jingru, et al.
Published: (2026)
by: Li, Jingru, et al.
Published: (2026)
Enhancing Legal Case Retrieval via Scaling High-quality Synthetic Query-Candidate Pairs
by: Gao, Cheng, et al.
Published: (2024)
by: Gao, Cheng, et al.
Published: (2024)
Enhancing Long-Chain Reasoning Distillation through Error-Aware Self-Reflection
by: Wu, Zhuoyang, et al.
Published: (2025)
by: Wu, Zhuoyang, et al.
Published: (2025)
Representation Learning for Natural Language Processing
by: Liu, Zhiyuan, et al.
Published: (2020)
by: Liu, Zhiyuan, et al.
Published: (2020)
A Review of Light-Field Imaging in Biomedical Sciences
by: Zhao, Ruixuan, et al.
Published: (2025)
by: Zhao, Ruixuan, et al.
Published: (2025)
PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning
by: Zhang, Qiran, et al.
Published: (2026)
by: Zhang, Qiran, et al.
Published: (2026)
TeachMaster: Generative Teaching via Code
by: Wang, Yuheng, et al.
Published: (2025)
by: Wang, Yuheng, et al.
Published: (2025)
Investigate the efficiency of incompressible flow simulations on CPUs and GPUs with BSAMR
by: Liu, Dewen, et al.
Published: (2024)
by: Liu, Dewen, et al.
Published: (2024)
Co-Saving: Resource Aware Multi-Agent Collaboration for Software Development
by: Qiu, Rennai, et al.
Published: (2025)
by: Qiu, Rennai, et al.
Published: (2025)
Optima: Optimizing Effectiveness and Efficiency for LLM-Based Multi-Agent System
by: Chen, Weize, et al.
Published: (2024)
by: Chen, Weize, et al.
Published: (2024)
H-Neurons: On the Existence, Impact, and Origin of Hallucination-Associated Neurons in LLMs
by: Gao, Cheng, et al.
Published: (2025)
by: Gao, Cheng, et al.
Published: (2025)
Blind Graph Matching Using Graph Signals
by: Liu, Hang, et al.
Published: (2023)
by: Liu, Hang, et al.
Published: (2023)
Advancing Quantum Information Science Pre-College Education: The Case for Learning Sciences Collaboration
by: Coelho, Raquel, et al.
Published: (2025)
by: Coelho, Raquel, et al.
Published: (2025)
Establishing Rigorous and Cost-effective Clinical Trials for Artificial Intelligence Models
by: Gao, Wanling, et al.
Published: (2024)
by: Gao, Wanling, et al.
Published: (2024)
StreamProfileBench: A Benchmark for Fine-Grained User Profile Inference in Real-World Streaming Scenarios
by: Wang, Sizhe, et al.
Published: (2026)
by: Wang, Sizhe, et al.
Published: (2026)
Clifford Perturbation Approximation for Quantum Error Mitigation
by: Zhang, Ruiqi, et al.
Published: (2024)
by: Zhang, Ruiqi, et al.
Published: (2024)
States Hidden in Hidden States: LLMs Emerge Discrete State Representations Implicitly
by: Chen, Junhao, et al.
Published: (2024)
by: Chen, Junhao, et al.
Published: (2024)
KG-Infused RAG: Augmenting Corpus-Based RAG with External Knowledge Graphs
by: Wu, Dingjun, et al.
Published: (2025)
by: Wu, Dingjun, et al.
Published: (2025)
Dual‐mode varifocal Moiré metalens for quantitative phase and edge‐enhanced imaging
by: Yi Lian, et al.
Published: (2025)
by: Yi Lian, et al.
Published: (2025)
OViP: Online Vision-Language Preference Learning for VLM Hallucination
by: Liu, Shujun, et al.
Published: (2025)
by: Liu, Shujun, et al.
Published: (2025)
Advancements in Carbon‐Based Piezoelectric Materials: Mechanism, Classification, and Applications in Energy Science
by: Mude Zhu, et al.
Published: (2025)
by: Mude Zhu, et al.
Published: (2025)
BurstAttention: An Efficient Distributed Attention Framework for Extremely Long Sequences
by: Sun, Ao, et al.
Published: (2024)
by: Sun, Ao, et al.
Published: (2024)
BurstEngine: an Efficient Distributed Framework for Training Transformers on Extremely Long Sequences of over 1M Tokens
by: Sun, Ao, et al.
Published: (2025)
by: Sun, Ao, et al.
Published: (2025)
Judge as A Judge: Improving the Evaluation of Retrieval-Augmented Generation through the Judge-Consistency of Large Language Models
by: Liu, Shuliang, et al.
Published: (2025)
by: Liu, Shuliang, et al.
Published: (2025)
Similar Items
-
AgentXRay: White-Boxing Agentic Systems via Workflow Reconstruction
by: Shi, Ruijie, et al.
Published: (2026) -
AppCopilot: Toward General, Accurate, Long-Horizon, and Efficient Mobile Agent
by: Fan, Jingru, et al.
Published: (2025) -
Multi-Agent Collaboration via Evolving Orchestration
by: Dang, Yufan, et al.
Published: (2025) -
Iterative Experience Refinement of Software-Developing Agents
by: Qian, Chen, et al.
Published: (2024) -
Metashinobius Wei & Liu 2025, gen. nov.
by: Cheng, Yufan, et al.
Published: (2025)