MAS-Orchestra: Understanding and Improving Multi-Agent Reasoning Through Holistic Orchestration and Controlled Benchmarks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ke, Zixuan, Ming, Yifei, Xu, Austin, Chin, Ryan, Nguyen, Xuan-Phi, Jwalapuram, Prathyusha, Wang, Jiayu, Yavuz, Semih, Xiong, Caiming, Joty, Shafiq |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MAS-ProVe: Understanding the Process Verification of Multi-Agent Systems
von: Venkataramani, Vishal, et al.
Veröffentlicht: (2026)
von: Venkataramani, Vishal, et al.
Veröffentlicht: (2026)
MAS-ZERO: Designing Multi-Agent Systems with Zero Supervision
von: Ke, Zixuan, et al.
Veröffentlicht: (2025)
von: Ke, Zixuan, et al.
Veröffentlicht: (2025)
Demystifying Domain-adaptive Post-training for Financial LLMs
von: Ke, Zixuan, et al.
Veröffentlicht: (2025)
von: Ke, Zixuan, et al.
Veröffentlicht: (2025)
Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows
von: Ming, Yifei, et al.
Veröffentlicht: (2025)
von: Ming, Yifei, et al.
Veröffentlicht: (2025)
SkillOrchestra: Learning to Route Agents via Skill Transfer
von: Wang, Jiayu, et al.
Veröffentlicht: (2026)
von: Wang, Jiayu, et al.
Veröffentlicht: (2026)
Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math
von: Pandit, Shrey, et al.
Veröffentlicht: (2025)
von: Pandit, Shrey, et al.
Veröffentlicht: (2025)
Synthesizing Agentic Data for Web Agents with Progressive Difficulty Enhancement Mechanisms
von: Pandit, Shrey, et al.
Veröffentlicht: (2025)
von: Pandit, Shrey, et al.
Veröffentlicht: (2025)
FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows"
von: Ming, Yifei, et al.
Veröffentlicht: (2024)
von: Ming, Yifei, et al.
Veröffentlicht: (2024)
ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestration
von: Su, Hongjin, et al.
Veröffentlicht: (2025)
von: Su, Hongjin, et al.
Veröffentlicht: (2025)
J4R: Learning to Judge with Equivalent Initial State Group Relative Policy Optimization
von: Xu, Austin, et al.
Veröffentlicht: (2025)
von: Xu, Austin, et al.
Veröffentlicht: (2025)
Least-Loaded Expert Parallelism: Load Balancing An Imbalanced Mixture-of-Experts
von: Nguyen, Xuan-Phi, et al.
Veröffentlicht: (2026)
von: Nguyen, Xuan-Phi, et al.
Veröffentlicht: (2026)
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
von: Xu, Austin, et al.
Veröffentlicht: (2025)
von: Xu, Austin, et al.
Veröffentlicht: (2025)
JudgeRank: Leveraging Large Language Models for Reasoning-Intensive Reranking
von: Niu, Tong, et al.
Veröffentlicht: (2024)
von: Niu, Tong, et al.
Veröffentlicht: (2024)
Beyond Accuracy: Dissecting Mathematical Reasoning for LLMs Under Reinforcement Learning
von: Wang, Jiayu, et al.
Veröffentlicht: (2025)
von: Wang, Jiayu, et al.
Veröffentlicht: (2025)
LiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild
von: Wang, Jiayu, et al.
Veröffentlicht: (2025)
von: Wang, Jiayu, et al.
Veröffentlicht: (2025)
NAACL2025 Tutorial: Adaptation of Large Language Models
von: Ke, Zixuan, et al.
Veröffentlicht: (2025)
von: Ke, Zixuan, et al.
Veröffentlicht: (2025)
Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains
von: Xu, Austin, et al.
Veröffentlicht: (2025)
von: Xu, Austin, et al.
Veröffentlicht: (2025)
CodeXEmbed: A Generalist Embedding Model Family for Multiligual and Multi-task Code Retrieval
von: Liu, Ye, et al.
Veröffentlicht: (2024)
von: Liu, Ye, et al.
Veröffentlicht: (2024)
Modeling Uncertainty and Using Post-fusion as Fallback Improves Retrieval Augmented Generation with LLMs
von: Liu, Ye, et al.
Veröffentlicht: (2023)
von: Liu, Ye, et al.
Veröffentlicht: (2023)
Efficiently Aligned Cross-Lingual Transfer Learning for Conversational Tasks using Prompt-Tuning
von: Tu, Lifu, et al.
Veröffentlicht: (2023)
von: Tu, Lifu, et al.
Veröffentlicht: (2023)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
von: Zhou, Yilun, et al.
Veröffentlicht: (2025)
von: Zhou, Yilun, et al.
Veröffentlicht: (2025)
SFR-RAG: Towards Contextually Faithful LLMs
von: Nguyen, Xuan-Phi, et al.
Veröffentlicht: (2024)
von: Nguyen, Xuan-Phi, et al.
Veröffentlicht: (2024)
Learning from Language Feedback via Variational Policy Distillation
von: Li, Yang, et al.
Veröffentlicht: (2026)
von: Li, Yang, et al.
Veröffentlicht: (2026)
PaperOrchestra: A Multi-Agent Framework for Automated AI Research Paper Writing
von: Song, Yiwen, et al.
Veröffentlicht: (2026)
von: Song, Yiwen, et al.
Veröffentlicht: (2026)
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents
von: Nguyen, Xuan-Phi, et al.
Veröffentlicht: (2025)
von: Nguyen, Xuan-Phi, et al.
Veröffentlicht: (2025)
A Survey of Frontiers in LLM Reasoning: Inference Scaling, Learning to Reason, and Agentic Systems
von: Ke, Zixuan, et al.
Veröffentlicht: (2025)
von: Ke, Zixuan, et al.
Veröffentlicht: (2025)
Tree Search for Simultaneous Move Games via Equilibrium Approximation
von: Yu, Ryan, et al.
Veröffentlicht: (2024)
von: Yu, Ryan, et al.
Veröffentlicht: (2024)
SweRank: Software Issue Localization with Code Ranking
von: Reddy, Revanth Gangi, et al.
Veröffentlicht: (2025)
von: Reddy, Revanth Gangi, et al.
Veröffentlicht: (2025)
Stronger-MAS: Multi-Agent Reinforcement Learning for Collaborative LLMs
von: Zhao, Yujie, et al.
Veröffentlicht: (2025)
von: Zhao, Yujie, et al.
Veröffentlicht: (2025)
ENGRAM: Effective, Lightweight Memory Orchestration for Conversational Agents
von: Patel, Daivik, et al.
Veröffentlicht: (2025)
von: Patel, Daivik, et al.
Veröffentlicht: (2025)
TranSPORTmer: A Holistic Approach to Trajectory Understanding in Multi-Agent Sports
von: Capellera, Guillem, et al.
Veröffentlicht: (2024)
von: Capellera, Guillem, et al.
Veröffentlicht: (2024)
AgentWebBench: Benchmarking Multi-Agent Coordination in Agentic Web
von: Zhong, Shanshan, et al.
Veröffentlicht: (2026)
von: Zhong, Shanshan, et al.
Veröffentlicht: (2026)
VIBEPASS: Can Vibe Coders Really Pass the Vibe Check?
von: Bansal, Srijan, et al.
Veröffentlicht: (2026)
von: Bansal, Srijan, et al.
Veröffentlicht: (2026)
On the external concurrency of current BDI frameworks for MAS
von: Baiardi, Martina, et al.
Veröffentlicht: (2024)
von: Baiardi, Martina, et al.
Veröffentlicht: (2024)
Adaptive Orchestration: Scalable Self-Evolving Multi-Agent Systems
von: Sampath, Sathish, et al.
Veröffentlicht: (2026)
von: Sampath, Sathish, et al.
Veröffentlicht: (2026)
Navigating Complexity: Orchestrated Problem Solving with Multi-Agent LLMs
von: Rasal, Sumedh, et al.
Veröffentlicht: (2024)
von: Rasal, Sumedh, et al.
Veröffentlicht: (2024)
A Multi-LLM Orchestration Engine for Personalized, Context-Rich Assistance
von: Rasal, Sumedh
Veröffentlicht: (2024)
von: Rasal, Sumedh
Veröffentlicht: (2024)
Using Distributed Reinforcement Learning for Resource Orchestration in a Network Slicing Scenario
von: Mason, Federico, et al.
Veröffentlicht: (2021)
von: Mason, Federico, et al.
Veröffentlicht: (2021)
MemFlow: Intent-Driven Memory Orchestration for Small Language Model Agents
von: Chen, Jiayi, et al.
Veröffentlicht: (2026)
von: Chen, Jiayi, et al.
Veröffentlicht: (2026)
Reuse, Don't Recompute: Efficient Large Reasoning Model Inference via Memory Orchestration
von: Patel, Daivik, et al.
Veröffentlicht: (2025)
von: Patel, Daivik, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MAS-ProVe: Understanding the Process Verification of Multi-Agent Systems
von: Venkataramani, Vishal, et al.
Veröffentlicht: (2026) -
MAS-ZERO: Designing Multi-Agent Systems with Zero Supervision
von: Ke, Zixuan, et al.
Veröffentlicht: (2025) -
Demystifying Domain-adaptive Post-training for Financial LLMs
von: Ke, Zixuan, et al.
Veröffentlicht: (2025) -
Helpful Agent Meets Deceptive Judge: Understanding Vulnerabilities in Agentic Workflows
von: Ming, Yifei, et al.
Veröffentlicht: (2025) -
SkillOrchestra: Learning to Route Agents via Skill Transfer
von: Wang, Jiayu, et al.
Veröffentlicht: (2026)