Bench-CoE: a Framework for Collaboration of Experts from Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yuanshuai, Zhang, Xingjian, Zhao, Jinkun, Wen, Siwei, Feng, Peilin, Liao, Shuhao, Huang, Lei, Wu, Wenjun |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CoE-Ops: Collaboration of LLM-based Experts for AIOps Question-Answering
by: Zhao, Jinkun, et al.
Published: (2025)
by: Zhao, Jinkun, et al.
Published: (2025)
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity
by: Zhang, Xingjian, et al.
Published: (2025)
by: Zhang, Xingjian, et al.
Published: (2025)
CoServe: Efficient Collaboration-of-Experts (CoE) Model Inference with Limited Memory
by: Suo, Jiashun, et al.
Published: (2025)
by: Suo, Jiashun, et al.
Published: (2025)
Sci-CoE: Co-evolving Scientific Reasoning LLMs via Geometric Consensus with Sparse Supervision
by: He, Xiaohan, et al.
Published: (2026)
by: He, Xiaohan, et al.
Published: (2026)
CoE: Collaborative Entropy for Uncertainty Quantification in Agentic Multi-LLM Systems
by: Sun, Kangkang, et al.
Published: (2026)
by: Sun, Kangkang, et al.
Published: (2026)
DCA-Bench: A Benchmark for Dataset Curation Agents
by: Huang, Benhao, et al.
Published: (2024)
by: Huang, Benhao, et al.
Published: (2024)
AQA-Bench: An Interactive Benchmark for Evaluating LLMs' Sequential Reasoning Ability
by: Yang, Siwei, et al.
Published: (2024)
by: Yang, Siwei, et al.
Published: (2024)
CoE: Chain-of-Explanation via Automatic Visual Concept Circuit Description and Polysemanticity Quantification
by: Yu, Wenlong, et al.
Published: (2025)
by: Yu, Wenlong, et al.
Published: (2025)
HMACE: Heterogeneous Multi-Agent Collaborative Evolution for Combinatorial Optimization
by: Yan, Yuping, et al.
Published: (2026)
by: Yan, Yuping, et al.
Published: (2026)
TopoEvo: A Topology-Aware Self-Evolving Multi-Agent Framework for Root Cause Analysis in Microservices
by: Wang, Junle, et al.
Published: (2026)
by: Wang, Junle, et al.
Published: (2026)
LLMRouterBench: A Massive Benchmark and Unified Framework for LLM Routing
by: Li, Hao, et al.
Published: (2026)
by: Li, Hao, et al.
Published: (2026)
Enhanced Recommendation Combining Collaborative Filtering and Large Language Models
by: Lin, Xueting, et al.
Published: (2024)
by: Lin, Xueting, et al.
Published: (2024)
TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning
by: Zhang, Xingjian, et al.
Published: (2025)
by: Zhang, Xingjian, et al.
Published: (2025)
CoCoReviewBench: A Completeness- and Correctness-Oriented Benchmark for AI Reviewers
by: Deng, Hexuan, et al.
Published: (2026)
by: Deng, Hexuan, et al.
Published: (2026)
MCCE: A Framework for Multi-LLM Collaborative Co-Evolution
by: Ran, Nian, et al.
Published: (2025)
by: Ran, Nian, et al.
Published: (2025)
SoundnessBench: A Soundness Benchmark for Neural Network Verifiers
by: Zhou, Xingjian, et al.
Published: (2024)
by: Zhou, Xingjian, et al.
Published: (2024)
FaithCoT-Bench: Benchmarking Instance-Level Faithfulness of Chain-of-Thought Reasoning
by: Shen, Xu, et al.
Published: (2025)
by: Shen, Xu, et al.
Published: (2025)
EmboCoach-Bench: Benchmarking AI Agents on Developing Embodied Robots
by: Lei, Zixing, et al.
Published: (2026)
by: Lei, Zixing, et al.
Published: (2026)
Web-Bench: A LLM Code Benchmark Based on Web Standards and Frameworks
by: Xu, Kai, et al.
Published: (2025)
by: Xu, Kai, et al.
Published: (2025)
GlobalDentBench: A Multinational Benchmark for Evaluating LLM Clinical Reasoning in Dentistry with Expert Calibration
by: Zhao, Junjie, et al.
Published: (2026)
by: Zhao, Junjie, et al.
Published: (2026)
PBT-Bench: Benchmarking AI Agents on Property-Based Testing
by: Jing, Lucas, et al.
Published: (2026)
by: Jing, Lucas, et al.
Published: (2026)
MedAgentsBench: Benchmarking Thinking Models and Agent Frameworks for Complex Medical Reasoning
by: Tang, Xiangru, et al.
Published: (2025)
by: Tang, Xiangru, et al.
Published: (2025)
TriCon-SF: A Triple-Shuffle and Contribution-Aware Serial Federated Learning Framework for Heterogeneous Healthcare Data
by: Yan, Yuping, et al.
Published: (2025)
by: Yan, Yuping, et al.
Published: (2025)
CCoE: A Compact and Efficient LLM Framework with Multi-Expert Collaboration for Resource-Limited Settings
by: Huang, Shaomang, et al.
Published: (2024)
by: Huang, Shaomang, et al.
Published: (2024)
Graph Neural Network Framework for Sentiment Analysis Using Syntactic Feature
by: Wu, Linxiao, et al.
Published: (2024)
by: Wu, Linxiao, et al.
Published: (2024)
CoMoE: Collaborative Optimization of Expert Aggregation and Offloading for MoE-based LLMs at Edge
by: Li, Muqing, et al.
Published: (2025)
by: Li, Muqing, et al.
Published: (2025)
Advancing ESG Intelligence: An Expert-level Agent and Comprehensive Benchmark for Sustainable Finance
by: Zhao, Yilei, et al.
Published: (2026)
by: Zhao, Yilei, et al.
Published: (2026)
CCrepairBench: A High-Fidelity Benchmark and Reinforcement Learning Framework for C++ Compilation Repair
by: Sun, Weixuan, et al.
Published: (2025)
by: Sun, Weixuan, et al.
Published: (2025)
MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM Agents
by: Zhang, Dongsen, et al.
Published: (2025)
by: Zhang, Dongsen, et al.
Published: (2025)
ComfyBench: Benchmarking LLM-based Agents in ComfyUI for Autonomously Designing Collaborative AI Systems
by: Xue, Xiangyuan, et al.
Published: (2024)
by: Xue, Xiangyuan, et al.
Published: (2024)
iWorld-Bench: A Benchmark for Interactive World Models with a Unified Action Generation Framework
by: Fang, Jianjie, et al.
Published: (2026)
by: Fang, Jianjie, et al.
Published: (2026)
CoCo-Bench: A Comprehensive Code Benchmark For Multi-task Large Language Model Evaluation
by: Yin, Wenjing, et al.
Published: (2025)
by: Yin, Wenjing, et al.
Published: (2025)
AlignBench: Benchmarking Chinese Alignment of Large Language Models
by: Liu, Xiao, et al.
Published: (2023)
by: Liu, Xiao, et al.
Published: (2023)
SIGMA: Sheaf-Informed Geometric Multi-Agent Pathfinding
by: Liao, Shuhao, et al.
Published: (2025)
by: Liao, Shuhao, et al.
Published: (2025)
ColorBench: Benchmarking Mobile Agents with Graph-Structured Framework for Complex Long-Horizon Tasks
by: Song, Yuanyi, et al.
Published: (2025)
by: Song, Yuanyi, et al.
Published: (2025)
SKA-Bench: A Fine-Grained Benchmark for Evaluating Structured Knowledge Understanding of LLMs
by: Liu, Zhiqiang, et al.
Published: (2025)
by: Liu, Zhiqiang, et al.
Published: (2025)
LoCoT2V-Bench: Benchmarking Long-Form and Complex Text-to-Video Generation
by: Zheng, Xiangqing, et al.
Published: (2025)
by: Zheng, Xiangqing, et al.
Published: (2025)
WritingBench: A Comprehensive Benchmark for Generative Writing
by: Wu, Yuning, et al.
Published: (2025)
by: Wu, Yuning, et al.
Published: (2025)
Health-ORSC-Bench: A Benchmark for Measuring Over-Refusal and Safety Completion in Health Context
by: Zhang, Zhihao, et al.
Published: (2026)
by: Zhang, Zhihao, et al.
Published: (2026)
Requesting Expert Reasoning: Augmenting LLM Agents with Learned Collaborative Intervention
by: Wang, Zhiming, et al.
Published: (2026)
by: Wang, Zhiming, et al.
Published: (2026)
Similar Items
-
CoE-Ops: Collaboration of LLM-based Experts for AIOps Question-Answering
by: Zhao, Jinkun, et al.
Published: (2025) -
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity
by: Zhang, Xingjian, et al.
Published: (2025) -
CoServe: Efficient Collaboration-of-Experts (CoE) Model Inference with Limited Memory
by: Suo, Jiashun, et al.
Published: (2025) -
Sci-CoE: Co-evolving Scientific Reasoning LLMs via Geometric Consensus with Sparse Supervision
by: He, Xiaohan, et al.
Published: (2026) -
CoE: Collaborative Entropy for Uncertainty Quantification in Agentic Multi-LLM Systems
by: Sun, Kangkang, et al.
Published: (2026)