ORXE: Orchestrating Experts for Dynamically Configurable Efficiency

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Qingyuan, Wang, Guoxin, Cardiff, Barry, John, Deepu
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916726019457024
author Wang, Qingyuan
Wang, Guoxin
Cardiff, Barry
John, Deepu
author_facet Wang, Qingyuan
Wang, Guoxin
Cardiff, Barry
John, Deepu
contents This paper presents ORXE, a modular and adaptable framework for achieving real-time configurable efficiency in AI models. By leveraging a collection of pre-trained experts with diverse computational costs and performance levels, ORXE dynamically adjusts inference pathways based on the complexity of input samples. Unlike conventional approaches that require complex metamodel training, ORXE achieves high efficiency and flexibility without complicating the development process. The proposed system utilizes a confidence-based gating mechanism to allocate appropriate computational resources for each input. ORXE also supports adjustments to the preference between inference cost and prediction performance across a wide range during runtime. We implemented a training-free ORXE system for image classification tasks, evaluating its efficiency and accuracy across various devices. The results demonstrate that ORXE achieves superior performance compared to individual experts and other dynamic models in most cases. This approach can be extended to other applications, providing a scalable solution for diverse real-world deployment scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2505_04850
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ORXE: Orchestrating Experts for Dynamically Configurable Efficiency
Wang, Qingyuan
Wang, Guoxin
Cardiff, Barry
John, Deepu
Computer Vision and Pattern Recognition
This paper presents ORXE, a modular and adaptable framework for achieving real-time configurable efficiency in AI models. By leveraging a collection of pre-trained experts with diverse computational costs and performance levels, ORXE dynamically adjusts inference pathways based on the complexity of input samples. Unlike conventional approaches that require complex metamodel training, ORXE achieves high efficiency and flexibility without complicating the development process. The proposed system utilizes a confidence-based gating mechanism to allocate appropriate computational resources for each input. ORXE also supports adjustments to the preference between inference cost and prediction performance across a wide range during runtime. We implemented a training-free ORXE system for image classification tasks, evaluating its efficiency and accuracy across various devices. The results demonstrate that ORXE achieves superior performance compared to individual experts and other dynamic models in most cases. This approach can be extended to other applications, providing a scalable solution for diverse real-world deployment scenarios.
title ORXE: Orchestrating Experts for Dynamically Configurable Efficiency
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.04850