Scalable Inference Architectures for Compound AI Systems: A Production Deployment Study
Fuente:
arXiv
Guardado en:
| Autores principales: | S V, Srikanta Prasad, Arora, Utkarsh |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A Novel Compound AI Model for 6G Networks in 3D Continuum
por: Gravara, Milos, et al.
Publicado: (2025)
por: Gravara, Milos, et al.
Publicado: (2025)
HECATE: An ECS-based Framework for Teaching and Developing Multi-Agent Systems
por: Casals, Arthur, et al.
Publicado: (2025)
por: Casals, Arthur, et al.
Publicado: (2025)
Rethinking AI Hardware: A Three-Layer Cognitive Architecture for Autonomous Agents
por: Chen, Li
Publicado: (2026)
por: Chen, Li
Publicado: (2026)
Counterweights and Complementarities: The Convergence of AI and Blockchain Powering a Decentralized Future
por: Li, Yibai, et al.
Publicado: (2026)
por: Li, Yibai, et al.
Publicado: (2026)
Ontology-Constrained Neural Reasoning in Enterprise Agentic Systems: A Neurosymbolic Architecture for Domain-Grounded AI Agents
por: Tuan, Thanh Luong, et al.
Publicado: (2026)
por: Tuan, Thanh Luong, et al.
Publicado: (2026)
On the Push-Based Asynchronous Federated Learning: A Bias-Correction Aggregation Approach
por: Bai, Jiahui, et al.
Publicado: (2026)
por: Bai, Jiahui, et al.
Publicado: (2026)
pFedDSH: Enabling Knowledge Transfer in Personalized Federated Learning through Data-free Sub-Hypernetwork
por: Nguyen, Thinh, et al.
Publicado: (2025)
por: Nguyen, Thinh, et al.
Publicado: (2025)
Byzantine-Resilient Federated Learning via QUBO-Based Client Selection on Quantum Annealers
por: Ferenczi, Andras, et al.
Publicado: (2026)
por: Ferenczi, Andras, et al.
Publicado: (2026)
Context Engineering: From Prompts to Corporate Multi-Agent Architecture
por: Vishnyakova, Vera V.
Publicado: (2026)
por: Vishnyakova, Vera V.
Publicado: (2026)
Swarm Learning: A Survey of Concepts, Applications, and Trends
por: Shammar, Elham, et al.
Publicado: (2024)
por: Shammar, Elham, et al.
Publicado: (2024)
Intelligent Product 3.0: Decentralised AI Agents and Web3 Intelligence Standards
por: Wong, Alex C. Y., et al.
Publicado: (2025)
por: Wong, Alex C. Y., et al.
Publicado: (2025)
Federated Learning and AI Regulation in the European Union: Who is Responsible? -- An Interdisciplinary Analysis
por: Woisetschläger, Herbert, et al.
Publicado: (2024)
por: Woisetschläger, Herbert, et al.
Publicado: (2024)
PRISM-Consult: A Panel-of-Experts Architecture for Clinician-Aligned Diagnosis
por: Levine, Lionel, et al.
Publicado: (2025)
por: Levine, Lionel, et al.
Publicado: (2025)
Knowledge Equivalence in Digital Twins of Intelligent Systems
por: Zhang, Nan, et al.
Publicado: (2022)
por: Zhang, Nan, et al.
Publicado: (2022)
Mechanical Conscience: A Mathematical Framework for Dependability of Machine Intelligenc
por: Batzorig, Munkhdegerekh, et al.
Publicado: (2026)
por: Batzorig, Munkhdegerekh, et al.
Publicado: (2026)
FSFM: A Biologically-Inspired Framework for Selective Forgetting of Agent Memory
por: Gu, Yingjie, et al.
Publicado: (2026)
por: Gu, Yingjie, et al.
Publicado: (2026)
Multi-Agent LLM Orchestration Achieves Deterministic, High-Quality Decision Support for Incident Response
por: Drammeh, Philip
Publicado: (2025)
por: Drammeh, Philip
Publicado: (2025)
FundaPod: A Multi-Persona Agent Pod Platform with Knowledge Graph Memory for AI-Assisted Fundamental Investment Research
por: Zhu, Di, et al.
Publicado: (2026)
por: Zhu, Di, et al.
Publicado: (2026)
HDP: A Lightweight Cryptographic Protocol for Human Delegation Provenance in Agentic AI Systems
por: Dalugoda, Asiri
Publicado: (2026)
por: Dalugoda, Asiri
Publicado: (2026)
How Machine Learning-Data Driven Replication Strategies Enhance Fault Tolerance in Large-Scale Distributed Systems
por: Murimi, Almond Kiruthu
Publicado: (2025)
por: Murimi, Almond Kiruthu
Publicado: (2025)
Multi-Paradigm Agent Interaction in Practice:A Systematic Analysis of Generator-Evaluator, ReAct Loop,and Adversarial Evaluation in the buddyMe Framework
por: Wang, Xiaohua, et al.
Publicado: (2026)
por: Wang, Xiaohua, et al.
Publicado: (2026)
ChargingBoul: A Competitive Negotiating Agent with Novel Opponent Modeling
por: Shymanski, Joe
Publicado: (2025)
por: Shymanski, Joe
Publicado: (2025)
Advancing Multimodal Agent Reasoning with Long-Term Neuro-Symbolic Memory
por: Jiang, Rongjie, et al.
Publicado: (2026)
por: Jiang, Rongjie, et al.
Publicado: (2026)
Mobile Traffic Prediction at the Edge Through Distributed and Deep Transfer Learning
por: Petrella, Alfredo, et al.
Publicado: (2023)
por: Petrella, Alfredo, et al.
Publicado: (2023)
Security Considerations for Multi-agent Systems
por: Nguyen, Tam, et al.
Publicado: (2026)
por: Nguyen, Tam, et al.
Publicado: (2026)
AMP4EC: Adaptive Model Partitioning Framework for Efficient Deep Learning Inference in Edge Computing Environments
por: Zhang, Guilin, et al.
Publicado: (2025)
por: Zhang, Guilin, et al.
Publicado: (2025)
Secure Decentralized Learning with Blockchain
por: Zhang, Xiaoxue, et al.
Publicado: (2023)
por: Zhang, Xiaoxue, et al.
Publicado: (2023)
Federated Learning Model Aggregation in Heterogenous Aerial and Space Networks
por: Dong, Fan, et al.
Publicado: (2023)
por: Dong, Fan, et al.
Publicado: (2023)
Adaptive GPU Resource Allocation for Multi-Agent Collaborative Reasoning in Serverless Environments
por: Zhang, Guilin, et al.
Publicado: (2025)
por: Zhang, Guilin, et al.
Publicado: (2025)
Token Coherence: Adapting MESI Cache Protocols to Minimize Synchronization Overhead in Multi-Agent LLM Systems
por: Parakhin, Vladyslav
Publicado: (2026)
por: Parakhin, Vladyslav
Publicado: (2026)
Evaluating Temporal Semantic Caching and Workflow Optimization in Agentic Plan-Execute Pipelines
por: Merchant, Alimurtaza Mustafa, et al.
Publicado: (2026)
por: Merchant, Alimurtaza Mustafa, et al.
Publicado: (2026)
Efficient Cloud-Edge-Device Query Execution Based on Collaborative Scan Operator
por: Zhao, Chunyu, et al.
Publicado: (2025)
por: Zhao, Chunyu, et al.
Publicado: (2025)
Impact of Network Topology on Byzantine Resilience in Decentralized Federated Learning
por: Bhattacharya, Siddhartha, et al.
Publicado: (2024)
por: Bhattacharya, Siddhartha, et al.
Publicado: (2024)
Large Language Models for Agentic NetOps and AIOps: Architectures, Evaluation, and Safety
por: Bilal, Muhammad, et al.
Publicado: (2026)
por: Bilal, Muhammad, et al.
Publicado: (2026)
From Multi-Agent Systems and the Semantic Web to Agentic AI: A Unified Narrative of the Web of Agents
por: Petrova, Tatiana, et al.
Publicado: (2025)
por: Petrova, Tatiana, et al.
Publicado: (2025)
$γ(3,4)$ `Attention' in Cognitive Agents: Ontology-Free Knowledge Representations With Promise Theoretic Semantics
por: Burgess, Mark
Publicado: (2025)
por: Burgess, Mark
Publicado: (2025)
PRISM: Perspective Reasoning for Integrated Synthesis and Mediation as a Multi-Perspective Framework for AI Alignment
por: Diamond, Anthony
Publicado: (2025)
por: Diamond, Anthony
Publicado: (2025)
A Taxonomy and Resolution Strategy for Client-Level Disagreements in Federated Learning
por: Rosendal, Daan, et al.
Publicado: (2026)
por: Rosendal, Daan, et al.
Publicado: (2026)
FedSPU: Personalized Federated Learning for Resource-constrained Devices with Stochastic Parameter Update
por: Niu, Ziru, et al.
Publicado: (2024)
por: Niu, Ziru, et al.
Publicado: (2024)
Owner-Harm: A Missing Threat Model for AI Agent Safety
por: Zhang, Dongcheng, et al.
Publicado: (2026)
por: Zhang, Dongcheng, et al.
Publicado: (2026)
Ejemplares similares
-
A Novel Compound AI Model for 6G Networks in 3D Continuum
por: Gravara, Milos, et al.
Publicado: (2025) -
HECATE: An ECS-based Framework for Teaching and Developing Multi-Agent Systems
por: Casals, Arthur, et al.
Publicado: (2025) -
Rethinking AI Hardware: A Three-Layer Cognitive Architecture for Autonomous Agents
por: Chen, Li
Publicado: (2026) -
Counterweights and Complementarities: The Convergence of AI and Blockchain Powering a Decentralized Future
por: Li, Yibai, et al.
Publicado: (2026) -
Ontology-Constrained Neural Reasoning in Enterprise Agentic Systems: A Neurosymbolic Architecture for Domain-Grounded AI Agents
por: Tuan, Thanh Luong, et al.
Publicado: (2026)