Hive: A Multi-Agent Infrastructure for Algorithm- and Task-Level Scaling
Fuente:
arXiv
Salvato in:
| Autori principali: | Luo, Zizhang, Luo, Yuhao, Xiao, Youwei, Xu, Yansong, Guo, Runlin, Liang, Yun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
FedMon: Federated eBPF Monitoring for Distributed Anomaly Detection in Multi-Cluster Cloud Environments
di: Zehra, Sehar, et al.
Pubblicazione: (2025)
di: Zehra, Sehar, et al.
Pubblicazione: (2025)
Flash-Fusion: Enabling Expressive, Low-Latency Queries on IoT Sensor Streams with LLMs
di: Patherya, Kausar, et al.
Pubblicazione: (2025)
di: Patherya, Kausar, et al.
Pubblicazione: (2025)
DAGER: Exact Gradient Inversion for Large Language Models
di: Petrov, Ivo, et al.
Pubblicazione: (2024)
di: Petrov, Ivo, et al.
Pubblicazione: (2024)
GraphBit: A Graph-based Agentic Framework for Non-Linear Agent Orchestration
di: Sarker, Yeahia, et al.
Pubblicazione: (2026)
di: Sarker, Yeahia, et al.
Pubblicazione: (2026)
CooperLLM: Cloud-Edge-End Cooperative Federated Fine-tuning for LLMs via ZOO-based Gradient Correction
di: Sun, He, et al.
Pubblicazione: (2026)
di: Sun, He, et al.
Pubblicazione: (2026)
Benchmarking Federated Learning for Throughput Prediction in 5G Live Streaming Applications
di: Dutta, Yuvraj, et al.
Pubblicazione: (2025)
di: Dutta, Yuvraj, et al.
Pubblicazione: (2025)
Prima.cpp: Fast 30-70B LLM Inference on Heterogeneous and Low-Resource Home Clusters
di: Li, Zonghang, et al.
Pubblicazione: (2025)
di: Li, Zonghang, et al.
Pubblicazione: (2025)
AAFLOW: Scalable Patterns for Agentic AI Workflows
di: Sarker, Arup Kumar, et al.
Pubblicazione: (2026)
di: Sarker, Arup Kumar, et al.
Pubblicazione: (2026)
Comparative Analysis of Large Language Model Inference Serving Systems: A Performance Study of vLLM and HuggingFace TGI
di: Kolluru, Saicharan
Pubblicazione: (2025)
di: Kolluru, Saicharan
Pubblicazione: (2025)
AIvailable: A Software-Defined Architecture for LLM-as-a-Service on Heterogeneous and Legacy GPUs
di: Antunes, Pedro, et al.
Pubblicazione: (2025)
di: Antunes, Pedro, et al.
Pubblicazione: (2025)
POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference
di: Kamath, Aditya K, et al.
Pubblicazione: (2024)
di: Kamath, Aditya K, et al.
Pubblicazione: (2024)
Augmenting the FedProx Algorithm by Minimizing Convergence
di: Sarkar, Anomitra, et al.
Pubblicazione: (2024)
di: Sarkar, Anomitra, et al.
Pubblicazione: (2024)
Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project
di: Penke, Carolin, et al.
Pubblicazione: (2025)
di: Penke, Carolin, et al.
Pubblicazione: (2025)
A Survey on Parallel Text Generation: From Parallel Decoding to Diffusion Language Models
di: Zhang, Lingzhe, et al.
Pubblicazione: (2025)
di: Zhang, Lingzhe, et al.
Pubblicazione: (2025)
Towards Building Private LLMs: Exploring Multi-Node Expert Parallelism on Apple Silicon for Mixture-of-Experts Large Language Model
di: Chen, Mu-Chi, et al.
Pubblicazione: (2025)
di: Chen, Mu-Chi, et al.
Pubblicazione: (2025)
Worldwide Federated Training of Language Models
di: Iacob, Alex, et al.
Pubblicazione: (2024)
di: Iacob, Alex, et al.
Pubblicazione: (2024)
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
di: Jo, Myeong Jun
Pubblicazione: (2026)
di: Jo, Myeong Jun
Pubblicazione: (2026)
Architecture-Aware LLM Inference Optimization on AMD Instinct GPUs: A Comprehensive Benchmark and Deployment Study
di: Georgiou, Athos
Pubblicazione: (2026)
di: Georgiou, Athos
Pubblicazione: (2026)
Federated Learning for Deforestation Detection: A Distributed Approach with Satellite Imagery
di: Dutta, Yuvraj, et al.
Pubblicazione: (2025)
di: Dutta, Yuvraj, et al.
Pubblicazione: (2025)
Parameter-Efficient and Personalized Federated Training of Generative Models at the Edge
di: Khan, Kabir, et al.
Pubblicazione: (2025)
di: Khan, Kabir, et al.
Pubblicazione: (2025)
Latency and Cost of Multi-Agent Intelligent Tutoring at Scale
di: Elhaimeur, Iizalaarab, et al.
Pubblicazione: (2026)
di: Elhaimeur, Iizalaarab, et al.
Pubblicazione: (2026)
Nanoscaling Floating-Point (NxFP): NanoMantissa, Adaptive Microexponents, and Code Recycling for Direct-Cast Compression of Large Language Models
di: Lo, Yun-Chen, et al.
Pubblicazione: (2024)
di: Lo, Yun-Chen, et al.
Pubblicazione: (2024)
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended
di: Kamath, Aditya K, et al.
Pubblicazione: (2026)
di: Kamath, Aditya K, et al.
Pubblicazione: (2026)
DSDE: Dynamic Speculative Decoding with KLD Stability for Real-World Serving
di: Yang, Mingyu, et al.
Pubblicazione: (2025)
di: Yang, Mingyu, et al.
Pubblicazione: (2025)
ATTNChecker: Highly-Optimized Fault Tolerant Attention for Large Language Model Training
di: Liang, Yuhang, et al.
Pubblicazione: (2024)
di: Liang, Yuhang, et al.
Pubblicazione: (2024)
MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices
di: Shakerdargah, Mohammadali, et al.
Pubblicazione: (2024)
di: Shakerdargah, Mohammadali, et al.
Pubblicazione: (2024)
Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference
di: Ganjihal, Sanjeev Rao
Pubblicazione: (2026)
di: Ganjihal, Sanjeev Rao
Pubblicazione: (2026)
Addressing tokens dynamic generation, propagation, storage and renewal to secure the GlideinWMS pilot based jobs and system
di: Coimbra, Bruno Moreira, et al.
Pubblicazione: (2025)
di: Coimbra, Bruno Moreira, et al.
Pubblicazione: (2025)
MeanCache: User-Centric Semantic Caching for LLM Web Services
di: Gill, Waris, et al.
Pubblicazione: (2024)
di: Gill, Waris, et al.
Pubblicazione: (2024)
Token Arena: A Continuous Benchmark Unifying Energy and Cognition in AI Inference
di: Gao, Yuxuan, et al.
Pubblicazione: (2026)
di: Gao, Yuxuan, et al.
Pubblicazione: (2026)
Scalable Engine and the Performance of Different LLM Models in a SLURM based HPC architecture
di: Luiz, Anderson de Lima, et al.
Pubblicazione: (2025)
di: Luiz, Anderson de Lima, et al.
Pubblicazione: (2025)
Cost Trade-offs of Reasoning and Non-Reasoning Large Language Models in Text-to-SQL
di: Deochake, Saurabh, et al.
Pubblicazione: (2025)
di: Deochake, Saurabh, et al.
Pubblicazione: (2025)
Parallelization Strategies for Dense LLM Deployment: Navigating Through Application-Specific Tradeoffs and Bottlenecks
di: Topcu, Burak, et al.
Pubblicazione: (2026)
di: Topcu, Burak, et al.
Pubblicazione: (2026)
Chronicals: A High-Performance Framework for LLM Fine-Tuning with 3.51x Speedup over Unsloth
di: Nair, Arjun S.
Pubblicazione: (2026)
di: Nair, Arjun S.
Pubblicazione: (2026)
Spark-LLM-Eval: A Distributed Framework for Statistically Rigorous Large Language Model Evaluation
di: Mitra, Subhadip
Pubblicazione: (2026)
di: Mitra, Subhadip
Pubblicazione: (2026)
StepCache: Step-Level Reuse with Lightweight Verification and Selective Patching for LLM Serving
di: Nouri, Azam
Pubblicazione: (2026)
di: Nouri, Azam
Pubblicazione: (2026)
Kant: An Efficient Unified Scheduling System for Large-Scale AI Clusters
di: Zeng, Lingling, et al.
Pubblicazione: (2025)
di: Zeng, Lingling, et al.
Pubblicazione: (2025)
De-DSI: Decentralised Differentiable Search Index
di: Neague, Petru, et al.
Pubblicazione: (2024)
di: Neague, Petru, et al.
Pubblicazione: (2024)
Towards Message Brokers for Generative AI: Survey, Challenges, and Opportunities
di: Saleh, Alaa, et al.
Pubblicazione: (2023)
di: Saleh, Alaa, et al.
Pubblicazione: (2023)
Combining Serverless and High-Performance Computing Paradigms to support ML Data-Intensive Applications
di: Staylor, Mills, et al.
Pubblicazione: (2025)
di: Staylor, Mills, et al.
Pubblicazione: (2025)
Documenti analoghi
-
FedMon: Federated eBPF Monitoring for Distributed Anomaly Detection in Multi-Cluster Cloud Environments
di: Zehra, Sehar, et al.
Pubblicazione: (2025) -
Flash-Fusion: Enabling Expressive, Low-Latency Queries on IoT Sensor Streams with LLMs
di: Patherya, Kausar, et al.
Pubblicazione: (2025) -
DAGER: Exact Gradient Inversion for Large Language Models
di: Petrov, Ivo, et al.
Pubblicazione: (2024) -
GraphBit: A Graph-based Agentic Framework for Non-Linear Agent Orchestration
di: Sarker, Yeahia, et al.
Pubblicazione: (2026) -
CooperLLM: Cloud-Edge-End Cooperative Federated Fine-tuning for LLMs via ZOO-based Gradient Correction
di: Sun, He, et al.
Pubblicazione: (2026)