Meta-Learning for Speeding Up Large Model Inference in Decentralized Environments
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Yuzhe, Du, Yipeng, Farhan, Ahmad, Angione, Claudio, Zhao, Yue, Yang, Harry, Johnston, Fielding, Buban, James, Colangelo, Patrick |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Model Agnostic Hybrid Sharding For Heterogeneous Distributed Inference
by: Angione, Claudio, et al.
Published: (2024)
by: Angione, Claudio, et al.
Published: (2024)
Meta-Learning for Speeding Up Large Model Inference in Decentralized Environments
by: Du, Yipeng, et al.
Published: (2025)
by: Du, Yipeng, et al.
Published: (2025)
Parallax: Efficient LLM Inference Service over Decentralized Environment
by: Tong, Chris, et al.
Published: (2025)
by: Tong, Chris, et al.
Published: (2025)
ExClique: An Express Consensus Algorithm for High-Speed Transaction Process in Blockchains
by: Zhao, Chonghe, et al.
Published: (2025)
by: Zhao, Chonghe, et al.
Published: (2025)
HexGen: Generative Inference of Large Language Model over Heterogeneous Environment
by: Jiang, Youhe, et al.
Published: (2023)
by: Jiang, Youhe, et al.
Published: (2023)
Lattica: A Decentralized Cross-NAT Communication Framework for Scalable AI Inference and Training
by: Yang, Ween, et al.
Published: (2025)
by: Yang, Ween, et al.
Published: (2025)
Bandwidth-Aware Network Topology Optimization for Decentralized Learning
by: Shen, Yipeng, et al.
Published: (2025)
by: Shen, Yipeng, et al.
Published: (2025)
Towards Secure and Private AI: A Framework for Decentralized Inference
by: Zhang, Hongyang, et al.
Published: (2024)
by: Zhang, Hongyang, et al.
Published: (2024)
Trusted Execution Environment for Decentralized Process Mining
by: Goretti, Valerio, et al.
Published: (2023)
by: Goretti, Valerio, et al.
Published: (2023)
λScale: Enabling Fast Scaling for Serverless Large Language Model Inference
by: Yu, Minchen, et al.
Published: (2025)
by: Yu, Minchen, et al.
Published: (2025)
Decentralized LLM Inference over Edge Networks with Energy Harvesting
by: Khoshsirat, Aria, et al.
Published: (2024)
by: Khoshsirat, Aria, et al.
Published: (2024)
A Decentralized Root Cause Localization Approach for Edge Computing Environments
by: Fernando, Duneesha, et al.
Published: (2025)
by: Fernando, Duneesha, et al.
Published: (2025)
CONFINE: Preserving Data Secrecy in Decentralized Process Mining
by: Goretti, Valerio, et al.
Published: (2024)
by: Goretti, Valerio, et al.
Published: (2024)
ParvaGPU: Efficient Spatial GPU Sharing for Large-Scale DNN Inference in Cloud Environments
by: Lee, Munkyu, et al.
Published: (2024)
by: Lee, Munkyu, et al.
Published: (2024)
Collaborative Inference for Large Models with Task Offloading and Early Exiting
by: Xie, Zuan, et al.
Published: (2024)
by: Xie, Zuan, et al.
Published: (2024)
GoodSpeed: Optimizing Fair Goodput with Adaptive Speculative Decoding in Distributed Edge Inference
by: Tran, Phuong, et al.
Published: (2025)
by: Tran, Phuong, et al.
Published: (2025)
ACE-GNN: Adaptive GNN Co-Inference with System-Aware Scheduling in Dynamic Edge Environments
by: Zhou, Ao, et al.
Published: (2025)
by: Zhou, Ao, et al.
Published: (2025)
Environment-Aware Dynamic Pruning for Pipelined Edge Inference
by: O'Quinn, Austin, et al.
Published: (2025)
by: O'Quinn, Austin, et al.
Published: (2025)
DFedSat: Communication-Efficient and Robust Decentralized Federated Learning for LEO Satellite Constellations
by: Yang, Minghao, et al.
Published: (2024)
by: Yang, Minghao, et al.
Published: (2024)
D-VRE: From a Jupyter-enabled Private Research Environment to Decentralized Collaborative Research Ecosystem
by: Wang, Yuandou, et al.
Published: (2024)
by: Wang, Yuandou, et al.
Published: (2024)
A TEE-based Approach for Preserving Data Secrecy in Process Mining with Decentralized Sources
by: Basile, Davide, et al.
Published: (2026)
by: Basile, Davide, et al.
Published: (2026)
MoEntwine: Unleashing the Potential of Wafer-scale Chips for Large-scale Expert Parallel Inference
by: Tang, Xinru, et al.
Published: (2025)
by: Tang, Xinru, et al.
Published: (2025)
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees
by: Ma, Chenxiang, et al.
Published: (2025)
by: Ma, Chenxiang, et al.
Published: (2025)
Minions: Accelerating Large Language Model Inference with Aggregated Speculative Execution
by: Wang, Siqi, et al.
Published: (2024)
by: Wang, Siqi, et al.
Published: (2024)
HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment
by: Jiang, Youhe, et al.
Published: (2025)
by: Jiang, Youhe, et al.
Published: (2025)
Scaling Up Throughput-oriented LLM Inference Applications on Heterogeneous Opportunistic GPU Clusters with Pervasive Context Management
by: Phung, Thanh Son, et al.
Published: (2025)
by: Phung, Thanh Son, et al.
Published: (2025)
Resource Allocation of Industry 4.0 Micro-Service Applications across Serverless Fog Federation
by: Hussain, Razin Farhan, et al.
Published: (2024)
by: Hussain, Razin Farhan, et al.
Published: (2024)
Litmus: Fair Pricing for Serverless Computing
by: Pei, Qi, et al.
Published: (2024)
by: Pei, Qi, et al.
Published: (2024)
Loki: A System for Serving ML Inference Pipelines with Hardware and Accuracy Scaling
by: Ahmad, Sohaib, et al.
Published: (2024)
by: Ahmad, Sohaib, et al.
Published: (2024)
Decentralized Semantic Federated Learning for Real-Time Public Safety Tasks: Challenges, Methods, and Directions
by: Li, Baosheng, et al.
Published: (2025)
by: Li, Baosheng, et al.
Published: (2025)
PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications
by: Du, Kuntai, et al.
Published: (2025)
by: Du, Kuntai, et al.
Published: (2025)
Containerization in Multi-Cloud Environment: Roles, Strategies, Challenges, and Solutions for Effective Implementation
by: Waseem, Muhammad, et al.
Published: (2024)
by: Waseem, Muhammad, et al.
Published: (2024)
Speculative Decoding in Decentralized LLM Inference: Turning Communication Latency into Computation Throughput
by: Song, Jingwei, et al.
Published: (2025)
by: Song, Jingwei, et al.
Published: (2025)
Communication-Efficient and Privacy-Preserving Decentralized Meta-Learning
by: Yang, Hansi, et al.
Published: (2024)
by: Yang, Hansi, et al.
Published: (2024)
Speeding up Model Loading with fastsafetensors
by: Yoshimura, Takeshi, et al.
Published: (2025)
by: Yoshimura, Takeshi, et al.
Published: (2025)
Rorqual: Speeding up Narwhal with TEEs
by: Freitas, Luciano, et al.
Published: (2024)
by: Freitas, Luciano, et al.
Published: (2024)
PolyLink: A Blockchain Based Decentralized Edge AI Platform for LLM Inference
by: Liu, Hongbo, et al.
Published: (2025)
by: Liu, Hongbo, et al.
Published: (2025)
SLO-Aware Scheduling for Large Language Model Inferences
by: Huang, Jinqi, et al.
Published: (2025)
by: Huang, Jinqi, et al.
Published: (2025)
Vault: Decentralized Storage Made Durable
by: Sun, Guangda, et al.
Published: (2023)
by: Sun, Guangda, et al.
Published: (2023)
Shelby: Decentralized Storage Designed to Serve
by: Goren, Guy, et al.
Published: (2025)
by: Goren, Guy, et al.
Published: (2025)
Similar Items
-
Model Agnostic Hybrid Sharding For Heterogeneous Distributed Inference
by: Angione, Claudio, et al.
Published: (2024) -
Meta-Learning for Speeding Up Large Model Inference in Decentralized Environments
by: Du, Yipeng, et al.
Published: (2025) -
Parallax: Efficient LLM Inference Service over Decentralized Environment
by: Tong, Chris, et al.
Published: (2025) -
ExClique: An Express Consensus Algorithm for High-Speed Transaction Process in Blockchains
by: Zhao, Chonghe, et al.
Published: (2025) -
HexGen: Generative Inference of Large Language Model over Heterogeneous Environment
by: Jiang, Youhe, et al.
Published: (2023)