Beyond End-to-End: Dynamic Chain Optimization for Private LLM Adaptation on the Edge
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Yebo, Li, Jingguang, Tian, Chunlin, Tam, Kahou, Guo, Zhijiang, Li, Li |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bridging Memory Gaps: Scaling Federated Learning for Heterogeneous Clients
by: Wu, Yebo, et al.
Published: (2024)
by: Wu, Yebo, et al.
Published: (2024)
Memory-Efficient Federated Fine-Tuning of Large Language Models via Layer Pruning
by: Wu, Yebo, et al.
Published: (2025)
by: Wu, Yebo, et al.
Published: (2025)
Elastic Mixture of Rank-Wise Experts for Knowledge Reuse in Federated Fine-Tuning
by: Wu, Yebo, et al.
Published: (2025)
by: Wu, Yebo, et al.
Published: (2025)
Breaking the Memory Wall for Heterogeneous Federated Learning via Model Splitting
by: Tian, Chunlin, et al.
Published: (2024)
by: Tian, Chunlin, et al.
Published: (2024)
A Survey on Federated Fine-tuning of Large Language Models
by: Wu, Yebo, et al.
Published: (2025)
by: Wu, Yebo, et al.
Published: (2025)
Floe: Federated Specialization for Real-Time LLM-SLM Inference
by: Tian, Chunlin, et al.
Published: (2026)
by: Tian, Chunlin, et al.
Published: (2026)
Learning Like Humans: Resource-Efficient Federated Fine-Tuning through Cognitive Developmental Stages
by: Wu, Yebo, et al.
Published: (2025)
by: Wu, Yebo, et al.
Published: (2025)
Heterogeneity-Aware Memory Efficient Federated Learning via Progressive Layer Freezing
by: Yebo, Wu, et al.
Published: (2024)
by: Yebo, Wu, et al.
Published: (2024)
Cost-Effective Edge Data Distribution with End-To-End Delay Guarantees in Edge Computing
by: Shankar, Ravi, et al.
Published: (2025)
by: Shankar, Ravi, et al.
Published: (2025)
LuWu: An End-to-End In-Network Out-of-Core Optimizer for 100B-Scale Model-in-Network Data-Parallel Training on Distributed GPUs
by: Sun, Mo, et al.
Published: (2024)
by: Sun, Mo, et al.
Published: (2024)
End-to-End and Phase-Level Performance Optimization for Hyperledger Fabric
by: Sollu, Pavan, et al.
Published: (2026)
by: Sollu, Pavan, et al.
Published: (2026)
KubeIntellect: A Modular LLM-Orchestrated Agent Framework for End-to-End Kubernetes Management
by: Ardebili, Mohsen Seyedkazemi, et al.
Published: (2025)
by: Ardebili, Mohsen Seyedkazemi, et al.
Published: (2025)
Beyond Model Scale Limits: End-Edge-Cloud Federated Learning with Self-Rectified Knowledge Agglomeration
by: Wu, Zhiyuan, et al.
Published: (2025)
by: Wu, Zhiyuan, et al.
Published: (2025)
ScaleLLM: A Resource-Frugal LLM Serving Framework by Optimizing End-to-End Efficiency
by: Yao, Yuhang, et al.
Published: (2024)
by: Yao, Yuhang, et al.
Published: (2024)
nncase: An End-to-End Compiler for Efficient LLM Deployment on Heterogeneous Storage Architectures
by: Guo, Hui, et al.
Published: (2025)
by: Guo, Hui, et al.
Published: (2025)
Breaking the Memory Wall for Heterogeneous Federated Learning via Progressive Training
by: Wu, Yebo, et al.
Published: (2024)
by: Wu, Yebo, et al.
Published: (2024)
A Proposed End-To-End Principle for Data Commons
by: Grossman, Robert L.
Published: (2025)
by: Grossman, Robert L.
Published: (2025)
Enshrined Proposer Builder Separation in the presence of Maximal Extractable Value
by: Wang, Yitian, et al.
Published: (2026)
by: Wang, Yitian, et al.
Published: (2026)
Saarthi: An End-to-End Intelligent Platform for Optimising Distributed Serverless Workloads
by: Agarwal, Siddharth, et al.
Published: (2025)
by: Agarwal, Siddharth, et al.
Published: (2025)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
by: Ruan, Chaoyi, et al.
Published: (2025)
by: Ruan, Chaoyi, et al.
Published: (2025)
A Survey of End-to-End Modeling for Distributed DNN Training: Workloads, Simulators, and TCO
by: Svedas, Jonas, et al.
Published: (2025)
by: Svedas, Jonas, et al.
Published: (2025)
Exploring Influence Factors on LLM Suitability for No-Code Development of End User IoT Applications
by: Wang, Minghe, et al.
Published: (2025)
by: Wang, Minghe, et al.
Published: (2025)
DynaShard: Secure and Adaptive Blockchain Sharding Protocol with Hybrid Consensus and Dynamic Shard Management
by: Liu, Ao, et al.
Published: (2024)
by: Liu, Ao, et al.
Published: (2024)
Accelerating End-Cloud Collaborative Inference via Near Bubble-free Pipeline Optimization
by: Gao, Luyao, et al.
Published: (2024)
by: Gao, Luyao, et al.
Published: (2024)
Optimizing LLM Inference Throughput via Memory-aware and SLA-constrained Dynamic Batching
by: Pang, Bowen, et al.
Published: (2025)
by: Pang, Bowen, et al.
Published: (2025)
PICE: A Semantic-Driven Progressive Inference System for LLM Serving in Cloud-Edge Networks
by: Zhan, Huiyou, et al.
Published: (2025)
by: Zhan, Huiyou, et al.
Published: (2025)
Odyssey: An End-to-End System for Pareto-Optimal Serverless Query Processing
by: Jesalpura, Shyam, et al.
Published: (2025)
by: Jesalpura, Shyam, et al.
Published: (2025)
Improving the End-to-End Efficiency of Offline Inference for Multi-LLM Applications Based on Sampling and Simulation
by: Fang, Jingzhi, et al.
Published: (2025)
by: Fang, Jingzhi, et al.
Published: (2025)
MSAO: Adaptive Modality Sparsity-Aware Offloading with Edge-Cloud Collaboration for Efficient Multimodal LLM Inference
by: Yang, Zheming, et al.
Published: (2026)
by: Yang, Zheming, et al.
Published: (2026)
ClusterRCA: An End-to-End Approach for Network Fault Localization and Classification for HPC System
by: Sun, Yongqian, et al.
Published: (2025)
by: Sun, Yongqian, et al.
Published: (2025)
Teola: Towards End-to-End Optimization of LLM-based Applications
by: Tan, Xin, et al.
Published: (2024)
by: Tan, Xin, et al.
Published: (2024)
ReaLB: Real-Time Load Balancing for Multimodal MoE Inference
by: Wang, Yingping, et al.
Published: (2026)
by: Wang, Yingping, et al.
Published: (2026)
HybridFlow: Resource-Adaptive Subtask Routing for Efficient Edge-Cloud LLM Inference
by: Dong, Jiangwen, et al.
Published: (2025)
by: Dong, Jiangwen, et al.
Published: (2025)
Efficient Fault Localization in a Cloud Stack Using End-to-End Application Service Topology
by: Mathews, Dhanya R, et al.
Published: (2025)
by: Mathews, Dhanya R, et al.
Published: (2025)
Agglomerative Federated Learning: Empowering Larger Model Training via End-Edge-Cloud Collaboration
by: Wu, Zhiyuan, et al.
Published: (2023)
by: Wu, Zhiyuan, et al.
Published: (2023)
A Pipelined Collaborative Speculative Decoding Framework for Efficient Edge-Cloud LLM Inference
by: Zhang, Yida, et al.
Published: (2026)
by: Zhang, Yida, et al.
Published: (2026)
FunLess: Functions-as-a-Service for Private Edge Cloud Systems
by: De Palma, Giuseppe, et al.
Published: (2024)
by: De Palma, Giuseppe, et al.
Published: (2024)
Online Optimization of DNN Inference Network Utility in Collaborative Edge Computing
by: Li, Rui, et al.
Published: (2024)
by: Li, Rui, et al.
Published: (2024)
Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement
by: Wu, Tian, et al.
Published: (2025)
by: Wu, Tian, et al.
Published: (2025)
LIME:Accelerating Collaborative Lossless LLM Inference on Memory-Constrained Edge Devices
by: Sun, Mingyu, et al.
Published: (2025)
by: Sun, Mingyu, et al.
Published: (2025)
Similar Items
-
Bridging Memory Gaps: Scaling Federated Learning for Heterogeneous Clients
by: Wu, Yebo, et al.
Published: (2024) -
Memory-Efficient Federated Fine-Tuning of Large Language Models via Layer Pruning
by: Wu, Yebo, et al.
Published: (2025) -
Elastic Mixture of Rank-Wise Experts for Knowledge Reuse in Federated Fine-Tuning
by: Wu, Yebo, et al.
Published: (2025) -
Breaking the Memory Wall for Heterogeneous Federated Learning via Model Splitting
by: Tian, Chunlin, et al.
Published: (2024) -
A Survey on Federated Fine-tuning of Large Language Models
by: Wu, Yebo, et al.
Published: (2025)