CooperLLM: Cloud-Edge-End Cooperative Federated Fine-tuning for LLMs via ZOO-based Gradient Correction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sun, He, Zhou, Jinrui, Li, Li, Xiao, Mingjun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FedMon: Federated eBPF Monitoring for Distributed Anomaly Detection in Multi-Cluster Cloud Environments
von: Zehra, Sehar, et al.
Veröffentlicht: (2025)
von: Zehra, Sehar, et al.
Veröffentlicht: (2025)
AAFLOW: Scalable Patterns for Agentic AI Workflows
von: Sarker, Arup Kumar, et al.
Veröffentlicht: (2026)
von: Sarker, Arup Kumar, et al.
Veröffentlicht: (2026)
Benchmarking Federated Learning for Throughput Prediction in 5G Live Streaming Applications
von: Dutta, Yuvraj, et al.
Veröffentlicht: (2025)
von: Dutta, Yuvraj, et al.
Veröffentlicht: (2025)
DAGER: Exact Gradient Inversion for Large Language Models
von: Petrov, Ivo, et al.
Veröffentlicht: (2024)
von: Petrov, Ivo, et al.
Veröffentlicht: (2024)
Augmenting the FedProx Algorithm by Minimizing Convergence
von: Sarkar, Anomitra, et al.
Veröffentlicht: (2024)
von: Sarkar, Anomitra, et al.
Veröffentlicht: (2024)
Parameter-Efficient and Personalized Federated Training of Generative Models at the Edge
von: Khan, Kabir, et al.
Veröffentlicht: (2025)
von: Khan, Kabir, et al.
Veröffentlicht: (2025)
ConfigSpec: Profiling-Based Configuration Selection for Distributed Edge--Cloud Speculative LLM Serving
von: Li, Xiangchen, et al.
Veröffentlicht: (2026)
von: Li, Xiangchen, et al.
Veröffentlicht: (2026)
Federated Learning for Deforestation Detection: A Distributed Approach with Satellite Imagery
von: Dutta, Yuvraj, et al.
Veröffentlicht: (2025)
von: Dutta, Yuvraj, et al.
Veröffentlicht: (2025)
Prima.cpp: Fast 30-70B LLM Inference on Heterogeneous and Low-Resource Home Clusters
von: Li, Zonghang, et al.
Veröffentlicht: (2025)
von: Li, Zonghang, et al.
Veröffentlicht: (2025)
Flash-Fusion: Enabling Expressive, Low-Latency Queries on IoT Sensor Streams with LLMs
von: Patherya, Kausar, et al.
Veröffentlicht: (2025)
von: Patherya, Kausar, et al.
Veröffentlicht: (2025)
Hive: A Multi-Agent Infrastructure for Algorithm- and Task-Level Scaling
von: Luo, Zizhang, et al.
Veröffentlicht: (2026)
von: Luo, Zizhang, et al.
Veröffentlicht: (2026)
Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project
von: Penke, Carolin, et al.
Veröffentlicht: (2025)
von: Penke, Carolin, et al.
Veröffentlicht: (2025)
Combining Serverless and High-Performance Computing Paradigms to support ML Data-Intensive Applications
von: Staylor, Mills, et al.
Veröffentlicht: (2025)
von: Staylor, Mills, et al.
Veröffentlicht: (2025)
Deep RC: A Scalable Data Engineering and Deep Learning Pipeline
von: Sarker, Arup Kumar, et al.
Veröffentlicht: (2025)
von: Sarker, Arup Kumar, et al.
Veröffentlicht: (2025)
Design and Implementation of an Analysis Pipeline for Heterogeneous Data
von: Sarker, Arup Kumar, et al.
Veröffentlicht: (2024)
von: Sarker, Arup Kumar, et al.
Veröffentlicht: (2024)
WISP: Waste- and Interference-Suppressed Distributed Speculative LLM Serving at the Edge via Dynamic Drafting and SLO-Aware Batching
von: Li, Xiangchen, et al.
Veröffentlicht: (2026)
von: Li, Xiangchen, et al.
Veröffentlicht: (2026)
Towards Building Private LLMs: Exploring Multi-Node Expert Parallelism on Apple Silicon for Mixture-of-Experts Large Language Model
von: Chen, Mu-Chi, et al.
Veröffentlicht: (2025)
von: Chen, Mu-Chi, et al.
Veröffentlicht: (2025)
POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference
von: Kamath, Aditya K, et al.
Veröffentlicht: (2024)
von: Kamath, Aditya K, et al.
Veröffentlicht: (2024)
AIvailable: A Software-Defined Architecture for LLM-as-a-Service on Heterogeneous and Legacy GPUs
von: Antunes, Pedro, et al.
Veröffentlicht: (2025)
von: Antunes, Pedro, et al.
Veröffentlicht: (2025)
Comparative Analysis of Large Language Model Inference Serving Systems: A Performance Study of vLLM and HuggingFace TGI
von: Kolluru, Saicharan
Veröffentlicht: (2025)
von: Kolluru, Saicharan
Veröffentlicht: (2025)
Worldwide Federated Training of Language Models
von: Iacob, Alex, et al.
Veröffentlicht: (2024)
von: Iacob, Alex, et al.
Veröffentlicht: (2024)
Chronicals: A High-Performance Framework for LLM Fine-Tuning with 3.51x Speedup over Unsloth
von: Nair, Arjun S.
Veröffentlicht: (2026)
von: Nair, Arjun S.
Veröffentlicht: (2026)
Network and Systems Performance Characterization of MCP-Enabled LLM Agents
von: Ding, Zihao, et al.
Veröffentlicht: (2025)
von: Ding, Zihao, et al.
Veröffentlicht: (2025)
GraphBit: A Graph-based Agentic Framework for Non-Linear Agent Orchestration
von: Sarker, Yeahia, et al.
Veröffentlicht: (2026)
von: Sarker, Yeahia, et al.
Veröffentlicht: (2026)
ATTNChecker: Highly-Optimized Fault Tolerant Attention for Large Language Model Training
von: Liang, Yuhang, et al.
Veröffentlicht: (2024)
von: Liang, Yuhang, et al.
Veröffentlicht: (2024)
Architecture-Aware LLM Inference Optimization on AMD Instinct GPUs: A Comprehensive Benchmark and Deployment Study
von: Georgiou, Athos
Veröffentlicht: (2026)
von: Georgiou, Athos
Veröffentlicht: (2026)
DSDE: Dynamic Speculative Decoding with KLD Stability for Real-World Serving
von: Yang, Mingyu, et al.
Veröffentlicht: (2025)
von: Yang, Mingyu, et al.
Veröffentlicht: (2025)
A Survey on Parallel Text Generation: From Parallel Decoding to Diffusion Language Models
von: Zhang, Lingzhe, et al.
Veröffentlicht: (2025)
von: Zhang, Lingzhe, et al.
Veröffentlicht: (2025)
Scalability Optimization in Cloud-Based AI Inference Services: Strategies for Real-Time Load Balancing and Automated Scaling
von: Jin, Yihong, et al.
Veröffentlicht: (2025)
von: Jin, Yihong, et al.
Veröffentlicht: (2025)
MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices
von: Shakerdargah, Mohammadali, et al.
Veröffentlicht: (2024)
von: Shakerdargah, Mohammadali, et al.
Veröffentlicht: (2024)
Token Arena: A Continuous Benchmark Unifying Energy and Cognition in AI Inference
von: Gao, Yuxuan, et al.
Veröffentlicht: (2026)
von: Gao, Yuxuan, et al.
Veröffentlicht: (2026)
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
von: Jo, Myeong Jun
Veröffentlicht: (2026)
von: Jo, Myeong Jun
Veröffentlicht: (2026)
Agentic Compilation: Mitigating the LLM Rerun Crisis for Minimized-Inference-Cost Web Automation
von: Chundru, Jagadeesh
Veröffentlicht: (2026)
von: Chundru, Jagadeesh
Veröffentlicht: (2026)
MATCH: Model-Aware TVM-based Compilation for Heterogeneous Edge Devices
von: Hamdi, Mohamed Amine, et al.
Veröffentlicht: (2024)
von: Hamdi, Mohamed Amine, et al.
Veröffentlicht: (2024)
Spark-LLM-Eval: A Distributed Framework for Statistically Rigorous Large Language Model Evaluation
von: Mitra, Subhadip
Veröffentlicht: (2026)
von: Mitra, Subhadip
Veröffentlicht: (2026)
MeanCache: User-Centric Semantic Caching for LLM Web Services
von: Gill, Waris, et al.
Veröffentlicht: (2024)
von: Gill, Waris, et al.
Veröffentlicht: (2024)
Cost Trade-offs of Reasoning and Non-Reasoning Large Language Models in Text-to-SQL
von: Deochake, Saurabh, et al.
Veröffentlicht: (2025)
von: Deochake, Saurabh, et al.
Veröffentlicht: (2025)
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended
von: Kamath, Aditya K, et al.
Veröffentlicht: (2026)
von: Kamath, Aditya K, et al.
Veröffentlicht: (2026)
ADF-LoRA: Alternating Low-Rank Aggregation for Decentralized Federated Fine-Tuning
von: Wang, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Wang, Xiaoyu, et al.
Veröffentlicht: (2025)
De-DSI: Decentralised Differentiable Search Index
von: Neague, Petru, et al.
Veröffentlicht: (2024)
von: Neague, Petru, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
FedMon: Federated eBPF Monitoring for Distributed Anomaly Detection in Multi-Cluster Cloud Environments
von: Zehra, Sehar, et al.
Veröffentlicht: (2025) -
AAFLOW: Scalable Patterns for Agentic AI Workflows
von: Sarker, Arup Kumar, et al.
Veröffentlicht: (2026) -
Benchmarking Federated Learning for Throughput Prediction in 5G Live Streaming Applications
von: Dutta, Yuvraj, et al.
Veröffentlicht: (2025) -
DAGER: Exact Gradient Inversion for Large Language Models
von: Petrov, Ivo, et al.
Veröffentlicht: (2024) -
Augmenting the FedProx Algorithm by Minimizing Convergence
von: Sarkar, Anomitra, et al.
Veröffentlicht: (2024)