Worldwide Federated Training of Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Iacob, Alex, Sani, Lorenzo, Marino, Bill, Aleksandrov, Preslav, Shen, William F., Lane, Nicholas Donald |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
The Future of Large Language Model Pre-training is Federated
par: Sani, Lorenzo, et autres
Publié: (2024)
par: Sani, Lorenzo, et autres
Publié: (2024)
Photon: Federated LLM Pre-Training
par: Sani, Lorenzo, et autres
Publié: (2024)
par: Sani, Lorenzo, et autres
Publié: (2024)
FedMon: Federated eBPF Monitoring for Distributed Anomaly Detection in Multi-Cluster Cloud Environments
par: Zehra, Sehar, et autres
Publié: (2025)
par: Zehra, Sehar, et autres
Publié: (2025)
DAGER: Exact Gradient Inversion for Large Language Models
par: Petrov, Ivo, et autres
Publié: (2024)
par: Petrov, Ivo, et autres
Publié: (2024)
Hive: A Multi-Agent Infrastructure for Algorithm- and Task-Level Scaling
par: Luo, Zizhang, et autres
Publié: (2026)
par: Luo, Zizhang, et autres
Publié: (2026)
Flash-Fusion: Enabling Expressive, Low-Latency Queries on IoT Sensor Streams with LLMs
par: Patherya, Kausar, et autres
Publié: (2025)
par: Patherya, Kausar, et autres
Publié: (2025)
Pollen: High-throughput Federated Learning Simulation via Resource-Aware Client Placement
par: Sani, Lorenzo, et autres
Publié: (2023)
par: Sani, Lorenzo, et autres
Publié: (2023)
Parameter-Efficient and Personalized Federated Training of Generative Models at the Edge
par: Khan, Kabir, et autres
Publié: (2025)
par: Khan, Kabir, et autres
Publié: (2025)
CooperLLM: Cloud-Edge-End Cooperative Federated Fine-tuning for LLMs via ZOO-based Gradient Correction
par: Sun, He, et autres
Publié: (2026)
par: Sun, He, et autres
Publié: (2026)
Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project
par: Penke, Carolin, et autres
Publié: (2025)
par: Penke, Carolin, et autres
Publié: (2025)
A Survey on Parallel Text Generation: From Parallel Decoding to Diffusion Language Models
par: Zhang, Lingzhe, et autres
Publié: (2025)
par: Zhang, Lingzhe, et autres
Publié: (2025)
Comparative Analysis of Large Language Model Inference Serving Systems: A Performance Study of vLLM and HuggingFace TGI
par: Kolluru, Saicharan
Publié: (2025)
par: Kolluru, Saicharan
Publié: (2025)
AAFLOW: Scalable Patterns for Agentic AI Workflows
par: Sarker, Arup Kumar, et autres
Publié: (2026)
par: Sarker, Arup Kumar, et autres
Publié: (2026)
AIvailable: A Software-Defined Architecture for LLM-as-a-Service on Heterogeneous and Legacy GPUs
par: Antunes, Pedro, et autres
Publié: (2025)
par: Antunes, Pedro, et autres
Publié: (2025)
POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference
par: Kamath, Aditya K, et autres
Publié: (2024)
par: Kamath, Aditya K, et autres
Publié: (2024)
GraphBit: A Graph-based Agentic Framework for Non-Linear Agent Orchestration
par: Sarker, Yeahia, et autres
Publié: (2026)
par: Sarker, Yeahia, et autres
Publié: (2026)
Prima.cpp: Fast 30-70B LLM Inference on Heterogeneous and Low-Resource Home Clusters
par: Li, Zonghang, et autres
Publié: (2025)
par: Li, Zonghang, et autres
Publié: (2025)
Towards Building Private LLMs: Exploring Multi-Node Expert Parallelism on Apple Silicon for Mixture-of-Experts Large Language Model
par: Chen, Mu-Chi, et autres
Publié: (2025)
par: Chen, Mu-Chi, et autres
Publié: (2025)
ATTNChecker: Highly-Optimized Fault Tolerant Attention for Large Language Model Training
par: Liang, Yuhang, et autres
Publié: (2024)
par: Liang, Yuhang, et autres
Publié: (2024)
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory
par: Jo, Myeong Jun
Publié: (2026)
par: Jo, Myeong Jun
Publié: (2026)
Architecture-Aware LLM Inference Optimization on AMD Instinct GPUs: A Comprehensive Benchmark and Deployment Study
par: Georgiou, Athos
Publié: (2026)
par: Georgiou, Athos
Publié: (2026)
Benchmarking Federated Learning for Throughput Prediction in 5G Live Streaming Applications
par: Dutta, Yuvraj, et autres
Publié: (2025)
par: Dutta, Yuvraj, et autres
Publié: (2025)
MeanCache: User-Centric Semantic Caching for LLM Web Services
par: Gill, Waris, et autres
Publié: (2024)
par: Gill, Waris, et autres
Publié: (2024)
Federated Learning for Deforestation Detection: A Distributed Approach with Satellite Imagery
par: Dutta, Yuvraj, et autres
Publié: (2025)
par: Dutta, Yuvraj, et autres
Publié: (2025)
Reducing the GPU Memory Bottleneck with Lossless Compression for ML -- Extended
par: Kamath, Aditya K, et autres
Publié: (2026)
par: Kamath, Aditya K, et autres
Publié: (2026)
DSDE: Dynamic Speculative Decoding with KLD Stability for Real-World Serving
par: Yang, Mingyu, et autres
Publié: (2025)
par: Yang, Mingyu, et autres
Publié: (2025)
Augmenting the FedProx Algorithm by Minimizing Convergence
par: Sarkar, Anomitra, et autres
Publié: (2024)
par: Sarkar, Anomitra, et autres
Publié: (2024)
MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices
par: Shakerdargah, Mohammadali, et autres
Publié: (2024)
par: Shakerdargah, Mohammadali, et autres
Publié: (2024)
Nanoscaling Floating-Point (NxFP): NanoMantissa, Adaptive Microexponents, and Code Recycling for Direct-Cast Compression of Large Language Models
par: Lo, Yun-Chen, et autres
Publié: (2024)
par: Lo, Yun-Chen, et autres
Publié: (2024)
Spark-LLM-Eval: A Distributed Framework for Statistically Rigorous Large Language Model Evaluation
par: Mitra, Subhadip
Publié: (2026)
par: Mitra, Subhadip
Publié: (2026)
Chronicals: A High-Performance Framework for LLM Fine-Tuning with 3.51x Speedup over Unsloth
par: Nair, Arjun S.
Publié: (2026)
par: Nair, Arjun S.
Publié: (2026)
Cost Trade-offs of Reasoning and Non-Reasoning Large Language Models in Text-to-SQL
par: Deochake, Saurabh, et autres
Publié: (2025)
par: Deochake, Saurabh, et autres
Publié: (2025)
Token Arena: A Continuous Benchmark Unifying Energy and Cognition in AI Inference
par: Gao, Yuxuan, et autres
Publié: (2026)
par: Gao, Yuxuan, et autres
Publié: (2026)
Addressing tokens dynamic generation, propagation, storage and renewal to secure the GlideinWMS pilot based jobs and system
par: Coimbra, Bruno Moreira, et autres
Publié: (2025)
par: Coimbra, Bruno Moreira, et autres
Publié: (2025)
Federated Learning Priorities Under the European Union Artificial Intelligence Act
par: Woisetschläger, Herbert, et autres
Publié: (2024)
par: Woisetschläger, Herbert, et autres
Publié: (2024)
Parallelization Strategies for Dense LLM Deployment: Navigating Through Application-Specific Tradeoffs and Bottlenecks
par: Topcu, Burak, et autres
Publié: (2026)
par: Topcu, Burak, et autres
Publié: (2026)
Training Diffusion Models with Federated Learning
par: de Goede, Matthijs, et autres
Publié: (2024)
par: de Goede, Matthijs, et autres
Publié: (2024)
Latency and Cost of Multi-Agent Intelligent Tutoring at Scale
par: Elhaimeur, Iizalaarab, et autres
Publié: (2026)
par: Elhaimeur, Iizalaarab, et autres
Publié: (2026)
Scalable Engine and the Performance of Different LLM Models in a SLURM based HPC architecture
par: Luiz, Anderson de Lima, et autres
Publié: (2025)
par: Luiz, Anderson de Lima, et autres
Publié: (2025)
Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference
par: Ganjihal, Sanjeev Rao
Publié: (2026)
par: Ganjihal, Sanjeev Rao
Publié: (2026)
Documents similaires
-
The Future of Large Language Model Pre-training is Federated
par: Sani, Lorenzo, et autres
Publié: (2024) -
Photon: Federated LLM Pre-Training
par: Sani, Lorenzo, et autres
Publié: (2024) -
FedMon: Federated eBPF Monitoring for Distributed Anomaly Detection in Multi-Cluster Cloud Environments
par: Zehra, Sehar, et autres
Publié: (2025) -
DAGER: Exact Gradient Inversion for Large Language Models
par: Petrov, Ivo, et autres
Publié: (2024) -
Hive: A Multi-Agent Infrastructure for Algorithm- and Task-Level Scaling
par: Luo, Zizhang, et autres
Publié: (2026)