Efficient Data-Parallel Continual Learning with Asynchronous Distributed Rehearsal Buffers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bouvier, Thomas, Nicolae, Bogdan, Chaugier, Hugo, Costan, Alexandru, Foster, Ian, Antoniu, Gabriel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DataStates-LLM: Lazy Asynchronous Checkpointing for Large Language Models
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
D-Rex: Heterogeneity-Aware Reliability Framework and Adaptive Algorithms for Distributed Storage
von: Gonthier, Maxime, et al.
Veröffentlicht: (2025)
von: Gonthier, Maxime, et al.
Veröffentlicht: (2025)
Kavier: Exploring Performance, Sustainability, and Efficiency of LLM Ecosystems under Inference through Cache-Aware Discrete-Event Simulation
von: Nicolae, Radu, et al.
Veröffentlicht: (2026)
von: Nicolae, Radu, et al.
Veröffentlicht: (2026)
M3SA: Exploring Datacenter Performance and Climate-Impact with Multi- and Meta-Model Simulation and Analysis
von: Nicolae, Radu, et al.
Veröffentlicht: (2026)
von: Nicolae, Radu, et al.
Veröffentlicht: (2026)
An Asynchronous Distributed-Memory Parallel Algorithm for k-mer Counting
von: Hati, Souvadra, et al.
Veröffentlicht: (2025)
von: Hati, Souvadra, et al.
Veröffentlicht: (2025)
DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers
von: Maurya, Avinash, et al.
Veröffentlicht: (2026)
von: Maurya, Avinash, et al.
Veröffentlicht: (2026)
Rehearsal-Free Continual Federated Learning with Synergistic Synaptic Intelligence
von: Li, Yichen, et al.
Veröffentlicht: (2024)
von: Li, Yichen, et al.
Veröffentlicht: (2024)
Wilkins: HPC In Situ Workflows Made Easy
von: Yildiz, Orcun, et al.
Veröffentlicht: (2024)
von: Yildiz, Orcun, et al.
Veröffentlicht: (2024)
Understanding LLM Checkpoint/Restore I/O Strategies and Patterns
von: Gossman, Mikaila J., et al.
Veröffentlicht: (2025)
von: Gossman, Mikaila J., et al.
Veröffentlicht: (2025)
Distributed Download from an External Data Source in Asynchronous Faulty Settings
von: Augustine, John, et al.
Veröffentlicht: (2025)
von: Augustine, John, et al.
Veröffentlicht: (2025)
Boosting Blockchain Throughput: Parallel EVM Execution with Asynchronous Storage for Reddio
von: Qi, Xiaodong, et al.
Veröffentlicht: (2025)
von: Qi, Xiaodong, et al.
Veröffentlicht: (2025)
Breaking the Memory Wall: A Study of I/O Patterns and GPU Memory Utilization for Hybrid CPU-GPU Offloaded Optimizers
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
von: Maurya, Avinash, et al.
Veröffentlicht: (2024)
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles
von: Arif, Moiz, et al.
Veröffentlicht: (2026)
von: Arif, Moiz, et al.
Veröffentlicht: (2026)
OpenDT: Exploring Datacenter Performance and Sustainability with a Self-Calibrating Digital Twin
von: Nicolae, Radu, et al.
Veröffentlicht: (2026)
von: Nicolae, Radu, et al.
Veröffentlicht: (2026)
Fault-Tolerant Decentralized Distributed Asynchronous Federated Learning with Adaptive Termination Detection
von: Akkinepally, Phani Sahasra, et al.
Veröffentlicht: (2025)
von: Akkinepally, Phani Sahasra, et al.
Veröffentlicht: (2025)
Computational Grids
von: Foster, Ian, et al.
Veröffentlicht: (2025)
von: Foster, Ian, et al.
Veröffentlicht: (2025)
APEX: Asynchronous Parallel CPU-GPU Execution for Online LLM Inference on Constrained GPUs
von: Fan, Jiakun, et al.
Veröffentlicht: (2025)
von: Fan, Jiakun, et al.
Veröffentlicht: (2025)
FedStaleWeight: Buffered Asynchronous Federated Learning with Fair Aggregation via Staleness Reweighting
von: Ma, Jeffrey, et al.
Veröffentlicht: (2024)
von: Ma, Jeffrey, et al.
Veröffentlicht: (2024)
AsyncMesh: Fully Asynchronous Optimization for Data and Pipeline Parallelism
von: Ajanthan, Thalaiyasingam, et al.
Veröffentlicht: (2026)
von: Ajanthan, Thalaiyasingam, et al.
Veröffentlicht: (2026)
Asynchronous Fault-Tolerant Distributed Proper Coloring of Graphs
von: Balliu, Alkida, et al.
Veröffentlicht: (2024)
von: Balliu, Alkida, et al.
Veröffentlicht: (2024)
AGILE: Lightweight and Efficient Asynchronous GPU-SSD Integration
von: Yang, Zhuoping, et al.
Veröffentlicht: (2025)
von: Yang, Zhuoping, et al.
Veröffentlicht: (2025)
UniFaaS: Programming across Distributed Cyberinfrastructure with Federated Function Serving
von: Li, Yifei, et al.
Veröffentlicht: (2024)
von: Li, Yifei, et al.
Veröffentlicht: (2024)
Buffer-based Gradient Projection for Continual Federated Learning
von: Dai, Shenghong, et al.
Veröffentlicht: (2024)
von: Dai, Shenghong, et al.
Veröffentlicht: (2024)
SiDP: Memory-Efficient Data Parallelism for Offline LLM Inference
von: Zhao, Alan, et al.
Veröffentlicht: (2026)
von: Zhao, Alan, et al.
Veröffentlicht: (2026)
WRATH: Workload Resilience Across Task Hierarchies in Task-based Parallel Programming Frameworks
von: Zhou, Sicheng, et al.
Veröffentlicht: (2025)
von: Zhou, Sicheng, et al.
Veröffentlicht: (2025)
Bandwidth-Aware and Cost-Efficient Pipeline Parallel Scheduling in Geo-Distributed LLM Training
von: Zhang, Han, et al.
Veröffentlicht: (2026)
von: Zhang, Han, et al.
Veröffentlicht: (2026)
Asynchronous Fault-Tolerant Language Decidability for Runtime Verification of Distributed Systems
von: Castañeda, Armando, et al.
Veröffentlicht: (2025)
von: Castañeda, Armando, et al.
Veröffentlicht: (2025)
Asynchronous Secure Federated Learning with Byzantine aggregators
von: Del Pozzo, Antonella, et al.
Veröffentlicht: (2026)
von: Del Pozzo, Antonella, et al.
Veröffentlicht: (2026)
SwitchDelta: Asynchronous Metadata Updating for Distributed Storage with In-Network Data Visibility
von: Li, Junru, et al.
Veröffentlicht: (2025)
von: Li, Junru, et al.
Veröffentlicht: (2025)
Delphi: Efficient Asynchronous Approximate Agreement for Distributed Oracles
von: Bandarupalli, Akhil, et al.
Veröffentlicht: (2024)
von: Bandarupalli, Akhil, et al.
Veröffentlicht: (2024)
Object Proxy Patterns for Accelerating Distributed Applications
von: Pauloski, J. Gregory, et al.
Veröffentlicht: (2024)
von: Pauloski, J. Gregory, et al.
Veröffentlicht: (2024)
Nesterov Method for Asynchronous Pipeline Parallel Optimization
von: Ajanthan, Thalaiyasingam, et al.
Veröffentlicht: (2025)
von: Ajanthan, Thalaiyasingam, et al.
Veröffentlicht: (2025)
LCI: a Lightweight Communication Interface for Efficient Asynchronous Multithreaded Communication
von: Yan, Jiakun, et al.
Veröffentlicht: (2025)
von: Yan, Jiakun, et al.
Veröffentlicht: (2025)
SPARe: Stacked Parallelism with Adaptive Reordering for Fault-Tolerant LLM Pretraining Systems with 100k+ GPUs
von: Lee, Jin, et al.
Veröffentlicht: (2026)
von: Lee, Jin, et al.
Veröffentlicht: (2026)
Experiences Porting Distributed Applications to Asynchronous Tasks: A Multidimensional FFT Case-study
von: Strack, Alexander, et al.
Veröffentlicht: (2024)
von: Strack, Alexander, et al.
Veröffentlicht: (2024)
Automated, Reliable, and Efficient Continental-Scale Replication of 7.3 Petabytes of Climate Simulation Data: A Case Study
von: Lacinski, Lukasz, et al.
Veröffentlicht: (2024)
von: Lacinski, Lukasz, et al.
Veröffentlicht: (2024)
JanusPipe: Efficient Pipeline Parallel Training for Machine Learning Interatomic Potentials
von: Wang, Hongyu, et al.
Veröffentlicht: (2026)
von: Wang, Hongyu, et al.
Veröffentlicht: (2026)
DreamDDP: Accelerating Data Parallel Distributed LLM Training with Layer-wise Scheduled Partial Synchronization
von: Tang, Zhenheng, et al.
Veröffentlicht: (2025)
von: Tang, Zhenheng, et al.
Veröffentlicht: (2025)
Equivalence and Separation between Heard-Of and Asynchronous Message-Passing Models
von: Attiya, Hagit, et al.
Veröffentlicht: (2025)
von: Attiya, Hagit, et al.
Veröffentlicht: (2025)
Hecate: Unlocking Efficient Sparse Model Training via Fully Sharded Sparse Data Parallelism
von: Qing, Yuhao, et al.
Veröffentlicht: (2025)
von: Qing, Yuhao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
DataStates-LLM: Lazy Asynchronous Checkpointing for Large Language Models
von: Maurya, Avinash, et al.
Veröffentlicht: (2024) -
D-Rex: Heterogeneity-Aware Reliability Framework and Adaptive Algorithms for Distributed Storage
von: Gonthier, Maxime, et al.
Veröffentlicht: (2025) -
Kavier: Exploring Performance, Sustainability, and Efficiency of LLM Ecosystems under Inference through Cache-Aware Discrete-Event Simulation
von: Nicolae, Radu, et al.
Veröffentlicht: (2026) -
M3SA: Exploring Datacenter Performance and Climate-Impact with Multi- and Meta-Model Simulation and Analysis
von: Nicolae, Radu, et al.
Veröffentlicht: (2026) -
An Asynchronous Distributed-Memory Parallel Algorithm for k-mer Counting
von: Hati, Souvadra, et al.
Veröffentlicht: (2025)