SSDTrain: An Activation Offloading Framework to SSDs for Faster Large Language Model Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Kun, Park, Jeongmin Brian, Zhang, Xiaofan, Hidayetoğlu, Mert, Mailthody, Vikram Sharma, Huang, Sitao, Lumetta, Steven Sam, Hwu, Wen-mei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hector: An Efficient Programming and Compilation Framework for Implementing Relational Graph Neural Networks in GPU Architectures
von: Wu, Kun, et al.
Veröffentlicht: (2023)
von: Wu, Kun, et al.
Veröffentlicht: (2023)
LSM-GNN: Large-scale Storage-based Multi-GPU GNN Training by Optimizing Data Transfer Scheme
von: Park, Jeongmin Brian, et al.
Veröffentlicht: (2024)
von: Park, Jeongmin Brian, et al.
Veröffentlicht: (2024)
Accelerating Sampling and Aggregation Operations in GNN Frameworks with GPU Initiated Direct Storage Accesses
von: Park, Jeongmin Brian, et al.
Veröffentlicht: (2023)
von: Park, Jeongmin Brian, et al.
Veröffentlicht: (2023)
Code generation and runtime techniques for enabling data-efficient deep learning training on GPUs
von: Wu, Kun
Veröffentlicht: (2024)
von: Wu, Kun
Veröffentlicht: (2024)
D-CODE: Data Colony Optimization for Dynamic Network Efficiency
von: Pandey, Tannu, et al.
Veröffentlicht: (2024)
von: Pandey, Tannu, et al.
Veröffentlicht: (2024)
A Frequency-based Parent Selection for Reducing the Effect of Evaluation Time Bias in Asynchronous Parallel Multi-objective Evolutionary Algorithms
von: Harada, Tomohiro
Veröffentlicht: (2021)
von: Harada, Tomohiro
Veröffentlicht: (2021)
Neuromorphic Simulation of Drosophila Melanogaster Brain Connectome on Loihi 2
von: Wang, Felix, et al.
Veröffentlicht: (2025)
von: Wang, Felix, et al.
Veröffentlicht: (2025)
A Fresh Approach to Evaluate Performance in Distributed Parallel Genetic Algorithms
von: Harada, Tomohiro, et al.
Veröffentlicht: (2021)
von: Harada, Tomohiro, et al.
Veröffentlicht: (2021)
Evaluation and Efficiency Comparison of Evolutionary Algorithms for Service Placement Optimization in Fog Architectures
von: Guerrero, Carlos, et al.
Veröffentlicht: (2025)
von: Guerrero, Carlos, et al.
Veröffentlicht: (2025)
Reducing Data Bottlenecks in Distributed, Heterogeneous Neural Networks
von: Lin, Ruhai, et al.
Veröffentlicht: (2024)
von: Lin, Ruhai, et al.
Veröffentlicht: (2024)
Trackable Agent-based Evolution Models at Wafer Scale
von: Moreno, Matthew Andres, et al.
Veröffentlicht: (2024)
von: Moreno, Matthew Andres, et al.
Veröffentlicht: (2024)
Trackable Island-model Genetic Algorithms at Wafer Scale
von: Moreno, Matthew Andres, et al.
Veröffentlicht: (2024)
von: Moreno, Matthew Andres, et al.
Veröffentlicht: (2024)
Sparse Spiking Neural-like Membrane Systems on Graphics Processing Units
von: Hernández-Tello, Javier, et al.
Veröffentlicht: (2024)
von: Hernández-Tello, Javier, et al.
Veröffentlicht: (2024)
Neuromorphic Computing: A Theoretical Framework for Time, Space, and Energy Scaling
von: Aimone, James B
Veröffentlicht: (2025)
von: Aimone, James B
Veröffentlicht: (2025)
HiAER-Spike: Hardware-Software Co-Design for Large-Scale Reconfigurable Event-Driven Neuromorphic Computing
von: Frank, Gwenevere, et al.
Veröffentlicht: (2025)
von: Frank, Gwenevere, et al.
Veröffentlicht: (2025)
Neuromorphic hardware for sustainable AI data centers
von: Vogginger, Bernhard, et al.
Veröffentlicht: (2024)
von: Vogginger, Bernhard, et al.
Veröffentlicht: (2024)
Federated Learning in Chemical Engineering: A Tutorial on a Framework for Privacy-Preserving Collaboration Across Distributed Data Sources
von: Dutta, Siddhant, et al.
Veröffentlicht: (2024)
von: Dutta, Siddhant, et al.
Veröffentlicht: (2024)
phys-MCP: A Control Plane for Heterogeneous Physical Neural Networks
von: Fischer, Stefan, et al.
Veröffentlicht: (2026)
von: Fischer, Stefan, et al.
Veröffentlicht: (2026)
GAP2WSS: A Genetic Algorithm based on the Pareto Principle for Web Service Selection
von: Khatoonabadi, SayedHassan, et al.
Veröffentlicht: (2021)
von: Khatoonabadi, SayedHassan, et al.
Veröffentlicht: (2021)
The AI Shadow War: SaaS vs. Edge Computing Architectures
von: Marpu, Rhea Pritham, et al.
Veröffentlicht: (2025)
von: Marpu, Rhea Pritham, et al.
Veröffentlicht: (2025)
A Reinforced Evolution-Based Approach to Multi-Resource Load Balancing
von: Sliwko, Leszek
Veröffentlicht: (2025)
von: Sliwko, Leszek
Veröffentlicht: (2025)
Wireless Sensor Networks as Parallel and Distributed Hardware Platform for Artificial Neural Networks
von: Serpen, Gursel
Veröffentlicht: (2025)
von: Serpen, Gursel
Veröffentlicht: (2025)
NeuroRing: Scaling Spiking Neural Networks via Multi-FPGA Bidirectional Ring Topologies and Stream-Dataflow Architectures
von: Hafiz, Muhammad Ihsan Al, et al.
Veröffentlicht: (2026)
von: Hafiz, Muhammad Ihsan Al, et al.
Veröffentlicht: (2026)
MAC-DO: An Efficient Output-Stationary GEMM Accelerator for CNNs Using DRAM Technology
von: Jeong, Minki, et al.
Veröffentlicht: (2022)
von: Jeong, Minki, et al.
Veröffentlicht: (2022)
Asynchronous Evolution of Deep Neural Network Architectures
von: Liang, Jason, et al.
Veröffentlicht: (2023)
von: Liang, Jason, et al.
Veröffentlicht: (2023)
Distributed genetic algorithm for application placement in the compute continuum leveraging infrastructure nodes for optimization
von: Guerrero, Carlos, et al.
Veröffentlicht: (2024)
von: Guerrero, Carlos, et al.
Veröffentlicht: (2024)
Edge Intelligence with Spiking Neural Networks
von: Deng, Shuiguang, et al.
Veröffentlicht: (2025)
von: Deng, Shuiguang, et al.
Veröffentlicht: (2025)
ParEVO: Synthesizing Code for Irregular Data: High-Performance Parallelism through Agentic Evolution
von: Yang, Liu, et al.
Veröffentlicht: (2026)
von: Yang, Liu, et al.
Veröffentlicht: (2026)
OpenRASE: Service Function Chain Emulation
von: Krishnamohan, Theviyanthan, et al.
Veröffentlicht: (2025)
von: Krishnamohan, Theviyanthan, et al.
Veröffentlicht: (2025)
Towards a Decentralised Application-Centric Orchestration Framework in the Cloud-Edge Continuum
von: Ullah, Amjad, et al.
Veröffentlicht: (2025)
von: Ullah, Amjad, et al.
Veröffentlicht: (2025)
HiCCL: A Hierarchical Collective Communication Library
von: Hidayetoglu, Mert, et al.
Veröffentlicht: (2024)
von: Hidayetoglu, Mert, et al.
Veröffentlicht: (2024)
Scalable Construction of Spiking Neural Networks using up to thousands of GPUs
von: Golosio, Bruno, et al.
Veröffentlicht: (2025)
von: Golosio, Bruno, et al.
Veröffentlicht: (2025)
Split Federated Learning Over Heterogeneous Edge Devices: Algorithm and Optimization
von: Sun, Yunrui, et al.
Veröffentlicht: (2024)
von: Sun, Yunrui, et al.
Veröffentlicht: (2024)
Comparison of Microservice Call Rate Predictions for Replication in the Cloud
von: Mehran, Narges, et al.
Veröffentlicht: (2023)
von: Mehran, Narges, et al.
Veröffentlicht: (2023)
Linear Reservoir: A Diagonalization-Based Optimization
von: de Coudenhove, Romain, et al.
Veröffentlicht: (2026)
von: de Coudenhove, Romain, et al.
Veröffentlicht: (2026)
Cyclic Data Parallelism for Efficient Parallelism of Deep Neural Networks
von: Fournier, Louis, et al.
Veröffentlicht: (2024)
von: Fournier, Louis, et al.
Veröffentlicht: (2024)
Online Continual Learning on Intel Loihi 2 via a Co-designed Spiking Neural Network
von: Hajizada, Elvin, et al.
Veröffentlicht: (2025)
von: Hajizada, Elvin, et al.
Veröffentlicht: (2025)
NeuraChip: Accelerating GNN Computations with a Hash-based Decoupled Spatial Accelerator
von: Shivdikar, Kaustubh, et al.
Veröffentlicht: (2024)
von: Shivdikar, Kaustubh, et al.
Veröffentlicht: (2024)
All-to-all reconfigurability with sparse and higher-order Ising machines
von: Nikhar, Srijan, et al.
Veröffentlicht: (2023)
von: Nikhar, Srijan, et al.
Veröffentlicht: (2023)
Man-Made Heuristics Are Dead. Long Live Code Generators!
von: Dwivedula, Rohit, et al.
Veröffentlicht: (2025)
von: Dwivedula, Rohit, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Hector: An Efficient Programming and Compilation Framework for Implementing Relational Graph Neural Networks in GPU Architectures
von: Wu, Kun, et al.
Veröffentlicht: (2023) -
LSM-GNN: Large-scale Storage-based Multi-GPU GNN Training by Optimizing Data Transfer Scheme
von: Park, Jeongmin Brian, et al.
Veröffentlicht: (2024) -
Accelerating Sampling and Aggregation Operations in GNN Frameworks with GPU Initiated Direct Storage Accesses
von: Park, Jeongmin Brian, et al.
Veröffentlicht: (2023) -
Code generation and runtime techniques for enabling data-efficient deep learning training on GPUs
von: Wu, Kun
Veröffentlicht: (2024) -
D-CODE: Data Colony Optimization for Dynamic Network Efficiency
von: Pandey, Tannu, et al.
Veröffentlicht: (2024)