PreSto: An In-Storage Data Preprocessing System for Training Recommendation Models
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Yunjae, Kim, Hyeseong, Rhu, Minsoo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
vTrain: A Simulation Framework for Evaluating Cost-effective and Compute-optimal Large Language Model Training
by: Bang, Jehyeon, et al.
Published: (2023)
by: Bang, Jehyeon, et al.
Published: (2023)
SwarmIO: Towards 100 Million IOPS SSD Emulation for Next-generation GPU-centric Storage Systems
by: Kim, Hyeseong, et al.
Published: (2026)
by: Kim, Hyeseong, et al.
Published: (2026)
PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers
by: Yeo, Gwangoo, et al.
Published: (2024)
by: Yeo, Gwangoo, et al.
Published: (2024)
The Cost of Dynamic Reasoning: Demystifying AI Agents and Test-Time Scaling from an AI Infrastructure Perspective
by: Kim, Jiin, et al.
Published: (2025)
by: Kim, Jiin, et al.
Published: (2025)
PASCAL: A Phase-Aware Scheduling Algorithm for Serving Reasoning-based Large Language Models
by: Cho, Eunyeong, et al.
Published: (2026)
by: Cho, Eunyeong, et al.
Published: (2026)
Mamba-X: An End-to-End Vision Mamba Accelerator for Edge Computing Devices
by: Yoon, Dongho, et al.
Published: (2025)
by: Yoon, Dongho, et al.
Published: (2025)
Heterogeneous Acceleration Pipeline for Recommendation System Training
by: Adnan, Muhammad, et al.
Published: (2022)
by: Adnan, Muhammad, et al.
Published: (2022)
PIM-MMU: A Memory Management Unit for Accelerating Data Transfers in Commercial PIM Systems
by: Lee, Dongjae, et al.
Published: (2024)
by: Lee, Dongjae, et al.
Published: (2024)
VerilogDB: The Largest, Highest-Quality Dataset with a Preprocessing Framework for LLM-based RTL Generation
by: Calzada, Paul E., et al.
Published: (2025)
by: Calzada, Paul E., et al.
Published: (2025)
Debunking the CUDA Myth Towards GPU-based AI Systems
by: Lee, Yunjae, et al.
Published: (2024)
by: Lee, Yunjae, et al.
Published: (2024)
MoNDE: Mixture of Near-Data Experts for Large-Scale Sparse Models
by: Kim, Taehyun, et al.
Published: (2024)
by: Kim, Taehyun, et al.
Published: (2024)
Pathfinding Future PIM Architectures by Demystifying a Commercial PIM Technology
by: Hyun, Bongjoon, et al.
Published: (2023)
by: Hyun, Bongjoon, et al.
Published: (2023)
Pre-gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference
by: Hwang, Ranggi, et al.
Published: (2023)
by: Hwang, Ranggi, et al.
Published: (2023)
Improving Quantization with Post-Training Model Expansion
by: Franco, Giuseppe, et al.
Published: (2025)
by: Franco, Giuseppe, et al.
Published: (2025)
EXION: Exploiting Inter- and Intra-Iteration Output Sparsity for Diffusion Models
by: Heo, Jaehoon, et al.
Published: (2025)
by: Heo, Jaehoon, et al.
Published: (2025)
SCRec: A Scalable Computational Storage System with Statistical Sharding and Tensor-train Decomposition for Recommendation Models
by: Yang, Jinho, et al.
Published: (2025)
by: Yang, Jinho, et al.
Published: (2025)
PrefixGPT: Prefix Adder Optimization by a Generative Pre-trained Transformer
by: Ding, Ruogu, et al.
Published: (2025)
by: Ding, Ruogu, et al.
Published: (2025)
Smart-Infinity: Fast Large Language Model Training using Near-Storage Processing on a Real System
by: Jang, Hongsun, et al.
Published: (2024)
by: Jang, Hongsun, et al.
Published: (2024)
PIM-malloc: A Fast and Scalable Dynamic Memory Allocator for Processing-In-Memory (PIM) Architectures
by: Lee, Dongjae, et al.
Published: (2025)
by: Lee, Dongjae, et al.
Published: (2025)
Skip the Benchmark: Generating System-Level High-Level Synthesis Data using Generative Machine Learning
by: Liao, Yuchao, et al.
Published: (2024)
by: Liao, Yuchao, et al.
Published: (2024)
Cocoon: A System Architecture for Differentially Private Training with Correlated Noises
by: Kim, Donghwan, et al.
Published: (2025)
by: Kim, Donghwan, et al.
Published: (2025)
Efficient Tabular Data Preprocessing of ML Pipelines
by: Zhu, Yu, et al.
Published: (2024)
by: Zhu, Yu, et al.
Published: (2024)
Mirage: An RNS-Based Photonic Accelerator for DNN Training
by: Demirkiran, Cansu, et al.
Published: (2023)
by: Demirkiran, Cansu, et al.
Published: (2023)
Chip Placement with Diffusion Models
by: Lee, Vint, et al.
Published: (2024)
by: Lee, Vint, et al.
Published: (2024)
Tensor-Compressed and Fully-Quantized Training of Neural PDE Solvers
by: Lu, Jinming, et al.
Published: (2025)
by: Lu, Jinming, et al.
Published: (2025)
TRAM: Training Approximate Multiplier Structures for Low-Power AI Accelerators
by: Meng, Chang, et al.
Published: (2026)
by: Meng, Chang, et al.
Published: (2026)
Graph Computation Meets Circuit Algebra: A Task-Aligned Analysis of Graph Neural Networks for Electronic Design Automation
by: Kim, Hyunmog
Published: (2026)
by: Kim, Hyunmog
Published: (2026)
NeFT: Negative Feedback Training to Improve Robustness of Compute-In-Memory DNN Accelerators
by: Qin, Yifan, et al.
Published: (2023)
by: Qin, Yifan, et al.
Published: (2023)
FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design
by: Xia, Haojun, et al.
Published: (2024)
by: Xia, Haojun, et al.
Published: (2024)
Column-wise Quantization of Weights and Partial Sums for Accurate and Efficient Compute-In-Memory Accelerators
by: Kim, Jiyoon, et al.
Published: (2025)
by: Kim, Jiyoon, et al.
Published: (2025)
Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns
by: Bambhaniya, Abhimanyu, et al.
Published: (2026)
by: Bambhaniya, Abhimanyu, et al.
Published: (2026)
LayerPipe2: Multistage Pipelining and Weight Recompute via Improved Exponential Moving Average for Training Neural Networks
by: Unnikrishnan, Nanda K., et al.
Published: (2025)
by: Unnikrishnan, Nanda K., et al.
Published: (2025)
Continuous-Flow Data-Rate-Aware CNN Inference on FPGA
by: Habermann, Tobias, et al.
Published: (2026)
by: Habermann, Tobias, et al.
Published: (2026)
MEADOW: Memory-efficient Dataflow and Data Packing for Low Power Edge LLMs
by: Moitra, Abhishek, et al.
Published: (2025)
by: Moitra, Abhishek, et al.
Published: (2025)
Enabling Physical AI at the Edge: Hardware-Accelerated Recovery of System Dynamics
by: Xu, Bin, et al.
Published: (2025)
by: Xu, Bin, et al.
Published: (2025)
Anda: Unlocking Efficient LLM Inference with a Variable-Length Grouped Activation Data Format
by: Fang, Chao, et al.
Published: (2024)
by: Fang, Chao, et al.
Published: (2024)
SwiftRL: Towards Efficient Reinforcement Learning on Real Processing-In-Memory Systems
by: Gogineni, Kailash, et al.
Published: (2024)
by: Gogineni, Kailash, et al.
Published: (2024)
MATADOR: Automated System-on-Chip Tsetlin Machine Design Generation for Edge Applications
by: Rahman, Tousif, et al.
Published: (2024)
by: Rahman, Tousif, et al.
Published: (2024)
Time-Series Forecasting and Sequence Learning Using Memristor-based Reservoir System
by: Zyarah, Abdullah M., et al.
Published: (2024)
by: Zyarah, Abdullah M., et al.
Published: (2024)
MetaWearS: A Shortcut in Wearable Systems Lifecycle with Only a Few Shots
by: Amirshahi, Alireza, et al.
Published: (2024)
by: Amirshahi, Alireza, et al.
Published: (2024)
Similar Items
-
vTrain: A Simulation Framework for Evaluating Cost-effective and Compute-optimal Large Language Model Training
by: Bang, Jehyeon, et al.
Published: (2023) -
SwarmIO: Towards 100 Million IOPS SSD Emulation for Next-generation GPU-centric Storage Systems
by: Kim, Hyeseong, et al.
Published: (2026) -
PREBA: A Hardware/Software Co-Design for Multi-Instance GPU based AI Inference Servers
by: Yeo, Gwangoo, et al.
Published: (2024) -
The Cost of Dynamic Reasoning: Demystifying AI Agents and Test-Time Scaling from an AI Infrastructure Perspective
by: Kim, Jiin, et al.
Published: (2025) -
PASCAL: A Phase-Aware Scheduling Algorithm for Serving Reasoning-based Large Language Models
by: Cho, Eunyeong, et al.
Published: (2026)