Intent-Driven Storage Systems: From Low-Level Tuning to High-Level Understanding
Fuente:
arXiv
Guardado en:
| Autores principales: | Bergman, Shai, Song, Won Wook, Cavigelli, Lukas, Berestizshevsky, Konstantin, Zhou, Ke, Zhang, Ji |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving
por: Yüzügüler, Ahmet Caner, et al.
Publicado: (2025)
por: Yüzügüler, Ahmet Caner, et al.
Publicado: (2025)
PUDTune: Multi-Level Charging for High-Precision Calibration in Processing-Using-DRAM
por: Kubo, Tatsuya, et al.
Publicado: (2025)
por: Kubo, Tatsuya, et al.
Publicado: (2025)
DP-HLS: A High-Level Synthesis Framework for Accelerating Dynamic Programming Algorithms in Bioinformatics
por: Cao, Yingqi, et al.
Publicado: (2024)
por: Cao, Yingqi, et al.
Publicado: (2024)
MLDSE: Scaling Design Space Exploration Infrastructure for Multi-Level Hardware
por: Qu, Huanyu, et al.
Publicado: (2025)
por: Qu, Huanyu, et al.
Publicado: (2025)
MegIS: High-Performance, Energy-Efficient, and Low-Cost Metagenomic Analysis with In-Storage Processing
por: Ghiasi, Nika Mansouri, et al.
Publicado: (2024)
por: Ghiasi, Nika Mansouri, et al.
Publicado: (2024)
Chopper: A Multi-Level GPU Characterization Tool & Derived Insights Into LLM Training Inefficiency
por: Kurzynski, Marco, et al.
Publicado: (2025)
por: Kurzynski, Marco, et al.
Publicado: (2025)
On-Package Memory with Universal Chiplet Interconnect Express (UCIe): A Low Power, High Bandwidth, Low Latency and Low Cost Approach
por: Sharma, Debendra Das, et al.
Publicado: (2025)
por: Sharma, Debendra Das, et al.
Publicado: (2025)
Cloud-Native Operation of Roadside Infrastructure Enabling Demand-Driven Collective Perception via V2X
por: Zanger, Lukas, et al.
Publicado: (2026)
por: Zanger, Lukas, et al.
Publicado: (2026)
Revisiting Computational Storage for Data Integrity and Security
por: Shi, Chao, et al.
Publicado: (2025)
por: Shi, Chao, et al.
Publicado: (2025)
Mitigating Shared Storage Congestion Using Control Theory
por: Collignon, Thomas, et al.
Publicado: (2025)
por: Collignon, Thomas, et al.
Publicado: (2025)
Adaptive Multi-Objective Tiered Storage Configuration for KV Cache in LLM Service
por: Zheng, Xianzhe, et al.
Publicado: (2026)
por: Zheng, Xianzhe, et al.
Publicado: (2026)
Dissecting the NVIDIA Hopper Architecture through Microbenchmarking and Multiple Level Analysis
por: Luo, Weile, et al.
Publicado: (2025)
por: Luo, Weile, et al.
Publicado: (2025)
Knowledge-Guided Attention-Inspired Learning for Task Offloading in Vehicle Edge Computing
por: Ma, Ke, et al.
Publicado: (2025)
por: Ma, Ke, et al.
Publicado: (2025)
SwarmIO: Towards 100 Million IOPS SSD Emulation for Next-generation GPU-centric Storage Systems
por: Kim, Hyeseong, et al.
Publicado: (2026)
por: Kim, Hyeseong, et al.
Publicado: (2026)
Efficient Batch Search Algorithm for B+ Tree Index Structures with Level-Wise Traversal on FPGAs
por: Tzschoppe, Max, et al.
Publicado: (2026)
por: Tzschoppe, Max, et al.
Publicado: (2026)
Understanding Bottlenecks for Efficiently Serving LLM Inference With KV Offloading
por: Meng, William, et al.
Publicado: (2025)
por: Meng, William, et al.
Publicado: (2025)
Microbenchmark-Driven Analytical Performance Modeling Across Modern GPU Architectures
por: Jarmusch, Aaron, et al.
Publicado: (2026)
por: Jarmusch, Aaron, et al.
Publicado: (2026)
Simopt-Power: Leveraging Simulation Metadata for Low-Power Design Synthesis
por: Wadhwa, Eashan, et al.
Publicado: (2025)
por: Wadhwa, Eashan, et al.
Publicado: (2025)
The DMA Streaming Framework: Kernel-Level Buffer Orchestration for High-Performance AI Data Paths
por: Graziano, Marco
Publicado: (2026)
por: Graziano, Marco
Publicado: (2026)
RapidStream IR: Infrastructure for FPGA High-Level Physical Synthesis
por: Lau, Jason, et al.
Publicado: (2024)
por: Lau, Jason, et al.
Publicado: (2024)
MVDRAM: Enabling GeMV Execution in Unmodified DRAM for Low-Bit LLM Acceleration
por: Kubo, Tatsuya, et al.
Publicado: (2025)
por: Kubo, Tatsuya, et al.
Publicado: (2025)
Sequence-Aware Split Heuristic to Mitigate SM Underutilization in FlashAttention-3 Low-Head-Count Decoding
por: Font, Martí Llopart, et al.
Publicado: (2026)
por: Font, Martí Llopart, et al.
Publicado: (2026)
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
por: Zhang, Chen, et al.
Publicado: (2026)
por: Zhang, Chen, et al.
Publicado: (2026)
SpArch: Efficient Architecture for Sparse Matrix Multiplication
por: Zhang, Zhekai, et al.
Publicado: (2020)
por: Zhang, Zhekai, et al.
Publicado: (2020)
Observation, Not Prediction: Conversation-Level Disaggregated Scheduling for Agentic Serving
por: Ding, Jianru, et al.
Publicado: (2026)
por: Ding, Jianru, et al.
Publicado: (2026)
A Lightweight High-Throughput Collective-Capable NoC for Large-Scale ML Accelerators
por: Colagrande, Luca, et al.
Publicado: (2026)
por: Colagrande, Luca, et al.
Publicado: (2026)
TAPA-CS: Enabling Scalable Accelerator Design on Distributed HBM-FPGAs
por: Prakriya, Neha, et al.
Publicado: (2023)
por: Prakriya, Neha, et al.
Publicado: (2023)
PULSAR: Simultaneous Many-Row Activation for Reliable and High-Performance Computing in Off-the-Shelf DRAM Chips
por: Yuksel, Ismail Emir, et al.
Publicado: (2023)
por: Yuksel, Ismail Emir, et al.
Publicado: (2023)
Proteus: Enabling High-Performance Processing-Using-DRAM with Dynamic Bit-Precision, Adaptive Data Representation, and Flexible Arithmetic
por: Oliveira, Geraldo F., et al.
Publicado: (2025)
por: Oliveira, Geraldo F., et al.
Publicado: (2025)
ALPHA-PIM: Analysis of Linear Algebraic Processing for High-Performance Graph Applications on a Real Processing-In-Memory System
por: Barkhordar, Marzieh, et al.
Publicado: (2026)
por: Barkhordar, Marzieh, et al.
Publicado: (2026)
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
por: li, Fei, et al.
Publicado: (2026)
por: li, Fei, et al.
Publicado: (2026)
MIMDRAM: An End-to-End Processing-Using-DRAM System for High-Throughput, Energy-Efficient and Programmer-Transparent Multiple-Instruction Multiple-Data Processing
por: Oliveira, Geraldo F., et al.
Publicado: (2024)
por: Oliveira, Geraldo F., et al.
Publicado: (2024)
TeraPool: A Physical Design Aware, 1024 RISC-V Cores Shared-L1-Memory Scaled-up Cluster Design with High Bandwidth Main Memory Link
por: Zhang, Yichao, et al.
Publicado: (2026)
por: Zhang, Yichao, et al.
Publicado: (2026)
MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems
por: Zhou, Zhuoshan, et al.
Publicado: (2026)
por: Zhou, Zhuoshan, et al.
Publicado: (2026)
Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs
por: Zhang, Qijun, et al.
Publicado: (2026)
por: Zhang, Qijun, et al.
Publicado: (2026)
Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
por: Lin, Bin, et al.
Publicado: (2024)
por: Lin, Bin, et al.
Publicado: (2024)
Deep Learning and Machine Learning with GPGPU and CUDA: Unlocking the Power of Parallel Computing
por: Li, Ming, et al.
Publicado: (2024)
por: Li, Ming, et al.
Publicado: (2024)
TT-Edge: A Hardware-Software Co-Design for Energy-Efficient Tensor-Train Decomposition on Edge AI
por: Kwak, Hyunseok, et al.
Publicado: (2025)
por: Kwak, Hyunseok, et al.
Publicado: (2025)
The DEEP-ER project: I/O and resiliency extensions for the Cluster-Booster architecture
por: Kreuzer, Anke, et al.
Publicado: (2019)
por: Kreuzer, Anke, et al.
Publicado: (2019)
An Evaluation and Comparison of GPU Hardware and Solver Libraries for Accelerating the OPM Flow Reservoir Simulator
por: Qiu, Tong Dong, et al.
Publicado: (2023)
por: Qiu, Tong Dong, et al.
Publicado: (2023)
Ejemplares similares
-
PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving
por: Yüzügüler, Ahmet Caner, et al.
Publicado: (2025) -
PUDTune: Multi-Level Charging for High-Precision Calibration in Processing-Using-DRAM
por: Kubo, Tatsuya, et al.
Publicado: (2025) -
DP-HLS: A High-Level Synthesis Framework for Accelerating Dynamic Programming Algorithms in Bioinformatics
por: Cao, Yingqi, et al.
Publicado: (2024) -
MLDSE: Scaling Design Space Exploration Infrastructure for Multi-Level Hardware
por: Qu, Huanyu, et al.
Publicado: (2025) -
MegIS: High-Performance, Energy-Efficient, and Low-Cost Metagenomic Analysis with In-Storage Processing
por: Ghiasi, Nika Mansouri, et al.
Publicado: (2024)