vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Prabhu, Ramya, Nayak, Ajay, Mohan, Jayashree, Ramjee, Ramachandran, Panwar, Ashish |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference
von: Kamath, Aditya K, et al.
Veröffentlicht: (2024)
von: Kamath, Aditya K, et al.
Veröffentlicht: (2024)
vAttention: Verified Sparse Attention
von: Desai, Aditya, et al.
Veröffentlicht: (2025)
von: Desai, Aditya, et al.
Veröffentlicht: (2025)
Reinforcement Learning for Dynamic Memory Allocation
von: Lim, Arisrei, et al.
Veröffentlicht: (2024)
von: Lim, Arisrei, et al.
Veröffentlicht: (2024)
Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
von: Agrawal, Amey, et al.
Veröffentlicht: (2024)
von: Agrawal, Amey, et al.
Veröffentlicht: (2024)
PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
von: Song, Yixin, et al.
Veröffentlicht: (2023)
von: Song, Yixin, et al.
Veröffentlicht: (2023)
FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation
von: Wu, Yinpeng, et al.
Veröffentlicht: (2026)
von: Wu, Yinpeng, et al.
Veröffentlicht: (2026)
Herding LLaMaS: Using LLMs as an OS Module
von: Kamath, Aditya K, et al.
Veröffentlicht: (2024)
von: Kamath, Aditya K, et al.
Veröffentlicht: (2024)
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
von: Feng, Shaoting, et al.
Veröffentlicht: (2025)
von: Feng, Shaoting, et al.
Veröffentlicht: (2025)
Sawtooth Wavefront Reordering: Enhanced CuTile FlashAttention on NVIDIA GB10
von: Zhu, Yifan, et al.
Veröffentlicht: (2026)
von: Zhu, Yifan, et al.
Veröffentlicht: (2026)
AdaptCache: KV Cache Native Storage Hierarchy for Low-Delay and High-Quality Language Model Serving
von: Feng, Shaoting, et al.
Veröffentlicht: (2025)
von: Feng, Shaoting, et al.
Veröffentlicht: (2025)
Niyama : Breaking the Silos of LLM Inference Serving
von: Goel, Kanishk, et al.
Veröffentlicht: (2025)
von: Goel, Kanishk, et al.
Veröffentlicht: (2025)
Attention, Distillation, and Tabularization: Towards Practical Neural Network-Based Prefetching
von: Zhang, Pengmiao, et al.
Veröffentlicht: (2023)
von: Zhang, Pengmiao, et al.
Veröffentlicht: (2023)
Nomad: Non-Exclusive Memory Tiering via Transactional Page Migration
von: Xiang, Lingfeng, et al.
Veröffentlicht: (2024)
von: Xiang, Lingfeng, et al.
Veröffentlicht: (2024)
numaPTE: Managing Page-Tables and TLBs on NUMA Systems
von: Gao, Bin, et al.
Veröffentlicht: (2024)
von: Gao, Bin, et al.
Veröffentlicht: (2024)
Vidur: A Large-Scale Simulation Framework For LLM Inference
von: Agrawal, Amey, et al.
Veröffentlicht: (2024)
von: Agrawal, Amey, et al.
Veröffentlicht: (2024)
MaLV-OS: Rethinking the Operating System Architecture for Machine Learning in Virtualized Clouds
von: Bitchebe, Stella, et al.
Veröffentlicht: (2025)
von: Bitchebe, Stella, et al.
Veröffentlicht: (2025)
Crash-Consistent Checkpointing for AI Training on macOS/APFS
von: Jeon, Juha
Veröffentlicht: (2025)
von: Jeon, Juha
Veröffentlicht: (2025)
Puzzle: Scheduling Multiple Deep Learning Models on Mobile Device with Heterogeneous Processors
von: Kang, Duseok, et al.
Veröffentlicht: (2025)
von: Kang, Duseok, et al.
Veröffentlicht: (2025)
Machine Learning (ML) library in Linux kernel
von: Dubeyko, Viacheslav
Veröffentlicht: (2026)
von: Dubeyko, Viacheslav
Veröffentlicht: (2026)
Energy-Efficient Computation with DVFS using Deep Reinforcement Learning for Multi-Task Systems in Edge Computing
von: Li, Xinyi, et al.
Veröffentlicht: (2024)
von: Li, Xinyi, et al.
Veröffentlicht: (2024)
LithOS: An Operating System for Efficient Machine Learning on GPUs
von: Coppock, Patrick H., et al.
Veröffentlicht: (2025)
von: Coppock, Patrick H., et al.
Veröffentlicht: (2025)
Accelerated Training on Low-Power Edge Devices
von: Ahmed, Mohamed Aboelenien, et al.
Veröffentlicht: (2025)
von: Ahmed, Mohamed Aboelenien, et al.
Veröffentlicht: (2025)
TierBPF: Page Migration Admission Control for Tiered Memory via eBPF
von: Wang, Xi, et al.
Veröffentlicht: (2026)
von: Wang, Xi, et al.
Veröffentlicht: (2026)
Jenga: Responsive Tiered Memory Management without Thrashing
von: Kadekodi, Rohan, et al.
Veröffentlicht: (2025)
von: Kadekodi, Rohan, et al.
Veröffentlicht: (2025)
TempoNet: Slack-Quantized Transformer-Guided Reinforcement Scheduler for Adaptive Deadline-Centric Real-Time Dispatchs
von: Fu, Rong, et al.
Veröffentlicht: (2026)
von: Fu, Rong, et al.
Veröffentlicht: (2026)
Learning Semantics, Not Addresses: Runtime Neural Prefetching for Far Memory
von: Huang, Yutong, et al.
Veröffentlicht: (2025)
von: Huang, Yutong, et al.
Veröffentlicht: (2025)
LearnedCache: An eBPF-Integrated Perceptron-Based Eviction Policy for the Linux Page Cache
von: Qi, Zejia
Veröffentlicht: (2026)
von: Qi, Zejia
Veröffentlicht: (2026)
Bauplan: zero-copy, scale-up FaaS for data pipelines
von: Tagliabue, Jacopo, et al.
Veröffentlicht: (2024)
von: Tagliabue, Jacopo, et al.
Veröffentlicht: (2024)
ASTRA: Accurate and Scalable ANNS-based Training of Extreme Classifiers
von: Mehta, Sonu, et al.
Veröffentlicht: (2024)
von: Mehta, Sonu, et al.
Veröffentlicht: (2024)
Token Management in Multi-Tenant AI Inference Platforms
von: Cunningham, William J.
Veröffentlicht: (2026)
von: Cunningham, William J.
Veröffentlicht: (2026)
LLM-42: Enabling Determinism in LLM Inference with Verified Speculation
von: Gond, Raja, et al.
Veröffentlicht: (2026)
von: Gond, Raja, et al.
Veröffentlicht: (2026)
Adaptive and Efficient Dynamic Memory Management for Hardware Enclaves
von: Dhanraj, Vijay, et al.
Veröffentlicht: (2025)
von: Dhanraj, Vijay, et al.
Veröffentlicht: (2025)
Virtual-Memory Assisted Buffer Management In Tiered Memory
von: Rayhan, Yeasir, et al.
Veröffentlicht: (2026)
von: Rayhan, Yeasir, et al.
Veröffentlicht: (2026)
Cache is King: Smart Page Eviction with eBPF
von: Zussman, Tal, et al.
Veröffentlicht: (2025)
von: Zussman, Tal, et al.
Veröffentlicht: (2025)
On Evaluating Performance of LLM Inference Serving Systems
von: Agrawal, Amey, et al.
Veröffentlicht: (2025)
von: Agrawal, Amey, et al.
Veröffentlicht: (2025)
Blink: CPU-Free LLM Inference by Delegating the Serving Stack to GPU and SmartNIC
von: Siavashi, Mohammad, et al.
Veröffentlicht: (2026)
von: Siavashi, Mohammad, et al.
Veröffentlicht: (2026)
Dynamic Optimization of Storage Systems Using Reinforcement Learning Techniques
von: Cheng, Chiyu, et al.
Veröffentlicht: (2024)
von: Cheng, Chiyu, et al.
Veröffentlicht: (2024)
When eBPF Meets Machine Learning: On-the-fly OS Kernel Compartmentalization
von: Wang, Zicheng, et al.
Veröffentlicht: (2024)
von: Wang, Zicheng, et al.
Veröffentlicht: (2024)
OSWorld-Human: Benchmarking the Efficiency of Computer-Use Agents
von: Abhyankar, Reyna, et al.
Veröffentlicht: (2025)
von: Abhyankar, Reyna, et al.
Veröffentlicht: (2025)
An Integrated Artificial Intelligence Operating System for Advanced Low-Altitude Aviation Applications
von: Tan, Minzhe, et al.
Veröffentlicht: (2024)
von: Tan, Minzhe, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference
von: Kamath, Aditya K, et al.
Veröffentlicht: (2024) -
vAttention: Verified Sparse Attention
von: Desai, Aditya, et al.
Veröffentlicht: (2025) -
Reinforcement Learning for Dynamic Memory Allocation
von: Lim, Arisrei, et al.
Veröffentlicht: (2024) -
Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
von: Agrawal, Amey, et al.
Veröffentlicht: (2024) -
PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
von: Song, Yixin, et al.
Veröffentlicht: (2023)