Reinforcement Learning for Dynamic Memory Allocation
Fuente:
arXiv
Saved in:
| Main Authors: | Lim, Arisrei, Maddukuri, Abhiram |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention
by: Prabhu, Ramya, et al.
Published: (2024)
by: Prabhu, Ramya, et al.
Published: (2024)
Energy-Efficient Computation with DVFS using Deep Reinforcement Learning for Multi-Task Systems in Edge Computing
by: Li, Xinyi, et al.
Published: (2024)
by: Li, Xinyi, et al.
Published: (2024)
Dynamic Optimization of Storage Systems Using Reinforcement Learning Techniques
by: Cheng, Chiyu, et al.
Published: (2024)
by: Cheng, Chiyu, et al.
Published: (2024)
Machine Learning (ML) library in Linux kernel
by: Dubeyko, Viacheslav
Published: (2026)
by: Dubeyko, Viacheslav
Published: (2026)
TempoNet: Slack-Quantized Transformer-Guided Reinforcement Scheduler for Adaptive Deadline-Centric Real-Time Dispatchs
by: Fu, Rong, et al.
Published: (2026)
by: Fu, Rong, et al.
Published: (2026)
LithOS: An Operating System for Efficient Machine Learning on GPUs
by: Coppock, Patrick H., et al.
Published: (2025)
by: Coppock, Patrick H., et al.
Published: (2025)
SARA: A Stall-Aware Memory Allocation Strategy for Mixed-Criticality Systems
by: Lee, Meng-Chia, et al.
Published: (2025)
by: Lee, Meng-Chia, et al.
Published: (2025)
Generative Profiling for Soft Real-Time Systems and its Applications to Resource Allocation
by: Bondar, Georgiy A., et al.
Published: (2026)
by: Bondar, Georgiy A., et al.
Published: (2026)
Puzzle: Scheduling Multiple Deep Learning Models on Mobile Device with Heterogeneous Processors
by: Kang, Duseok, et al.
Published: (2025)
by: Kang, Duseok, et al.
Published: (2025)
MaLV-OS: Rethinking the Operating System Architecture for Machine Learning in Virtualized Clouds
by: Bitchebe, Stella, et al.
Published: (2025)
by: Bitchebe, Stella, et al.
Published: (2025)
Enhancing Battery Storage Energy Arbitrage with Deep Reinforcement Learning and Time-Series Forecasting
by: Sage, Manuel, et al.
Published: (2024)
by: Sage, Manuel, et al.
Published: (2024)
Learning Semantics, Not Addresses: Runtime Neural Prefetching for Far Memory
by: Huang, Yutong, et al.
Published: (2025)
by: Huang, Yutong, et al.
Published: (2025)
Herding LLaMaS: Using LLMs as an OS Module
by: Kamath, Aditya K, et al.
Published: (2024)
by: Kamath, Aditya K, et al.
Published: (2024)
Crash-Consistent Checkpointing for AI Training on macOS/APFS
by: Jeon, Juha
Published: (2025)
by: Jeon, Juha
Published: (2025)
PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
by: Song, Yixin, et al.
Published: (2023)
by: Song, Yixin, et al.
Published: (2023)
Accelerated Training on Low-Power Edge Devices
by: Ahmed, Mohamed Aboelenien, et al.
Published: (2025)
by: Ahmed, Mohamed Aboelenien, et al.
Published: (2025)
When eBPF Meets Machine Learning: On-the-fly OS Kernel Compartmentalization
by: Wang, Zicheng, et al.
Published: (2024)
by: Wang, Zicheng, et al.
Published: (2024)
Dynamic Adaptation in Data Storage: Real-Time Machine Learning for Enhanced Prefetching
by: Cheng, Chiyu, et al.
Published: (2024)
by: Cheng, Chiyu, et al.
Published: (2024)
Bauplan: zero-copy, scale-up FaaS for data pipelines
by: Tagliabue, Jacopo, et al.
Published: (2024)
by: Tagliabue, Jacopo, et al.
Published: (2024)
Leveraging Machine Learning for Accurate IoT Device Identification in Dynamic Wireless Contexts
by: Tushir, Bhagyashri, et al.
Published: (2024)
by: Tushir, Bhagyashri, et al.
Published: (2024)
An Integrated Artificial Intelligence Operating System for Advanced Low-Altitude Aviation Applications
by: Tan, Minzhe, et al.
Published: (2024)
by: Tan, Minzhe, et al.
Published: (2024)
FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation
by: Wu, Yinpeng, et al.
Published: (2026)
by: Wu, Yinpeng, et al.
Published: (2026)
OSWorld-Human: Benchmarking the Efficiency of Computer-Use Agents
by: Abhyankar, Reyna, et al.
Published: (2025)
by: Abhyankar, Reyna, et al.
Published: (2025)
AdaptCache: KV Cache Native Storage Hierarchy for Low-Delay and High-Quality Language Model Serving
by: Feng, Shaoting, et al.
Published: (2025)
by: Feng, Shaoting, et al.
Published: (2025)
Semantic Scheduling for LLM Inference
by: Hua, Wenyue, et al.
Published: (2025)
by: Hua, Wenyue, et al.
Published: (2025)
Selective KV-Cache Sharing to Mitigate Timing Side-Channels in LLM Inference
by: Chu, Kexin, et al.
Published: (2025)
by: Chu, Kexin, et al.
Published: (2025)
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
by: Feng, Shaoting, et al.
Published: (2025)
by: Feng, Shaoting, et al.
Published: (2025)
From Imperative to Declarative: Towards LLM-friendly OS Interfaces for Boosted Computer-Use Agents
by: Wang, Yuan, et al.
Published: (2025)
by: Wang, Yuan, et al.
Published: (2025)
Preparation Meets Opportunity: Enhancing Data Preprocessing for ML Training With Seneca
by: Desai, Omkar, et al.
Published: (2025)
by: Desai, Omkar, et al.
Published: (2025)
Optimizing SSD Caches for Cloud Block Storage Systems Using Machine Learning Approaches
by: Cheng, Chiyu, et al.
Published: (2024)
by: Cheng, Chiyu, et al.
Published: (2024)
Head-First Memory Allocation on Best-Fit with Space-Fitting
by: Hakarsa, Adam Noto
Published: (2024)
by: Hakarsa, Adam Noto
Published: (2024)
An Online Gradient-Based Caching Policy with Logarithmic Complexity and Regret Guarantees
by: Carra, Damiano, et al.
Published: (2024)
by: Carra, Damiano, et al.
Published: (2024)
Neuralink: Fast LLM Inference on Smartphones with Neuron Co-Activation Linking
by: Wang, Tuowei, et al.
Published: (2024)
by: Wang, Tuowei, et al.
Published: (2024)
Sawtooth Wavefront Reordering: Enhanced CuTile FlashAttention on NVIDIA GB10
by: Zhu, Yifan, et al.
Published: (2026)
by: Zhu, Yifan, et al.
Published: (2026)
An Early Exploration of Deep-Learning-Driven Prefetching for Far Memory
by: Huang, Yutong, et al.
Published: (2025)
by: Huang, Yutong, et al.
Published: (2025)
Enhancing Adaptive Mixed-Criticality Scheduling with Deep Reinforcement Learning
by: Mendes, Bruno, et al.
Published: (2024)
by: Mendes, Bruno, et al.
Published: (2024)
Talyxion: From Speculation to Optimization in Risk Managed Crypto Portfolio Allocation
by: Nguyen, Thanh
Published: (2025)
by: Nguyen, Thanh
Published: (2025)
E-Mapper: Energy-Efficient Resource Allocation for Traditional Operating Systems on Heterogeneous Processors
by: Smejkal, Till, et al.
Published: (2024)
by: Smejkal, Till, et al.
Published: (2024)
Everything You Always Wanted to Know About Storage Compressibility of Pre-Trained ML Models but Were Afraid to Ask
by: Su, Zhaoyuan, et al.
Published: (2024)
by: Su, Zhaoyuan, et al.
Published: (2024)
Hardware-Assisted Virtualization of Neural Processing Units for Cloud Platforms
by: Xue, Yuqi, et al.
Published: (2024)
by: Xue, Yuqi, et al.
Published: (2024)
Similar Items
-
vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention
by: Prabhu, Ramya, et al.
Published: (2024) -
Energy-Efficient Computation with DVFS using Deep Reinforcement Learning for Multi-Task Systems in Edge Computing
by: Li, Xinyi, et al.
Published: (2024) -
Dynamic Optimization of Storage Systems Using Reinforcement Learning Techniques
by: Cheng, Chiyu, et al.
Published: (2024) -
Machine Learning (ML) library in Linux kernel
by: Dubeyko, Viacheslav
Published: (2026) -
TempoNet: Slack-Quantized Transformer-Guided Reinforcement Scheduler for Adaptive Deadline-Centric Real-Time Dispatchs
by: Fu, Rong, et al.
Published: (2026)