Herding LLaMaS: Using LLMs as an OS Module
Fuente:
arXiv
Salvato in:
| Autori principali: | Kamath, Aditya K, Yadalam, Sujay |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ARMS: Adaptive and Robust Memory Tiering System
di: Yadalam, Sujay, et al.
Pubblicazione: (2025)
di: Yadalam, Sujay, et al.
Pubblicazione: (2025)
MaLV-OS: Rethinking the Operating System Architecture for Machine Learning in Virtualized Clouds
di: Bitchebe, Stella, et al.
Pubblicazione: (2025)
di: Bitchebe, Stella, et al.
Pubblicazione: (2025)
From Good to Great: Improving Memory Tiering Performance Through Parameter Tuning
di: Kanellis, Konstantinos, et al.
Pubblicazione: (2025)
di: Kanellis, Konstantinos, et al.
Pubblicazione: (2025)
Vulcan: Instance-Optimal Systems Heuristics Through LLM-Driven Search
di: Dwivedula, Rohit, et al.
Pubblicazione: (2025)
di: Dwivedula, Rohit, et al.
Pubblicazione: (2025)
Crash-Consistent Checkpointing for AI Training on macOS/APFS
di: Jeon, Juha
Pubblicazione: (2025)
di: Jeon, Juha
Pubblicazione: (2025)
LithOS: An Operating System for Efficient Machine Learning on GPUs
di: Coppock, Patrick H., et al.
Pubblicazione: (2025)
di: Coppock, Patrick H., et al.
Pubblicazione: (2025)
vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention
di: Prabhu, Ramya, et al.
Pubblicazione: (2024)
di: Prabhu, Ramya, et al.
Pubblicazione: (2024)
When eBPF Meets Machine Learning: On-the-fly OS Kernel Compartmentalization
di: Wang, Zicheng, et al.
Pubblicazione: (2024)
di: Wang, Zicheng, et al.
Pubblicazione: (2024)
From Imperative to Declarative: Towards LLM-friendly OS Interfaces for Boosted Computer-Use Agents
di: Wang, Yuan, et al.
Pubblicazione: (2025)
di: Wang, Yuan, et al.
Pubblicazione: (2025)
Reinforcement Learning for Dynamic Memory Allocation
di: Lim, Arisrei, et al.
Pubblicazione: (2024)
di: Lim, Arisrei, et al.
Pubblicazione: (2024)
Energy-Efficient Computation with DVFS using Deep Reinforcement Learning for Multi-Task Systems in Edge Computing
di: Li, Xinyi, et al.
Pubblicazione: (2024)
di: Li, Xinyi, et al.
Pubblicazione: (2024)
Puzzle: Scheduling Multiple Deep Learning Models on Mobile Device with Heterogeneous Processors
di: Kang, Duseok, et al.
Pubblicazione: (2025)
di: Kang, Duseok, et al.
Pubblicazione: (2025)
PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
di: Song, Yixin, et al.
Pubblicazione: (2023)
di: Song, Yixin, et al.
Pubblicazione: (2023)
Machine Learning (ML) library in Linux kernel
di: Dubeyko, Viacheslav
Pubblicazione: (2026)
di: Dubeyko, Viacheslav
Pubblicazione: (2026)
Accelerated Training on Low-Power Edge Devices
di: Ahmed, Mohamed Aboelenien, et al.
Pubblicazione: (2025)
di: Ahmed, Mohamed Aboelenien, et al.
Pubblicazione: (2025)
TempoNet: Slack-Quantized Transformer-Guided Reinforcement Scheduler for Adaptive Deadline-Centric Real-Time Dispatchs
di: Fu, Rong, et al.
Pubblicazione: (2026)
di: Fu, Rong, et al.
Pubblicazione: (2026)
Bauplan: zero-copy, scale-up FaaS for data pipelines
di: Tagliabue, Jacopo, et al.
Pubblicazione: (2024)
di: Tagliabue, Jacopo, et al.
Pubblicazione: (2024)
An Integrated Artificial Intelligence Operating System for Advanced Low-Altitude Aviation Applications
di: Tan, Minzhe, et al.
Pubblicazione: (2024)
di: Tan, Minzhe, et al.
Pubblicazione: (2024)
FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation
di: Wu, Yinpeng, et al.
Pubblicazione: (2026)
di: Wu, Yinpeng, et al.
Pubblicazione: (2026)
OSWorld-Human: Benchmarking the Efficiency of Computer-Use Agents
di: Abhyankar, Reyna, et al.
Pubblicazione: (2025)
di: Abhyankar, Reyna, et al.
Pubblicazione: (2025)
AdaptCache: KV Cache Native Storage Hierarchy for Low-Delay and High-Quality Language Model Serving
di: Feng, Shaoting, et al.
Pubblicazione: (2025)
di: Feng, Shaoting, et al.
Pubblicazione: (2025)
Semantic Scheduling for LLM Inference
di: Hua, Wenyue, et al.
Pubblicazione: (2025)
di: Hua, Wenyue, et al.
Pubblicazione: (2025)
Selective KV-Cache Sharing to Mitigate Timing Side-Channels in LLM Inference
di: Chu, Kexin, et al.
Pubblicazione: (2025)
di: Chu, Kexin, et al.
Pubblicazione: (2025)
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
di: Feng, Shaoting, et al.
Pubblicazione: (2025)
di: Feng, Shaoting, et al.
Pubblicazione: (2025)
Preparation Meets Opportunity: Enhancing Data Preprocessing for ML Training With Seneca
di: Desai, Omkar, et al.
Pubblicazione: (2025)
di: Desai, Omkar, et al.
Pubblicazione: (2025)
Dynamic Optimization of Storage Systems Using Reinforcement Learning Techniques
di: Cheng, Chiyu, et al.
Pubblicazione: (2024)
di: Cheng, Chiyu, et al.
Pubblicazione: (2024)
Optimizing SSD Caches for Cloud Block Storage Systems Using Machine Learning Approaches
di: Cheng, Chiyu, et al.
Pubblicazione: (2024)
di: Cheng, Chiyu, et al.
Pubblicazione: (2024)
Enhancing Battery Storage Energy Arbitrage with Deep Reinforcement Learning and Time-Series Forecasting
di: Sage, Manuel, et al.
Pubblicazione: (2024)
di: Sage, Manuel, et al.
Pubblicazione: (2024)
Generative Profiling for Soft Real-Time Systems and its Applications to Resource Allocation
di: Bondar, Georgiy A., et al.
Pubblicazione: (2026)
di: Bondar, Georgiy A., et al.
Pubblicazione: (2026)
An Online Gradient-Based Caching Policy with Logarithmic Complexity and Regret Guarantees
di: Carra, Damiano, et al.
Pubblicazione: (2024)
di: Carra, Damiano, et al.
Pubblicazione: (2024)
Neuralink: Fast LLM Inference on Smartphones with Neuron Co-Activation Linking
di: Wang, Tuowei, et al.
Pubblicazione: (2024)
di: Wang, Tuowei, et al.
Pubblicazione: (2024)
Sawtooth Wavefront Reordering: Enhanced CuTile FlashAttention on NVIDIA GB10
di: Zhu, Yifan, et al.
Pubblicazione: (2026)
di: Zhu, Yifan, et al.
Pubblicazione: (2026)
Exploiting Application-to-Architecture Dependencies for Designing Scalable OS
di: Xiao, Yao, et al.
Pubblicazione: (2025)
di: Xiao, Yao, et al.
Pubblicazione: (2025)
Man-Made Heuristics Are Dead. Long Live Code Generators!
di: Dwivedula, Rohit, et al.
Pubblicazione: (2025)
di: Dwivedula, Rohit, et al.
Pubblicazione: (2025)
TenonOS: A Self-Generating LibOS-on-LibOS Framework for Time-Critical Embedded Operating Systems
di: Zhao, Xinkui, et al.
Pubblicazione: (2025)
di: Zhao, Xinkui, et al.
Pubblicazione: (2025)
Leveraging Machine Learning for Accurate IoT Device Identification in Dynamic Wireless Contexts
di: Tushir, Bhagyashri, et al.
Pubblicazione: (2024)
di: Tushir, Bhagyashri, et al.
Pubblicazione: (2024)
Attention, Distillation, and Tabularization: Towards Practical Neural Network-Based Prefetching
di: Zhang, Pengmiao, et al.
Pubblicazione: (2023)
di: Zhang, Pengmiao, et al.
Pubblicazione: (2023)
Everything You Always Wanted to Know About Storage Compressibility of Pre-Trained ML Models but Were Afraid to Ask
di: Su, Zhaoyuan, et al.
Pubblicazione: (2024)
di: Su, Zhaoyuan, et al.
Pubblicazione: (2024)
Dynamic Adaptation in Data Storage: Real-Time Machine Learning for Enhanced Prefetching
di: Cheng, Chiyu, et al.
Pubblicazione: (2024)
di: Cheng, Chiyu, et al.
Pubblicazione: (2024)
Hardware-Assisted Virtualization of Neural Processing Units for Cloud Platforms
di: Xue, Yuqi, et al.
Pubblicazione: (2024)
di: Xue, Yuqi, et al.
Pubblicazione: (2024)
Documenti analoghi
-
ARMS: Adaptive and Robust Memory Tiering System
di: Yadalam, Sujay, et al.
Pubblicazione: (2025) -
MaLV-OS: Rethinking the Operating System Architecture for Machine Learning in Virtualized Clouds
di: Bitchebe, Stella, et al.
Pubblicazione: (2025) -
From Good to Great: Improving Memory Tiering Performance Through Parameter Tuning
di: Kanellis, Konstantinos, et al.
Pubblicazione: (2025) -
Vulcan: Instance-Optimal Systems Heuristics Through LLM-Driven Search
di: Dwivedula, Rohit, et al.
Pubblicazione: (2025) -
Crash-Consistent Checkpointing for AI Training on macOS/APFS
di: Jeon, Juha
Pubblicazione: (2025)