Efficient Function-as-a-Service for Large Language Models with TIDAL
Fuente:
arXiv
Saved in:
| Main Authors: | Cui, Weihao, Xu, Ziyi, Zhao, Han, Chen, Quan, Li, Zijun, He, Bingsheng, Guo, Minyi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Flare: Anomaly Diagnostics for Divergent LLM Training in GPU Clusters of Thousand-Plus Scale
by: Cui, Weihao, et al.
Published: (2025)
by: Cui, Weihao, et al.
Published: (2025)
Towards High-Goodput LLM Serving with Prefill-decode Multiplexing
by: Chen, Yukang, et al.
Published: (2025)
by: Chen, Yukang, et al.
Published: (2025)
Trustworthy and Controllable Professional Knowledge Utilization in Large Language Models with TEE-GPU Execution
by: Cai, Yifeng, et al.
Published: (2025)
by: Cai, Yifeng, et al.
Published: (2025)
LLM as a System Service on Mobile Devices
by: Yin, Wangsong, et al.
Published: (2024)
by: Yin, Wangsong, et al.
Published: (2024)
TIDAL: Recovering Temporal Phase for Cloud Block Storage Placement from LLM-Derived Semantics
by: Tan, Difan, et al.
Published: (2026)
by: Tan, Difan, et al.
Published: (2026)
Exploiting Dependency and Parallelism: Real-Time Scheduling and Analysis for GPU Tasks
by: Zhang, Yuanhai, et al.
Published: (2026)
by: Zhang, Yuanhai, et al.
Published: (2026)
Nova: Real-Time Agentic Vision-Language Model Serving with Adaptive Cross-Stage Parallelization
by: Xu, Yuhang, et al.
Published: (2025)
by: Xu, Yuhang, et al.
Published: (2025)
ROSfs: A User-Level File System for ROS
by: Xu, Zijun, et al.
Published: (2024)
by: Xu, Zijun, et al.
Published: (2024)
GoCkpt: Gradient-Assisted Multi-Step overlapped Checkpointing for Efficient LLM Training
by: Zhang, Keyao, et al.
Published: (2025)
by: Zhang, Keyao, et al.
Published: (2025)
PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
by: Song, Yixin, et al.
Published: (2023)
by: Song, Yixin, et al.
Published: (2023)
Configuration Validation with Large Language Models
by: Lian, Xinyu, et al.
Published: (2023)
by: Lian, Xinyu, et al.
Published: (2023)
Concurrency Testing in the Linux Kernel via eBPF
by: Xu, Jiacheng, et al.
Published: (2025)
by: Xu, Jiacheng, et al.
Published: (2025)
FRAP: A Flexible Resource Accessing Protocol for Multiprocessor Real-Time Systems
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
Dissecting CXL Memory Performance at Scale: Analysis, Modeling, and Optimization
by: Liu, Jinshu, et al.
Published: (2024)
by: Liu, Jinshu, et al.
Published: (2024)
MVVM: Deploy Your AI Agents-Securely, Efficiently, Everywhere
by: Yang, Yiwei, et al.
Published: (2024)
by: Yang, Yiwei, et al.
Published: (2024)
Taiji: A DPU Memory Elasticity Solution for In-production Cloud Environments
by: Zheng, Hao, et al.
Published: (2025)
by: Zheng, Hao, et al.
Published: (2025)
Efficient Memory Tiering in a Virtual Machine
by: Prakash, Chandra, et al.
Published: (2025)
by: Prakash, Chandra, et al.
Published: (2025)
MigGPT: Harnessing Large Language Models for Automated Migration of Out-of-Tree Linux Kernel Patches Across Versions
by: Dang, Pucheng, et al.
Published: (2025)
by: Dang, Pucheng, et al.
Published: (2025)
PipeANN-Filter: An Efficient Filtered Vector Search System on SSD
by: Guo, Hao, et al.
Published: (2026)
by: Guo, Hao, et al.
Published: (2026)
BYOS: Knowledge-driven Large Language Models Bring Your Own Operating System More Excellent
by: Lin, Hongyu, et al.
Published: (2025)
by: Lin, Hongyu, et al.
Published: (2025)
Holistic Heterogeneous Scheduling for Autonomous Applications using Fine-grained, Multi-XPU Abstraction
by: Han, Mingcong, et al.
Published: (2025)
by: Han, Mingcong, et al.
Published: (2025)
XBOF: A Cost-Efficient CXL JBOF with Inter-SSD Compute Resource Sharing
by: Yi, Shushu, et al.
Published: (2025)
by: Yi, Shushu, et al.
Published: (2025)
Leveraging OS-Level Primitives for Robotic Action Management
by: Zheng, Wenxin, et al.
Published: (2025)
by: Zheng, Wenxin, et al.
Published: (2025)
Squeezy: Rapid VM Memory Reclamation for Serverless Functions
by: Nikolos, Orestis Lagkas, et al.
Published: (2024)
by: Nikolos, Orestis Lagkas, et al.
Published: (2024)
Idleness is Relative: Exploiting Tool-Call Idle Windows for Offloading in Agentic Systems with MORI
by: Xia, Tian, et al.
Published: (2026)
by: Xia, Tian, et al.
Published: (2026)
HeteroPod: XPU-Accelerated Infrastructure Offloading for Commodity Cloud-Native Applications
by: Yang, Bicheng, et al.
Published: (2025)
by: Yang, Bicheng, et al.
Published: (2025)
Vmem: A Lightweight Hot-Upgradable Memory Management for In-production Cloud Environment
by: Zheng, Hao, et al.
Published: (2025)
by: Zheng, Hao, et al.
Published: (2025)
Optimizing Tree-structure Indexes for CXL-based Heterogeneous Memory with SINLK
by: Zhao, Haoru, et al.
Published: (2025)
by: Zhao, Haoru, et al.
Published: (2025)
DeLog: An Efficient Log Compression Framework with Pattern Signature Synthesis
by: Yu, Siyu, et al.
Published: (2026)
by: Yu, Siyu, et al.
Published: (2026)
SSV: Sparse Speculative Verification for Efficient LLM Inference
by: Wang, Zhibin, et al.
Published: (2026)
by: Wang, Zhibin, et al.
Published: (2026)
ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference System
by: Kang, Hao, et al.
Published: (2026)
by: Kang, Hao, et al.
Published: (2026)
Boosting File Systems Elegantly: A Transparent NVM Write-ahead Log for Disk File Systems
by: Wang, Guoyu, et al.
Published: (2024)
by: Wang, Guoyu, et al.
Published: (2024)
SVFF: An Automated Framework for SR-IOV Virtual Function Management in FPGA Accelerated Virtualized Environments
by: Cirici, Stefano, et al.
Published: (2024)
by: Cirici, Stefano, et al.
Published: (2024)
Asterinas: A Linux ABI-Compatible, Rust-Based Framekernel OS with a Small and Sound TCB
by: Peng, Yuke, et al.
Published: (2025)
by: Peng, Yuke, et al.
Published: (2025)
E-Mapper: Energy-Efficient Resource Allocation for Traditional Operating Systems on Heterogeneous Processors
by: Smejkal, Till, et al.
Published: (2024)
by: Smejkal, Till, et al.
Published: (2024)
AERO: Adaptive and Efficient Runtime-Aware OTA Updates for Energy-Harvesting IoT
by: Wei, Wei, et al.
Published: (2026)
by: Wei, Wei, et al.
Published: (2026)
SemaTune: Semantic-Aware Online OS Tuning with Large Language Models
by: Liargkovas, Georgios, et al.
Published: (2026)
by: Liargkovas, Georgios, et al.
Published: (2026)
AppFlow: Memory Scheduling for Cold Launch of Large Apps on Mobile and Vehicle Systems
by: Li, Xiaochen, et al.
Published: (2026)
by: Li, Xiaochen, et al.
Published: (2026)
An Early Exploration of Deep-Learning-Driven Prefetching for Far Memory
by: Huang, Yutong, et al.
Published: (2025)
by: Huang, Yutong, et al.
Published: (2025)
BULKHEAD: Secure, Scalable, and Efficient Kernel Compartmentalization with PKS
by: Guo, Yinggang, et al.
Published: (2024)
by: Guo, Yinggang, et al.
Published: (2024)
Similar Items
-
Flare: Anomaly Diagnostics for Divergent LLM Training in GPU Clusters of Thousand-Plus Scale
by: Cui, Weihao, et al.
Published: (2025) -
Towards High-Goodput LLM Serving with Prefill-decode Multiplexing
by: Chen, Yukang, et al.
Published: (2025) -
Trustworthy and Controllable Professional Knowledge Utilization in Large Language Models with TEE-GPU Execution
by: Cai, Yifeng, et al.
Published: (2025) -
LLM as a System Service on Mobile Devices
by: Yin, Wangsong, et al.
Published: (2024) -
TIDAL: Recovering Temporal Phase for Cloud Block Storage Placement from LLM-Derived Semantics
by: Tan, Difan, et al.
Published: (2026)