Nova: Real-Time Agentic Vision-Language Model Serving with Adaptive Cross-Stage Parallelization
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Yuhang, Liu, Shengzhong, Zhang, Dong, Yan, Bingheng, Wu, Fan, Chen, Guihai |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploiting Dependency and Parallelism: Real-Time Scheduling and Analysis for GPU Tasks
by: Zhang, Yuanhai, et al.
Published: (2026)
by: Zhang, Yuanhai, et al.
Published: (2026)
Towards High-Goodput LLM Serving with Prefill-decode Multiplexing
by: Chen, Yukang, et al.
Published: (2025)
by: Chen, Yukang, et al.
Published: (2025)
Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving
by: Qiu, Shi, et al.
Published: (2026)
by: Qiu, Shi, et al.
Published: (2026)
FRAP: A Flexible Resource Accessing Protocol for Multiprocessor Real-Time Systems
by: Zhao, Shuai, et al.
Published: (2024)
by: Zhao, Shuai, et al.
Published: (2024)
PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
by: Song, Yixin, et al.
Published: (2023)
by: Song, Yixin, et al.
Published: (2023)
A Task Equalization Allocation Algorithm Incorporating Blocking Estimation and Resource Similarity Analysis for Vehicle Control Real-Time Systems
by: Duan, Qianlong, et al.
Published: (2025)
by: Duan, Qianlong, et al.
Published: (2025)
ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference System
by: Kang, Hao, et al.
Published: (2026)
by: Kang, Hao, et al.
Published: (2026)
CHRONOS: Compensating Hardware Related Overheads with Native Multi Timer Support for Real-Time Operating Systems
by: Heider, Kay, et al.
Published: (2025)
by: Heider, Kay, et al.
Published: (2025)
Idleness is Relative: Exploiting Tool-Call Idle Windows for Offloading in Agentic Systems with MORI
by: Xia, Tian, et al.
Published: (2026)
by: Xia, Tian, et al.
Published: (2026)
Hybrid Adaptive Tuning for Tiered Memory Systems
by: Wang, Xi, et al.
Published: (2026)
by: Wang, Xi, et al.
Published: (2026)
Guidelines for Building Indexes on Partially Cache-Coherent CXL Shared Memory
by: Wu, Fangnuo, et al.
Published: (2025)
by: Wu, Fangnuo, et al.
Published: (2025)
MNN-AECS: Energy Optimization for LLM Decoding on Mobile Devices via Adaptive Core Selection
by: Huang, Zhengxiang, et al.
Published: (2025)
by: Huang, Zhengxiang, et al.
Published: (2025)
Oneiros: KV Cache Optimization through Parameter Remapping for Multi-tenant LLM Serving
by: Li, Ruihao, et al.
Published: (2025)
by: Li, Ruihao, et al.
Published: (2025)
Efficient Function-as-a-Service for Large Language Models with TIDAL
by: Cui, Weihao, et al.
Published: (2025)
by: Cui, Weihao, et al.
Published: (2025)
FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation
by: Wu, Yinpeng, et al.
Published: (2026)
by: Wu, Yinpeng, et al.
Published: (2026)
C2CServe: Leveraging NVLink-C2C for Elastic Serverless LLM Serving on MIG
by: Luo, Shutian, et al.
Published: (2026)
by: Luo, Shutian, et al.
Published: (2026)
Global Scheduling of Weakly-Hard Real-Time Tasks using Job-Level Priority Classes
by: Moyano, V. Gabriel, et al.
Published: (2024)
by: Moyano, V. Gabriel, et al.
Published: (2024)
TempoNet: Slack-Quantized Transformer-Guided Reinforcement Scheduler for Adaptive Deadline-Centric Real-Time Dispatchs
by: Fu, Rong, et al.
Published: (2026)
by: Fu, Rong, et al.
Published: (2026)
Foreactor: Exploiting Storage I/O Parallelism with Explicit Speculation
by: Hu, Guanzhou, et al.
Published: (2024)
by: Hu, Guanzhou, et al.
Published: (2024)
Optimizing Tree-structure Indexes for CXL-based Heterogeneous Memory with SINLK
by: Zhao, Haoru, et al.
Published: (2025)
by: Zhao, Haoru, et al.
Published: (2025)
CARTOS: A Charging-Aware Real-Time Operating System for Intermittent Batteryless Devices
by: Karimi, Mohsen, et al.
Published: (2023)
by: Karimi, Mohsen, et al.
Published: (2023)
MARS: Efficient, Adaptive Co-Scheduling for Heterogeneous Agentic Systems
by: Wang, Yifei, et al.
Published: (2026)
by: Wang, Yifei, et al.
Published: (2026)
Toward Systems Foundations for Agentic Exploration
by: Xu, Jiakai, et al.
Published: (2025)
by: Xu, Jiakai, et al.
Published: (2025)
AdaptCache: KV Cache Native Storage Hierarchy for Low-Delay and High-Quality Language Model Serving
by: Feng, Shaoting, et al.
Published: (2025)
by: Feng, Shaoting, et al.
Published: (2025)
HyperGraph ROS: An Open-Source Robot Operating System for Hybrid Parallel Computing based on Computational HyperGraph
by: Zhang, Shufang, et al.
Published: (2025)
by: Zhang, Shufang, et al.
Published: (2025)
Trustworthy and Controllable Professional Knowledge Utilization in Large Language Models with TEE-GPU Execution
by: Cai, Yifeng, et al.
Published: (2025)
by: Cai, Yifeng, et al.
Published: (2025)
SSV: Sparse Speculative Verification for Efficient LLM Inference
by: Wang, Zhibin, et al.
Published: (2026)
by: Wang, Zhibin, et al.
Published: (2026)
Dissecting CXL Memory Performance at Scale: Analysis, Modeling, and Optimization
by: Liu, Jinshu, et al.
Published: (2024)
by: Liu, Jinshu, et al.
Published: (2024)
DynamicAdaptiveClimb: Adaptive Cache Replacement with Dynamic Resizing
by: Berend, Daniel, et al.
Published: (2025)
by: Berend, Daniel, et al.
Published: (2025)
vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention
by: Prabhu, Ramya, et al.
Published: (2024)
by: Prabhu, Ramya, et al.
Published: (2024)
ARMS: Adaptive and Robust Memory Tiering System
by: Yadalam, Sujay, et al.
Published: (2025)
by: Yadalam, Sujay, et al.
Published: (2025)
Adaptive Migration Decision for Multi-Tenant Memory Systems
by: Cho, Hyungjun, et al.
Published: (2025)
by: Cho, Hyungjun, et al.
Published: (2025)
Mitigating Timing-Based Attacks in Real-Time Cyber-Physical Systems
by: Sain, Arkaprava, et al.
Published: (2026)
by: Sain, Arkaprava, et al.
Published: (2026)
Formal Definitions and Performance Comparison of Consistency Models for Parallel File Systems
by: Wang, Chen, et al.
Published: (2024)
by: Wang, Chen, et al.
Published: (2024)
TierBPF: Page Migration Admission Control for Tiered Memory via eBPF
by: Wang, Xi, et al.
Published: (2026)
by: Wang, Xi, et al.
Published: (2026)
HeteroPod: XPU-Accelerated Infrastructure Offloading for Commodity Cloud-Native Applications
by: Yang, Bicheng, et al.
Published: (2025)
by: Yang, Bicheng, et al.
Published: (2025)
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
by: Feng, Shaoting, et al.
Published: (2025)
by: Feng, Shaoting, et al.
Published: (2025)
AERO: Adaptive and Efficient Runtime-Aware OTA Updates for Energy-Harvesting IoT
by: Wei, Wei, et al.
Published: (2026)
by: Wei, Wei, et al.
Published: (2026)
Optimizing Logical Execution Time Model for Both Determinism and Low Latency
by: Wang, Sen, et al.
Published: (2023)
by: Wang, Sen, et al.
Published: (2023)
Principled Performance Tunability in Operating System Kernels
by: Chen, Zhongjie, et al.
Published: (2025)
by: Chen, Zhongjie, et al.
Published: (2025)
Similar Items
-
Exploiting Dependency and Parallelism: Real-Time Scheduling and Analysis for GPU Tasks
by: Zhang, Yuanhai, et al.
Published: (2026) -
Towards High-Goodput LLM Serving with Prefill-decode Multiplexing
by: Chen, Yukang, et al.
Published: (2025) -
Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving
by: Qiu, Shi, et al.
Published: (2026) -
FRAP: A Flexible Resource Accessing Protocol for Multiprocessor Real-Time Systems
by: Zhao, Shuai, et al.
Published: (2024) -
PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
by: Song, Yixin, et al.
Published: (2023)