Saved in:
Bibliographic Details
Main Authors: Makwana, Darshan, Jogi, Yash, Kotta, Harsh, Kubba, Aayush
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2603.11273
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908880359915520
author Makwana, Darshan
Jogi, Yash
Kotta, Harsh
Kubba, Aayush
author_facet Makwana, Darshan
Jogi, Yash
Kotta, Harsh
Kubba, Aayush
contents Scheduling policies in large-scale Automatic Speech Recognition (ASR) serving pipelines play a key role in determining end-to-end (E2E) latency. Yet, widely used serving engines rely on first-come-first-served (FCFS) scheduling, which ignores variability in request duration and leads to head-of-line blocking under workload drift. We show that audio duration is an accurate proxy for job processing time in ASR models such as Whisper, and use this insight to enable duration-aware scheduling. We integrate two classical algorithms, Shortest Job First (SJF) and Highest Response Ratio Next (HRRN), into vLLM and evaluate them under realistic and drifted workloads. On LibriSpeech test-clean, compared to baseline, SJF reduces median E2E latency by up to $73\%$ at high load, but increases $90$th-percentile tail latency by up to $97\%$ due to starvation of long requests. HRRN addresses this trade-off: it reduces median E2E latency by up to $28\%$ while bounding tail-latency degradation to at most $24\%$. These gains persist under workload drift, with no throughput penalty and $<0.1$\,ms scheduling overhead per request.
format Preprint
id arxiv_https___arxiv_org_abs_2603_11273
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Duration Aware Scheduling for ASR Serving Under Workload Drift
Makwana, Darshan
Jogi, Yash
Kotta, Harsh
Kubba, Aayush
Machine Learning
Scheduling policies in large-scale Automatic Speech Recognition (ASR) serving pipelines play a key role in determining end-to-end (E2E) latency. Yet, widely used serving engines rely on first-come-first-served (FCFS) scheduling, which ignores variability in request duration and leads to head-of-line blocking under workload drift. We show that audio duration is an accurate proxy for job processing time in ASR models such as Whisper, and use this insight to enable duration-aware scheduling. We integrate two classical algorithms, Shortest Job First (SJF) and Highest Response Ratio Next (HRRN), into vLLM and evaluate them under realistic and drifted workloads. On LibriSpeech test-clean, compared to baseline, SJF reduces median E2E latency by up to $73\%$ at high load, but increases $90$th-percentile tail latency by up to $97\%$ due to starvation of long requests. HRRN addresses this trade-off: it reduces median E2E latency by up to $28\%$ while bounding tail-latency degradation to at most $24\%$. These gains persist under workload drift, with no throughput penalty and $<0.1$\,ms scheduling overhead per request.
title Duration Aware Scheduling for ASR Serving Under Workload Drift
topic Machine Learning
url https://arxiv.org/abs/2603.11273