Saved in:
Bibliographic Details
Main Authors: Lucca, Alessandro, Pierri, Francesco
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2512.19161
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912782355529728
author Lucca, Alessandro
Pierri, Francesco
author_facet Lucca, Alessandro
Pierri, Francesco
contents Subtitles are essential for video accessibility and audience engagement. Modern Automatic Speech Recognition (ASR) systems, built upon Encoder-Decoder neural network architectures and trained on massive amounts of data, have progressively reduced transcription errors on standard benchmark datasets. However, their performance in real-world production environments, particularly for non-English content like long-form Italian videos, remains largely unexplored. This paper presents a case study on developing a professional subtitling system for an Italian media company. To inform our system design, we evaluated four state-of-the-art ASR models (Whisper Large v2, AssemblyAI Universal, Parakeet TDT v3 0.6b, and WhisperX) on a 50-hour dataset of Italian television programs. The study highlights their strengths and limitations, benchmarking their performance against the work of professional human subtitlers. The findings indicate that, while current models cannot meet the media industry's accuracy needs for full autonomy, they can serve as highly effective tools for enhancing human productivity. We conclude that a human-in-the-loop (HITL) approach is crucial and present the production-grade, cloud-based infrastructure we designed to support this workflow.
format Preprint
id arxiv_https___arxiv_org_abs_2512_19161
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From Speech to Subtitles: Evaluating ASR Models in Subtitling Italian Television Programs
Lucca, Alessandro
Pierri, Francesco
Computation and Language
Subtitles are essential for video accessibility and audience engagement. Modern Automatic Speech Recognition (ASR) systems, built upon Encoder-Decoder neural network architectures and trained on massive amounts of data, have progressively reduced transcription errors on standard benchmark datasets. However, their performance in real-world production environments, particularly for non-English content like long-form Italian videos, remains largely unexplored. This paper presents a case study on developing a professional subtitling system for an Italian media company. To inform our system design, we evaluated four state-of-the-art ASR models (Whisper Large v2, AssemblyAI Universal, Parakeet TDT v3 0.6b, and WhisperX) on a 50-hour dataset of Italian television programs. The study highlights their strengths and limitations, benchmarking their performance against the work of professional human subtitlers. The findings indicate that, while current models cannot meet the media industry's accuracy needs for full autonomy, they can serve as highly effective tools for enhancing human productivity. We conclude that a human-in-the-loop (HITL) approach is crucial and present the production-grade, cloud-based infrastructure we designed to support this workflow.
title From Speech to Subtitles: Evaluating ASR Models in Subtitling Italian Television Programs
topic Computation and Language
url https://arxiv.org/abs/2512.19161