Saved in:
Bibliographic Details
Main Authors: Kundu, Arpita, Chakraborty, Joyita, Desarkar, Anindita, Sen, Aritra, Patil, Srushti Anil, Raman, Vishwanathan
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2510.24180
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914118012764160
author Kundu, Arpita
Chakraborty, Joyita
Desarkar, Anindita
Sen, Aritra
Patil, Srushti Anil
Raman, Vishwanathan
author_facet Kundu, Arpita
Chakraborty, Joyita
Desarkar, Anindita
Sen, Aritra
Patil, Srushti Anil
Raman, Vishwanathan
contents The surge of audiovisual content on streaming platforms and social media has heightened the demand for accurate and accessible subtitles. However, existing subtitle generation methods primarily speech-based transcription or OCR-based extraction suffer from several shortcomings, including poor synchronization, incorrect or harmful text, inconsistent formatting, inappropriate reading speeds, and the inability to adapt to dynamic audio-visual contexts. Current approaches often address isolated issues, leaving post-editing as a labor-intensive and time-consuming process. In this paper, we introduce V-SAT (Video Subtitle Annotation Tool), a unified framework that automatically detects and corrects a wide range of subtitle quality issues. By combining Large Language Models(LLMs), Vision-Language Models (VLMs), Image Processing, and Automatic Speech Recognition (ASR), V-SAT leverages contextual cues from both audio and video. Subtitle quality improved, with the SUBER score reduced from 9.6 to 3.54 after resolving all language mode issues and F1-scores of ~0.80 for image mode issues. Human-in-the-loop validation ensures high-quality results, providing the first comprehensive solution for robust subtitle annotation.
format Preprint
id arxiv_https___arxiv_org_abs_2510_24180
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle V-SAT: Video Subtitle Annotation Tool
Kundu, Arpita
Chakraborty, Joyita
Desarkar, Anindita
Sen, Aritra
Patil, Srushti Anil
Raman, Vishwanathan
Machine Learning
The surge of audiovisual content on streaming platforms and social media has heightened the demand for accurate and accessible subtitles. However, existing subtitle generation methods primarily speech-based transcription or OCR-based extraction suffer from several shortcomings, including poor synchronization, incorrect or harmful text, inconsistent formatting, inappropriate reading speeds, and the inability to adapt to dynamic audio-visual contexts. Current approaches often address isolated issues, leaving post-editing as a labor-intensive and time-consuming process. In this paper, we introduce V-SAT (Video Subtitle Annotation Tool), a unified framework that automatically detects and corrects a wide range of subtitle quality issues. By combining Large Language Models(LLMs), Vision-Language Models (VLMs), Image Processing, and Automatic Speech Recognition (ASR), V-SAT leverages contextual cues from both audio and video. Subtitle quality improved, with the SUBER score reduced from 9.6 to 3.54 after resolving all language mode issues and F1-scores of ~0.80 for image mode issues. Human-in-the-loop validation ensures high-quality results, providing the first comprehensive solution for robust subtitle annotation.
title V-SAT: Video Subtitle Annotation Tool
topic Machine Learning
url https://arxiv.org/abs/2510.24180