whisply: Cross-Platform Python App for Batch Transcription, Translation, Speaker Annotation and Subtitle Generation of Video and Audio Content

Fuente: Zenodo
Salvato in:
Dettagli Bibliografici
Autore principale: Schmidt, Thomas
Natura: Recurso digital
Pubblicazione: Zenodo 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866902093385695232
author Schmidt, Thomas
author_facet Schmidt, Thomas
contents Automated Transcription, Translation, and Speaker Diarization for Research whisply is a fast Python tool designed to transcribe, translate, and annotate audio and video files using OpenAI's open-weight Whisper models. It optimizes batch processing by integrating specialized implementations such as faster-whisper, mlx-whisper, and whisperX. Key Features: Hardware-Optimized Performance: Automatically selects the fastest Whisper implementation based on your hardware (NVIDIA GPU via CUDA, Apple Silicon M1-M5 via MLX, or CPU). Speaker Diarization: Provides word-level speaker annotations by integrating pyannote.audio, making it ideal for transcribing qualitative interviews and group discussions. Batch Processing: Efficiently handles single files, folders, URLs, or batch lists, with support for configuration files to streamline research workflows. Versatile Export Formats: Generates structured data (JSON, RTTM), plain or annotated TXT, HTML (compatible with the noScribe editor), and various subtitle formats (SRT, VTT, WebVTT). Data Privacy: Runs entirely locally on your own hardware, ensuring the confidentiality and security of sensitive research data. CLI and App: whisply is available as both a command-line interface (CLI) and a browser-based app (gradio), supporting Windows, Linux, and macOS. Installation Instructions: Refer the ReadMe of the GitHub repository for detailed installation instructions: https://github.com/tsmdt/whisply
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_20329799
institution Zenodo
language
publishDate 2026
publisher Zenodo
record_format zenodo
spellingShingle whisply: Cross-Platform Python App for Batch Transcription, Translation, Speaker Annotation and Subtitle Generation of Video and Audio Content
Schmidt, Thomas
whisper
asr
automatic-speech-recognition
speech-recognition
speech-to-text
audio-transcription
Automated Transcription, Translation, and Speaker Diarization for Research whisply is a fast Python tool designed to transcribe, translate, and annotate audio and video files using OpenAI's open-weight Whisper models. It optimizes batch processing by integrating specialized implementations such as faster-whisper, mlx-whisper, and whisperX. Key Features: Hardware-Optimized Performance: Automatically selects the fastest Whisper implementation based on your hardware (NVIDIA GPU via CUDA, Apple Silicon M1-M5 via MLX, or CPU). Speaker Diarization: Provides word-level speaker annotations by integrating pyannote.audio, making it ideal for transcribing qualitative interviews and group discussions. Batch Processing: Efficiently handles single files, folders, URLs, or batch lists, with support for configuration files to streamline research workflows. Versatile Export Formats: Generates structured data (JSON, RTTM), plain or annotated TXT, HTML (compatible with the noScribe editor), and various subtitle formats (SRT, VTT, WebVTT). Data Privacy: Runs entirely locally on your own hardware, ensuring the confidentiality and security of sensitive research data. CLI and App: whisply is available as both a command-line interface (CLI) and a browser-based app (gradio), supporting Windows, Linux, and macOS. Installation Instructions: Refer the ReadMe of the GitHub repository for detailed installation instructions: https://github.com/tsmdt/whisply
title whisply: Cross-Platform Python App for Batch Transcription, Translation, Speaker Annotation and Subtitle Generation of Video and Audio Content
topic whisper
asr
automatic-speech-recognition
speech-recognition
speech-to-text
audio-transcription
url https://doi.org/10.5281/zenodo.20329799