whisply: Cross-Platform Python App for Batch Transcription, Translation, Speaker Annotation and Subtitle Generation of Video and Audio Content
Fuente:
Zenodo
Salvato in:
| Autore principale: | |
|---|---|
| Natura: | Recurso digital |
| Pubblicazione: |
Zenodo
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866902093385695232 |
|---|---|
| author | Schmidt, Thomas |
| author_facet | Schmidt, Thomas |
| contents | Automated Transcription, Translation, and Speaker Diarization for Research whisply is a fast Python tool designed to transcribe, translate, and annotate audio and video files using OpenAI's open-weight Whisper models. It optimizes batch processing by integrating specialized implementations such as faster-whisper, mlx-whisper, and whisperX. Key Features: Hardware-Optimized Performance: Automatically selects the fastest Whisper implementation based on your hardware (NVIDIA GPU via CUDA, Apple Silicon M1-M5 via MLX, or CPU). Speaker Diarization: Provides word-level speaker annotations by integrating pyannote.audio, making it ideal for transcribing qualitative interviews and group discussions. Batch Processing: Efficiently handles single files, folders, URLs, or batch lists, with support for configuration files to streamline research workflows. Versatile Export Formats: Generates structured data (JSON, RTTM), plain or annotated TXT, HTML (compatible with the noScribe editor), and various subtitle formats (SRT, VTT, WebVTT). Data Privacy: Runs entirely locally on your own hardware, ensuring the confidentiality and security of sensitive research data. CLI and App: whisply is available as both a command-line interface (CLI) and a browser-based app (gradio), supporting Windows, Linux, and macOS. Installation Instructions: Refer the ReadMe of the GitHub repository for detailed installation instructions: https://github.com/tsmdt/whisply |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_20329799 |
| institution | Zenodo |
| language | |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | whisply: Cross-Platform Python App for Batch Transcription, Translation, Speaker Annotation and Subtitle Generation of Video and Audio Content Schmidt, Thomas whisper asr automatic-speech-recognition speech-recognition speech-to-text audio-transcription Automated Transcription, Translation, and Speaker Diarization for Research whisply is a fast Python tool designed to transcribe, translate, and annotate audio and video files using OpenAI's open-weight Whisper models. It optimizes batch processing by integrating specialized implementations such as faster-whisper, mlx-whisper, and whisperX. Key Features: Hardware-Optimized Performance: Automatically selects the fastest Whisper implementation based on your hardware (NVIDIA GPU via CUDA, Apple Silicon M1-M5 via MLX, or CPU). Speaker Diarization: Provides word-level speaker annotations by integrating pyannote.audio, making it ideal for transcribing qualitative interviews and group discussions. Batch Processing: Efficiently handles single files, folders, URLs, or batch lists, with support for configuration files to streamline research workflows. Versatile Export Formats: Generates structured data (JSON, RTTM), plain or annotated TXT, HTML (compatible with the noScribe editor), and various subtitle formats (SRT, VTT, WebVTT). Data Privacy: Runs entirely locally on your own hardware, ensuring the confidentiality and security of sensitive research data. CLI and App: whisply is available as both a command-line interface (CLI) and a browser-based app (gradio), supporting Windows, Linux, and macOS. Installation Instructions: Refer the ReadMe of the GitHub repository for detailed installation instructions: https://github.com/tsmdt/whisply |
| title | whisply: Cross-Platform Python App for Batch Transcription, Translation, Speaker Annotation and Subtitle Generation of Video and Audio Content |
| topic | whisper asr automatic-speech-recognition speech-recognition speech-to-text audio-transcription |
| url | https://doi.org/10.5281/zenodo.20329799 |