Whispy: Adapting STT Whisper Models to Real-Time Environments
Fuente:
arXiv
Saved in:
| Main Authors: | Bevilacqua, Antonio, Saviano, Paolo, Amirante, Alessandro, Romano, Simon Pietro |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Whisper-SV: Adapting Whisper for Low-data-resource Speaker Verification
by: Zhang, Li, et al.
Published: (2024)
by: Zhang, Li, et al.
Published: (2024)
Fine-Tuning Whisper for Inclusive Prosodic Stress Analysis
by: Sohn, Samuel S., et al.
Published: (2025)
by: Sohn, Samuel S., et al.
Published: (2025)
Domain Adapting Deep Reinforcement Learning for Real-world Speech Emotion Recognition
by: Rajapakshe, Thejan, et al.
Published: (2022)
by: Rajapakshe, Thejan, et al.
Published: (2022)
WhisperRT -- Turning Whisper into a Causal Streaming Model
by: Krichli, Tomer, et al.
Published: (2025)
by: Krichli, Tomer, et al.
Published: (2025)
Leveraging Whisper Embeddings for Audio-based Lyrics Matching
by: Mancini, Eleonora, et al.
Published: (2025)
by: Mancini, Eleonora, et al.
Published: (2025)
Non-Intrusive Speech Intelligibility Prediction for Hearing Aids using Whisper and Metadata
by: Zezario, Ryandhimas E., et al.
Published: (2023)
by: Zezario, Ryandhimas E., et al.
Published: (2023)
Behind the Scenes: Mechanistic Interpretability of LoRA-adapted Whisper for Speech Emotion Recognition
by: Ma, Yujian, et al.
Published: (2025)
by: Ma, Yujian, et al.
Published: (2025)
WhiSQA: Non-Intrusive Speech Quality Prediction Using Whisper Encoder Features
by: Close, George, et al.
Published: (2025)
by: Close, George, et al.
Published: (2025)
A Comparative Evaluation of Deep Learning Models for Speech Enhancement in Real-World Noisy Environments
by: Khondkar, Md Jahangir Alam, et al.
Published: (2025)
by: Khondkar, Md Jahangir Alam, et al.
Published: (2025)
Optimizing Multi-Stuttered Speech Classification: Leveraging Whisper's Encoder for Efficient Parameter Reduction in Automated Assessment
by: Ameer, Huma, et al.
Published: (2024)
by: Ameer, Huma, et al.
Published: (2024)
Towards Sub-millisecond Latency Real-Time Speech Enhancement Models on Hearables
by: Dementyev, Artem, et al.
Published: (2024)
by: Dementyev, Artem, et al.
Published: (2024)
Towards Efficient and Real-Time Piano Transcription Using Neural Autoregressive Models
by: Kwon, Taegyun, et al.
Published: (2024)
by: Kwon, Taegyun, et al.
Published: (2024)
Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding
by: Zhou, Haoran, et al.
Published: (2025)
by: Zhou, Haoran, et al.
Published: (2025)
Whilter: A Whisper-based Data Filter for "In-the-Wild" Speech Corpora Using Utterance-level Multi-Task Classification
by: Ravenscroft, William, et al.
Published: (2025)
by: Ravenscroft, William, et al.
Published: (2025)
WhisperD: Dementia Speech Recognition and Filler Word Detection with Whisper
by: Akinrintoyo, Emmanuel, et al.
Published: (2025)
by: Akinrintoyo, Emmanuel, et al.
Published: (2025)
Real-Time Streaming Mel Vocoding with Generative Flow Matching
by: Welker, Simon, et al.
Published: (2025)
by: Welker, Simon, et al.
Published: (2025)
Adapting WavLM for Speech Emotion Recognition
by: Diatlova, Daria, et al.
Published: (2024)
by: Diatlova, Daria, et al.
Published: (2024)
Temporal Convolution-based Hybrid Model Approach with Representation Learning for Real-Time Acoustic Anomaly Detection
by: Dissanayaka, Sahan, et al.
Published: (2024)
by: Dissanayaka, Sahan, et al.
Published: (2024)
Audio-to-Score Conversion Model Based on Whisper methodology
by: Zhang, Hongyao, et al.
Published: (2024)
by: Zhang, Hongyao, et al.
Published: (2024)
Improving Real-Time Music Accompaniment Separation with MMDenseNet
by: Wang, Chun-Hsiang, et al.
Published: (2024)
by: Wang, Chun-Hsiang, et al.
Published: (2024)
StreamVC: Real-Time Low-Latency Voice Conversion
by: Yang, Yang, et al.
Published: (2024)
by: Yang, Yang, et al.
Published: (2024)
TF-MLPNet: Tiny Real-Time Neural Speech Separation
by: Itani, Malek, et al.
Published: (2025)
by: Itani, Malek, et al.
Published: (2025)
UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching
by: Glazer, Neta, et al.
Published: (2025)
by: Glazer, Neta, et al.
Published: (2025)
Adapting Automatic Speech Recognition for Accented Air Traffic Control Communications
by: Wee, Marcus Yu Zhe, et al.
Published: (2025)
by: Wee, Marcus Yu Zhe, et al.
Published: (2025)
Exploring System Adaptations For Minimum Latency Real-Time Piano Transcription
by: Hu, Patricia, et al.
Published: (2025)
by: Hu, Patricia, et al.
Published: (2025)
ERSAM: Neural Architecture Search For Energy-Efficient and Real-Time Social Ambiance Measurement
by: Li, Chaojian, et al.
Published: (2023)
by: Li, Chaojian, et al.
Published: (2023)
Highly Efficient Real-Time Streaming and Fully On-Device Speaker Diarization with Multi-Stage Clustering
by: Wang, Quan, et al.
Published: (2022)
by: Wang, Quan, et al.
Published: (2022)
Edge Intelligence for Wildlife Conservation: Real-Time Hornbill Call Classification Using TinyML
by: Hing, Kong Ka, et al.
Published: (2025)
by: Hing, Kong Ka, et al.
Published: (2025)
A Real-Time Lyrics Alignment System Using Chroma And Phonetic Features For Classical Vocal Performance
by: Park, Jiyun, et al.
Published: (2024)
by: Park, Jiyun, et al.
Published: (2024)
SQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASR
by: Guo, Pengcheng, et al.
Published: (2024)
by: Guo, Pengcheng, et al.
Published: (2024)
Whisper Smarter, not Harder: Adversarial Attack on Partial Suppression
by: Wong, Zheng Jie, et al.
Published: (2025)
by: Wong, Zheng Jie, et al.
Published: (2025)
Whisper-PMFA: Partial Multi-Scale Feature Aggregation for Speaker Verification using Whisper Models
by: Zhao, Yiyang, et al.
Published: (2024)
by: Zhao, Yiyang, et al.
Published: (2024)
Data-Driven Room Acoustic Modeling Via Differentiable Feedback Delay Networks With Learnable Delay Lines
by: Mezza, Alessandro Ilic, et al.
Published: (2024)
by: Mezza, Alessandro Ilic, et al.
Published: (2024)
Distribution Preserving Source Separation With Time Frequency Predictive Models
by: T., Pedro J. Villasana, et al.
Published: (2023)
by: T., Pedro J. Villasana, et al.
Published: (2023)
Towards Lightweight Adaptation of Speech Enhancement Models in Real-World Environments
by: Cheng, Longbiao, et al.
Published: (2026)
by: Cheng, Longbiao, et al.
Published: (2026)
LLark: A Multimodal Instruction-Following Language Model for Music
by: Gardner, Josh, et al.
Published: (2023)
by: Gardner, Josh, et al.
Published: (2023)
Multi-Microphone and Multi-Modal Emotion Recognition in Reverberant Environment
by: Cohen, Ohad, et al.
Published: (2024)
by: Cohen, Ohad, et al.
Published: (2024)
Sound event localization and classification using WASN in Outdoor Environment
by: Zhang, Dongzhe, et al.
Published: (2024)
by: Zhang, Dongzhe, et al.
Published: (2024)
Diffusion Models for Audio Restoration
by: Lemercier, Jean-Marie, et al.
Published: (2024)
by: Lemercier, Jean-Marie, et al.
Published: (2024)
StoRM: A Diffusion-based Stochastic Regeneration Model for Speech Enhancement and Dereverberation
by: Lemercier, Jean-Marie, et al.
Published: (2022)
by: Lemercier, Jean-Marie, et al.
Published: (2022)
Similar Items
-
Whisper-SV: Adapting Whisper for Low-data-resource Speaker Verification
by: Zhang, Li, et al.
Published: (2024) -
Fine-Tuning Whisper for Inclusive Prosodic Stress Analysis
by: Sohn, Samuel S., et al.
Published: (2025) -
Domain Adapting Deep Reinforcement Learning for Real-world Speech Emotion Recognition
by: Rajapakshe, Thejan, et al.
Published: (2022) -
WhisperRT -- Turning Whisper into a Causal Streaming Model
by: Krichli, Tomer, et al.
Published: (2025) -
Leveraging Whisper Embeddings for Audio-based Lyrics Matching
by: Mancini, Eleonora, et al.
Published: (2025)