CrisperWhisper: Accurate Timestamps on Verbatim Speech Transcriptions
Fuente:
arXiv
Saved in:
| Main Authors: | Wagner, Laurin, Thallinger, Bernhard, Zusag, Mario |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Prompting Whisper for Improved Verbatim Transcription and End-to-end Miscue Detection
by: Smith, Griffin Dietz, et al.
Published: (2025)
by: Smith, Griffin Dietz, et al.
Published: (2025)
Analyzing and Fine-Tuning Whisper Models for Multilingual Pilot Speech Transcription in the Cockpit
by: Nareddy, Kartheek Kumar Reddy, et al.
Published: (2025)
by: Nareddy, Kartheek Kumar Reddy, et al.
Published: (2025)
Demystifying Verbatim Memorization in Large Language Models
by: Huang, Jing, et al.
Published: (2024)
by: Huang, Jing, et al.
Published: (2024)
Large Language Models Can Verbatim Reproduce Long Malicious Sequences
by: Lin, Sharon, et al.
Published: (2025)
by: Lin, Sharon, et al.
Published: (2025)
WhisperD: Dementia Speech Recognition and Filler Word Detection with Whisper
by: Akinrintoyo, Emmanuel, et al.
Published: (2025)
by: Akinrintoyo, Emmanuel, et al.
Published: (2025)
Reading the Finetuning Prior: Verbatim Content Recovery via Contrastive Decoding Diffing
by: Brzozowski, Michał, et al.
Published: (2026)
by: Brzozowski, Michał, et al.
Published: (2026)
WhisperNER: Unified Open Named Entity and Speech Recognition
by: Ayache, Gil, et al.
Published: (2024)
by: Ayache, Gil, et al.
Published: (2024)
A Semi-Supervised Framework for Speech Confidence Detection using Whisper
by: Wynn, Adam, et al.
Published: (2026)
by: Wynn, Adam, et al.
Published: (2026)
MaskCycleGAN-based Whisper to Normal Speech Conversion
by: Gupta, K. Rohith, et al.
Published: (2024)
by: Gupta, K. Rohith, et al.
Published: (2024)
In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions
by: Fan, Xulin, et al.
Published: (2026)
by: Fan, Xulin, et al.
Published: (2026)
Enhancing Aviation Communication Transcription: Fine-Tuning Distil-Whisper with LoRA
by: Mirzaei, Shokoufeh, et al.
Published: (2025)
by: Mirzaei, Shokoufeh, et al.
Published: (2025)
WhisperAlign: Word-Boundary-Aware ASR and WhisperX-Anchored Pyannote Diarization for Long-Form Bengali Speech
by: Chowdhury, Aurchi, et al.
Published: (2026)
by: Chowdhury, Aurchi, et al.
Published: (2026)
APT: Affine Prototype-Timestamp For Time Series Forecasting Under Distribution Shift
by: Li, Yujie, et al.
Published: (2025)
by: Li, Yujie, et al.
Published: (2025)
The Impact of Automatic Speech Transcription on Speaker Attribution
by: Aggazzotti, Cristina, et al.
Published: (2025)
by: Aggazzotti, Cristina, et al.
Published: (2025)
How Will My Business Process Unfold? Predicting Case Suffixes With Start and End Timestamps
by: Ali, Muhammad Awais, et al.
Published: (2025)
by: Ali, Muhammad Awais, et al.
Published: (2025)
Language Models May Verbatim Complete Text They Were Not Explicitly Trained On
by: Liu, Ken Ziyu, et al.
Published: (2025)
by: Liu, Ken Ziyu, et al.
Published: (2025)
Whispering in Amharic: Fine-tuning Whisper for Low-resource Language
by: Gete, Dawit Ketema, et al.
Published: (2025)
by: Gete, Dawit Ketema, et al.
Published: (2025)
Multi-Action Self-Improvement for Neural Combinatorial Optimization
by: Luttmann, Laurin, et al.
Published: (2025)
by: Luttmann, Laurin, et al.
Published: (2025)
IndexNet: Timestamp and Variable-Aware Modeling for Time Series Forecasting
by: Wu, Beiliang, et al.
Published: (2025)
by: Wu, Beiliang, et al.
Published: (2025)
ParaPO: Aligning Language Models to Reduce Verbatim Reproduction of Pre-training Data
by: Chen, Tong, et al.
Published: (2025)
by: Chen, Tong, et al.
Published: (2025)
Quartered Chirp Spectral Envelope for Whispered vs Normal Speech Classification
by: Joysingh, S. Johanan, et al.
Published: (2024)
by: Joysingh, S. Johanan, et al.
Published: (2024)
Rethinking the Power of Timestamps for Robust Time Series Forecasting: A Global-Local Fusion Perspective
by: Wang, Chengsen, et al.
Published: (2024)
by: Wang, Chengsen, et al.
Published: (2024)
Non-Intrusive Speech Intelligibility Prediction for Hearing Aids using Whisper and Metadata
by: Zezario, Ryandhimas E., et al.
Published: (2023)
by: Zezario, Ryandhimas E., et al.
Published: (2023)
Hallucination Level of Artificial Intelligence Whisperer: Case Speech Recognizing Pantterinousut Rap Song
by: Horppu, Ismo, et al.
Published: (2025)
by: Horppu, Ismo, et al.
Published: (2025)
Can Authorship Attribution Models Distinguish Speakers in Speech Transcripts?
by: Aggazzotti, Cristina, et al.
Published: (2023)
by: Aggazzotti, Cristina, et al.
Published: (2023)
Quartered Spectral Envelope and 1D-CNN-based Classification of Normally Phonated and Whispered Speech
by: Joysingh, S. Johanan, et al.
Published: (2024)
by: Joysingh, S. Johanan, et al.
Published: (2024)
Behind the Scenes: Mechanistic Interpretability of LoRA-adapted Whisper for Speech Emotion Recognition
by: Ma, Yujian, et al.
Published: (2025)
by: Ma, Yujian, et al.
Published: (2025)
WhiSQA: Non-Intrusive Speech Quality Prediction Using Whisper Encoder Features
by: Close, George, et al.
Published: (2025)
by: Close, George, et al.
Published: (2025)
Invitations to Publish in a Journal by Email: A Prospective Analysis
by: Christiane Thallinger, et al.
Published: (2026)
by: Christiane Thallinger, et al.
Published: (2026)
Learning to Solve the Min-Max Mixed-Shelves Picker-Routing Problem via Hierarchical and Parallel Decoding
by: Luttmann, Laurin, et al.
Published: (2025)
by: Luttmann, Laurin, et al.
Published: (2025)
Speech Emotion Recognition Leveraging OpenAI's Whisper Representations and Attentive Pooling Methods
by: Shendabadi, Ali, et al.
Published: (2026)
by: Shendabadi, Ali, et al.
Published: (2026)
LVSum: A Benchmark for Timestamp-Aware Long Video Summarization
by: Patel, Alkesh, et al.
Published: (2026)
by: Patel, Alkesh, et al.
Published: (2026)
Chunk-wise Attention Transducers for Fast and Accurate Streaming Speech-to-Text
by: Xu, Hainan, et al.
Published: (2026)
by: Xu, Hainan, et al.
Published: (2026)
SyncFed: Time-Aware Federated Learning through Explicit Timestamping and Synchronization
by: Gül, Baran Can, et al.
Published: (2025)
by: Gül, Baran Can, et al.
Published: (2025)
Edge-Based Speech Transcription and Synthesis for Kinyarwanda and Swahili Languages
by: Mbonimpa, Pacome Simon, et al.
Published: (2025)
by: Mbonimpa, Pacome Simon, et al.
Published: (2025)
Whispering Under the Eaves: Protecting User Privacy Against Commercial and LLM-powered Automatic Speech Recognition Systems
by: Jin, Weifei, et al.
Published: (2025)
by: Jin, Weifei, et al.
Published: (2025)
Adopting Whisper for Confidence Estimation
by: Aggarwal, Vaibhav, et al.
Published: (2025)
by: Aggarwal, Vaibhav, et al.
Published: (2025)
Optimizing Multi-Stuttered Speech Classification: Leveraging Whisper's Encoder for Efficient Parameter Reduction in Automated Assessment
by: Ameer, Huma, et al.
Published: (2024)
by: Ameer, Huma, et al.
Published: (2024)
Whispers in the Machine: Confidentiality in Agentic Systems
by: Evertz, Jonathan, et al.
Published: (2024)
by: Evertz, Jonathan, et al.
Published: (2024)
WhisperRT -- Turning Whisper into a Causal Streaming Model
by: Krichli, Tomer, et al.
Published: (2025)
by: Krichli, Tomer, et al.
Published: (2025)
Similar Items
-
Prompting Whisper for Improved Verbatim Transcription and End-to-end Miscue Detection
by: Smith, Griffin Dietz, et al.
Published: (2025) -
Analyzing and Fine-Tuning Whisper Models for Multilingual Pilot Speech Transcription in the Cockpit
by: Nareddy, Kartheek Kumar Reddy, et al.
Published: (2025) -
Demystifying Verbatim Memorization in Large Language Models
by: Huang, Jing, et al.
Published: (2024) -
Large Language Models Can Verbatim Reproduce Long Malicious Sequences
by: Lin, Sharon, et al.
Published: (2025) -
WhisperD: Dementia Speech Recognition and Filler Word Detection with Whisper
by: Akinrintoyo, Emmanuel, et al.
Published: (2025)