Gespeichert in:
| Hauptverfasser: | Aggazzotti, Cristina, Wiesner, Matthew, Smith, Elizabeth Allyn, Andrews, Nicholas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2507.08660 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Can Authorship Attribution Models Distinguish Speakers in Speech Transcripts?
von: Aggazzotti, Cristina, et al.
Veröffentlicht: (2023)
von: Aggazzotti, Cristina, et al.
Veröffentlicht: (2023)
A stylometric analysis of speaker attribution from speech transcripts
von: Aggazzotti, Cristina, et al.
Veröffentlicht: (2025)
von: Aggazzotti, Cristina, et al.
Veröffentlicht: (2025)
Content Anonymization for Privacy in Long-form Audio
von: Aggazzotti, Cristina, et al.
Veröffentlicht: (2025)
von: Aggazzotti, Cristina, et al.
Veröffentlicht: (2025)
Robust Long-Form Bangla Speech Processing: Automatic Speech Recognition and Speaker Diarization
von: Chowdhury, MD. Sagor, et al.
Veröffentlicht: (2026)
von: Chowdhury, MD. Sagor, et al.
Veröffentlicht: (2026)
AttributionBench: How Hard is Automatic Attribution Evaluation?
von: Li, Yifei, et al.
Veröffentlicht: (2024)
von: Li, Yifei, et al.
Veröffentlicht: (2024)
Modeling Overlapped Speech with Shuffles
von: Wiesner, Matthew, et al.
Veröffentlicht: (2026)
von: Wiesner, Matthew, et al.
Veröffentlicht: (2026)
Selective Augmentation: Improving Universal Automatic Phonetic Transcription via G2P Bootstrapping
von: Bystrich, Tobias, et al.
Veröffentlicht: (2026)
von: Bystrich, Tobias, et al.
Veröffentlicht: (2026)
Speech Rhythm-Based Speaker Embeddings Extraction from Phonemes and Phoneme Duration for Multi-Speaker Speech Synthesis
von: Fujita, Kenichi, et al.
Veröffentlicht: (2024)
von: Fujita, Kenichi, et al.
Veröffentlicht: (2024)
Speaker Style-Aware Phoneme Anchoring for Improved Cross-Lingual Speech Emotion Recognition
von: Upadhyay, Shreya G., et al.
Veröffentlicht: (2025)
von: Upadhyay, Shreya G., et al.
Veröffentlicht: (2025)
Continuously Learning New Words in Automatic Speech Recognition
von: Huber, Christian, et al.
Veröffentlicht: (2024)
von: Huber, Christian, et al.
Veröffentlicht: (2024)
Context Biasing for Pronunciation-Orthography Mismatch in Automatic Speech Recognition
von: Huber, Christian, et al.
Veröffentlicht: (2025)
von: Huber, Christian, et al.
Veröffentlicht: (2025)
FalAR: A Large-scale Speaker-Annotated European Portuguese Speech Corpus of Parliamentary Sessions
von: Teixeira, Francisco, et al.
Veröffentlicht: (2026)
von: Teixeira, Francisco, et al.
Veröffentlicht: (2026)
Early Attentive Sparsification Accelerates Neural Speech Transcription
von: Xu, Zifei, et al.
Veröffentlicht: (2025)
von: Xu, Zifei, et al.
Veröffentlicht: (2025)
Mitigating Paraphrase Attacks on Machine-Text Detectors via Paraphrase Inversion
von: Soto, Rafael Rivera, et al.
Veröffentlicht: (2024)
von: Soto, Rafael Rivera, et al.
Veröffentlicht: (2024)
Huntington Disease Automatic Speech Recognition with Biomarker Supervision
von: Wang, Charles L., et al.
Veröffentlicht: (2026)
von: Wang, Charles L., et al.
Veröffentlicht: (2026)
Moonshine: Speech Recognition for Live Transcription and Voice Commands
von: Jeffries, Nat, et al.
Veröffentlicht: (2024)
von: Jeffries, Nat, et al.
Veröffentlicht: (2024)
Analyzing and Improving Speaker Similarity Assessment for Speech Synthesis
von: Carbonneau, Marc-André, et al.
Veröffentlicht: (2025)
von: Carbonneau, Marc-André, et al.
Veröffentlicht: (2025)
Supplementary Resources and Analysis for Automatic Speech Recognition Systems Trained on the Loquacious Dataset
von: Rossenbach, Nick, et al.
Veröffentlicht: (2025)
von: Rossenbach, Nick, et al.
Veröffentlicht: (2025)
Joint Unsupervised and Supervised Training for Automatic Speech Recognition via Bilevel Optimization
von: Saif, A F M, et al.
Veröffentlicht: (2024)
von: Saif, A F M, et al.
Veröffentlicht: (2024)
A Survey on Automatic Online Hate Speech Detection in Low-Resource Languages
von: Das, Susmita, et al.
Veröffentlicht: (2024)
von: Das, Susmita, et al.
Veröffentlicht: (2024)
Learning Extrapolative Sequence Transformations from Markov Chains
von: Hager, Sophia, et al.
Veröffentlicht: (2025)
von: Hager, Sophia, et al.
Veröffentlicht: (2025)
Uncertainty Distillation: Teaching Language Models to Express Semantic Confidence
von: Hager, Sophia, et al.
Veröffentlicht: (2025)
von: Hager, Sophia, et al.
Veröffentlicht: (2025)
Unseen Speaker and Language Adaptation for Lightweight Text-To-Speech with Adapters
von: Falai, Alessio, et al.
Veröffentlicht: (2025)
von: Falai, Alessio, et al.
Veröffentlicht: (2025)
Exploring Speech Foundation Models for Speaker Diarization in Child-Adult Dyadic Interactions
von: Xu, Anfeng, et al.
Veröffentlicht: (2024)
von: Xu, Anfeng, et al.
Veröffentlicht: (2024)
Robust Unsupervised Adaptation of a Speech Recogniser Using Entropy Minimisation and Speaker Codes
von: van Dalen, Rogier C., et al.
Veröffentlicht: (2025)
von: van Dalen, Rogier C., et al.
Veröffentlicht: (2025)
Codec-ASR: Training Performant Automatic Speech Recognition Systems with Discrete Speech Representations
von: Dhawan, Kunal, et al.
Veröffentlicht: (2024)
von: Dhawan, Kunal, et al.
Veröffentlicht: (2024)
Language Models Optimized to Fool Detectors Still Have a Distinct Style (And How to Change It)
von: Soto, Rafael Rivera, et al.
Veröffentlicht: (2025)
von: Soto, Rafael Rivera, et al.
Veröffentlicht: (2025)
DC-Spin: A Speaker-invariant Speech Tokenizer for Spoken Language Models
von: Chang, Heng-Jui, et al.
Veröffentlicht: (2024)
von: Chang, Heng-Jui, et al.
Veröffentlicht: (2024)
A Semi-Automatic Approach to Create Large Gender- and Age-Balanced Speaker Corpora: Usefulness of Speaker Diarization & Identification
von: Uro, Rémi, et al.
Veröffentlicht: (2024)
von: Uro, Rémi, et al.
Veröffentlicht: (2024)
TaCo: Targeted Concept Erasure Prevents Non-Linear Classifiers From Detecting Protected Attributes
von: Jourdan, Fanny, et al.
Veröffentlicht: (2023)
von: Jourdan, Fanny, et al.
Veröffentlicht: (2023)
Continual Learning for Monolingual End-to-End Automatic Speech Recognition
von: Eeckt, Steven Vander, et al.
Veröffentlicht: (2021)
von: Eeckt, Steven Vander, et al.
Veröffentlicht: (2021)
Combining TF-GridNet and Mixture Encoder for Continuous Speech Separation for Meeting Transcription
von: Vieting, Peter, et al.
Veröffentlicht: (2023)
von: Vieting, Peter, et al.
Veröffentlicht: (2023)
Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems
von: Park, Taejin, et al.
Veröffentlicht: (2024)
von: Park, Taejin, et al.
Veröffentlicht: (2024)
On the Problem of Text-To-Speech Model Selection for Synthetic Data Generation in Automatic Speech Recognition
von: Rossenbach, Nick, et al.
Veröffentlicht: (2024)
von: Rossenbach, Nick, et al.
Veröffentlicht: (2024)
The ParlaSpeech Collection of Automatically Generated Speech and Text Datasets from Parliamentary Proceedings
von: Ljubešić, Nikola, et al.
Veröffentlicht: (2024)
von: Ljubešić, Nikola, et al.
Veröffentlicht: (2024)
Regularizing Learnable Feature Extraction for Automatic Speech Recognition
von: Vieting, Peter, et al.
Veröffentlicht: (2025)
von: Vieting, Peter, et al.
Veröffentlicht: (2025)
Analyzing and Fine-Tuning Whisper Models for Multilingual Pilot Speech Transcription in the Cockpit
von: Nareddy, Kartheek Kumar Reddy, et al.
Veröffentlicht: (2025)
von: Nareddy, Kartheek Kumar Reddy, et al.
Veröffentlicht: (2025)
Transducers with Pronunciation-aware Embeddings for Automatic Speech Recognition
von: Xu, Hainan, et al.
Veröffentlicht: (2024)
von: Xu, Hainan, et al.
Veröffentlicht: (2024)
BeaverTalk: Oregon State University's IWSLT 2025 Simultaneous Speech Translation System
von: Raffel, Matthew, et al.
Veröffentlicht: (2025)
von: Raffel, Matthew, et al.
Veröffentlicht: (2025)
Semantically Corrected Amharic Automatic Speech Recognition
von: Adnew, Samuael, et al.
Veröffentlicht: (2024)
von: Adnew, Samuael, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Can Authorship Attribution Models Distinguish Speakers in Speech Transcripts?
von: Aggazzotti, Cristina, et al.
Veröffentlicht: (2023) -
A stylometric analysis of speaker attribution from speech transcripts
von: Aggazzotti, Cristina, et al.
Veröffentlicht: (2025) -
Content Anonymization for Privacy in Long-form Audio
von: Aggazzotti, Cristina, et al.
Veröffentlicht: (2025) -
Robust Long-Form Bangla Speech Processing: Automatic Speech Recognition and Speaker Diarization
von: Chowdhury, MD. Sagor, et al.
Veröffentlicht: (2026) -
AttributionBench: How Hard is Automatic Attribution Evaluation?
von: Li, Yifei, et al.
Veröffentlicht: (2024)