Conversational Rubert for Detecting Competitive Interruptions in ASR-Transcribed Dialogues
Fuente:
arXiv
Saved in:
| Main Authors: | Galimzianov, Dmitrii, Vyshegorodtsev, Viacheslav |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Self-Train Before You Transcribe
by: Flynn, Robert, et al.
Published: (2024)
by: Flynn, Robert, et al.
Published: (2024)
ISPA: Inter-Species Phonetic Alphabet for Transcribing Animal Sounds
by: Hagiwara, Masato, et al.
Published: (2024)
by: Hagiwara, Masato, et al.
Published: (2024)
GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
by: Chen, Guoguo, et al.
Published: (2021)
by: Chen, Guoguo, et al.
Published: (2021)
Anatomy of Industrial Scale Multilingual ASR
by: Ramirez, Francis McCann, et al.
Published: (2024)
by: Ramirez, Francis McCann, et al.
Published: (2024)
Beyond Transcription: Mechanistic Interpretability in ASR
by: Glazer, Neta, et al.
Published: (2025)
by: Glazer, Neta, et al.
Published: (2025)
SALSA: Speedy ASR-LLM Synchronous Aggregation
by: Mittal, Ashish, et al.
Published: (2024)
by: Mittal, Ashish, et al.
Published: (2024)
Revisiting ASR Error Correction with Specialized Models
by: Gu, Zijin, et al.
Published: (2024)
by: Gu, Zijin, et al.
Published: (2024)
Federated Learning of Large ASR Models in the Real World
by: Xiao, Yonghui, et al.
Published: (2024)
by: Xiao, Yonghui, et al.
Published: (2024)
Unsupervised ASR via Cross-Lingual Pseudo-Labeling
by: Likhomanenko, Tatiana, et al.
Published: (2023)
by: Likhomanenko, Tatiana, et al.
Published: (2023)
Efficient Adapter Finetuning for Tail Languages in Streaming Multilingual ASR
by: Bai, Junwen, et al.
Published: (2024)
by: Bai, Junwen, et al.
Published: (2024)
Transformer-based Model for ASR N-Best Rescoring and Rewriting
by: Kang, Iwen E., et al.
Published: (2024)
by: Kang, Iwen E., et al.
Published: (2024)
Contextualization of ASR with LLM using phonetic retrieval-based augmentation
by: Lei, Zhihong, et al.
Published: (2024)
by: Lei, Zhihong, et al.
Published: (2024)
Cross-utterance ASR Rescoring with Graph-based Label Propagation
by: Tankasala, Srinath, et al.
Published: (2023)
by: Tankasala, Srinath, et al.
Published: (2023)
Unified Learnable 2D Convolutional Feature Extraction for ASR
by: Vieting, Peter, et al.
Published: (2025)
by: Vieting, Peter, et al.
Published: (2025)
Conformer-1: Robust ASR via Large-Scale Semisupervised Bootstrapping
by: Zhang, Kevin, et al.
Published: (2024)
by: Zhang, Kevin, et al.
Published: (2024)
Text-only adaptation in LLM-based ASR through text denoising
by: Carofilis, Andrés, et al.
Published: (2026)
by: Carofilis, Andrés, et al.
Published: (2026)
Dysarthria Normalization via Local Lie Group Transformations for Robust ASR
by: Osipov, Mikhail
Published: (2025)
by: Osipov, Mikhail
Published: (2025)
Learn and Don't Forget: Adding a New Language to ASR Foundation Models
by: Qian, Mengjie, et al.
Published: (2024)
by: Qian, Mengjie, et al.
Published: (2024)
Enhanced ASR Robustness to Packet Loss with a Front-End Adaptation Network
by: Dissen, Yehoshua, et al.
Published: (2024)
by: Dissen, Yehoshua, et al.
Published: (2024)
OLMoASR: Open Models and Data for Training Robust Speech Recognition Models
by: Ngo, Huong, et al.
Published: (2025)
by: Ngo, Huong, et al.
Published: (2025)
TOGGL: Transcribing Overlapping Speech with Staggered Labeling
by: Li, Chak-Fai, et al.
Published: (2024)
by: Li, Chak-Fai, et al.
Published: (2024)
A light-weight and efficient punctuation and word casing prediction model for on-device streaming ASR
by: You, Jian, et al.
Published: (2024)
by: You, Jian, et al.
Published: (2024)
In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions
by: Fan, Xulin, et al.
Published: (2026)
by: Fan, Xulin, et al.
Published: (2026)
Aligning Spoken Dialogue Models from User Interactions
by: Wu, Anne, et al.
Published: (2025)
by: Wu, Anne, et al.
Published: (2025)
XLS-R Deep Learning Model for Multilingual ASR on Low- Resource Languages: Indonesian, Javanese, and Sundanese
by: Arisaputra, Panji, et al.
Published: (2024)
by: Arisaputra, Panji, et al.
Published: (2024)
Enhancing Code-Switching ASR Leveraging Non-Peaky CTC Loss and Deep Language Posterior Injection
by: Yang, Tzu-Ting, et al.
Published: (2024)
by: Yang, Tzu-Ting, et al.
Published: (2024)
Text-only domain adaptation for end-to-end ASR using integrated text-to-mel-spectrogram generator
by: Bataev, Vladimir, et al.
Published: (2023)
by: Bataev, Vladimir, et al.
Published: (2023)
Exploring Pathological Speech Quality Assessment with ASR-Powered Wav2Vec2 in Data-Scarce Context
by: Nguyen, Tuan, et al.
Published: (2024)
by: Nguyen, Tuan, et al.
Published: (2024)
Evaluating Standard and Dialectal Frisian ASR: Multilingual Fine-tuning and Language Identification for Improved Low-resource Performance
by: Amooie, Reihaneh, et al.
Published: (2025)
by: Amooie, Reihaneh, et al.
Published: (2025)
Benchmarking Akan ASR Models Across Domain-Specific Datasets: A Comparative Evaluation of Performance, Scalability, and Adaptability
by: Mensah, Mark Atta, et al.
Published: (2025)
by: Mensah, Mark Atta, et al.
Published: (2025)
Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents
by: Veluri, Bandhav, et al.
Published: (2024)
by: Veluri, Bandhav, et al.
Published: (2024)
Task Oriented Dialogue as a Catalyst for Self-Supervised Automatic Speech Recognition
by: Chan, David M., et al.
Published: (2024)
by: Chan, David M., et al.
Published: (2024)
Transcribing Rhythmic Patterns of the Guitar Track in Polyphonic Music
by: Lukoianov, Aleksandr, et al.
Published: (2025)
by: Lukoianov, Aleksandr, et al.
Published: (2025)
WEE-Therapy: A Mixture of Weak Encoders Framework for Psychological Counseling Dialogue Analysis
by: Kang, Yongqi, et al.
Published: (2025)
by: Kang, Yongqi, et al.
Published: (2025)
Text-Based Detection of On-Hold Scripts in Contact Center Calls
by: Galimzianov, Dmitrii, et al.
Published: (2024)
by: Galimzianov, Dmitrii, et al.
Published: (2024)
ASR Benchmarking: Need for a More Representative Conversational Dataset
by: Maheshwari, Gaurav, et al.
Published: (2024)
by: Maheshwari, Gaurav, et al.
Published: (2024)
LibriConvo: Simulating Conversations from Read Literature for ASR and Diarization
by: Gedeon, Máté, et al.
Published: (2025)
by: Gedeon, Máté, et al.
Published: (2025)
Audio-to-Score Conversion Model Based on Whisper methodology
by: Zhang, Hongyao, et al.
Published: (2024)
by: Zhang, Hongyao, et al.
Published: (2024)
ChipChat: Low-Latency Cascaded Conversational Agent in MLX
by: Likhomanenko, Tatiana, et al.
Published: (2025)
by: Likhomanenko, Tatiana, et al.
Published: (2025)
Towards General-Purpose Text-Instruction-Guided Voice Conversion
by: Kuan, Chun-Yi, et al.
Published: (2023)
by: Kuan, Chun-Yi, et al.
Published: (2023)
Similar Items
-
Self-Train Before You Transcribe
by: Flynn, Robert, et al.
Published: (2024) -
ISPA: Inter-Species Phonetic Alphabet for Transcribing Animal Sounds
by: Hagiwara, Masato, et al.
Published: (2024) -
GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
by: Chen, Guoguo, et al.
Published: (2021) -
Anatomy of Industrial Scale Multilingual ASR
by: Ramirez, Francis McCann, et al.
Published: (2024) -
Beyond Transcription: Mechanistic Interpretability in ASR
by: Glazer, Neta, et al.
Published: (2025)