Written Term Detection Improves Spoken Term Detection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yusuf, Bolaji, Saraçlar, Murat |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Pretraining End-to-End Keyword Search with Automatically Discovered Acoustic Units
von: Yusuf, Bolaji, et al.
Veröffentlicht: (2024)
von: Yusuf, Bolaji, et al.
Veröffentlicht: (2024)
Joint Speech and Text Training for LLM-Based End-to-End Spoken Dialogue State Tracking
von: Vendrame, Katia, et al.
Veröffentlicht: (2025)
von: Vendrame, Katia, et al.
Veröffentlicht: (2025)
Enhancing Few-shot Keyword Spotting Performance through Pre-Trained Self-supervised Speech Models
von: Gok, Alican, et al.
Veröffentlicht: (2025)
von: Gok, Alican, et al.
Veröffentlicht: (2025)
BEST-STD: Bidirectional Mamba-Enhanced Speech Tokenization for Spoken Term Detection
von: Singh, Anup, et al.
Veröffentlicht: (2024)
von: Singh, Anup, et al.
Veröffentlicht: (2024)
Spirit LM: Interleaved Spoken and Written Language Model
von: Nguyen, Tu Anh, et al.
Veröffentlicht: (2024)
von: Nguyen, Tu Anh, et al.
Veröffentlicht: (2024)
Speculative Speech Recognition by Audio-Prefixed Low-Rank Adaptation of Language Models
von: Yusuf, Bolaji, et al.
Veröffentlicht: (2024)
von: Yusuf, Bolaji, et al.
Veröffentlicht: (2024)
Approaching Dialogue State Tracking via Aligning Speech Encoders and LLMs
von: Sedláček, Šimon, et al.
Veröffentlicht: (2025)
von: Sedláček, Šimon, et al.
Veröffentlicht: (2025)
H-QuEST: Accelerating Query-by-Example Spoken Term Detection with Hierarchical Indexing
von: Singh, Akanksha, et al.
Veröffentlicht: (2025)
von: Singh, Akanksha, et al.
Veröffentlicht: (2025)
BEST-STD2.0: Balanced and Efficient Speech Tokenizer for Spoken Term Detection
von: Singh, Anup, et al.
Veröffentlicht: (2025)
von: Singh, Anup, et al.
Veröffentlicht: (2025)
Predictive Speech Recognition and End-of-Utterance Detection Towards Spoken Dialog Systems
von: Zink, Oswald, et al.
Veröffentlicht: (2024)
von: Zink, Oswald, et al.
Veröffentlicht: (2024)
BERT-LID: Leveraging BERT to Improve Spoken Language Identification
von: Nie, Yuting, et al.
Veröffentlicht: (2022)
von: Nie, Yuting, et al.
Veröffentlicht: (2022)
Scaling and Prompting for Improved End-to-End Spoken Grammatical Error Correction
von: Qian, Mengjie, et al.
Veröffentlicht: (2025)
von: Qian, Mengjie, et al.
Veröffentlicht: (2025)
Towards Hierarchical Spoken Language Dysfluency Modeling
von: Lian, Jiachen, et al.
Veröffentlicht: (2024)
von: Lian, Jiachen, et al.
Veröffentlicht: (2024)
Spoken-Term Discovery using Discrete Speech Units
von: van Niekerk, Benjamin, et al.
Veröffentlicht: (2024)
von: van Niekerk, Benjamin, et al.
Veröffentlicht: (2024)
Voice Conversion Improves Cross-Domain Robustness for Spoken Arabic Dialect Identification
von: Abdullah, Badr M., et al.
Veröffentlicht: (2025)
von: Abdullah, Badr M., et al.
Veröffentlicht: (2025)
Paralinguistics-Enhanced Large Language Modeling of Spoken Dialogue
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2023)
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2023)
SHANKS: Simultaneous Hearing and Thinking for Spoken Language Models
von: Chiang, Cheng-Han, et al.
Veröffentlicht: (2025)
von: Chiang, Cheng-Han, et al.
Veröffentlicht: (2025)
Scaling Spoken Language Models with Syllabic Speech Tokenization
von: Lee, Nicholas, et al.
Veröffentlicht: (2025)
von: Lee, Nicholas, et al.
Veröffentlicht: (2025)
Improving Spoken Language Modeling with Phoneme Classification: A Simple Fine-tuning Approach
von: Poli, Maxime, et al.
Veröffentlicht: (2024)
von: Poli, Maxime, et al.
Veröffentlicht: (2024)
Spoken Language Modeling with Duration-Penalized Self-Supervised Units
von: Visser, Nicol, et al.
Veröffentlicht: (2025)
von: Visser, Nicol, et al.
Veröffentlicht: (2025)
Towards a Japanese Full-duplex Spoken Dialogue System
von: Ohashi, Atsumoto, et al.
Veröffentlicht: (2025)
von: Ohashi, Atsumoto, et al.
Veröffentlicht: (2025)
Finetuning End-to-End Models for Estonian Conversational Spoken Language Translation
von: Sildam, Tiia, et al.
Veröffentlicht: (2024)
von: Sildam, Tiia, et al.
Veröffentlicht: (2024)
Evaluating Emotion Recognition in Spoken Language Models on Emotionally Incongruent Speech
von: Corrêa, Pedro, et al.
Veröffentlicht: (2025)
von: Corrêa, Pedro, et al.
Veröffentlicht: (2025)
EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Spoken Dialogue Systems
von: Liu, Jingwen, et al.
Veröffentlicht: (2025)
von: Liu, Jingwen, et al.
Veröffentlicht: (2025)
STITCH: Simultaneous Thinking and Talking with Chunked Reasoning for Spoken Language Models
von: Chiang, Cheng-Han, et al.
Veröffentlicht: (2025)
von: Chiang, Cheng-Han, et al.
Veröffentlicht: (2025)
LLM-Enhanced Dialogue Management for Full-Duplex Spoken Dialogue Systems
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
J-CHAT: Japanese Large-scale Spoken Dialogue Corpus for Spoken Dialogue Language Modeling
von: Nakata, Wataru, et al.
Veröffentlicht: (2024)
von: Nakata, Wataru, et al.
Veröffentlicht: (2024)
Finding Task-specific Subnetworks in Multi-task Spoken Language Understanding Model
von: Futami, Hayato, et al.
Veröffentlicht: (2024)
von: Futami, Hayato, et al.
Veröffentlicht: (2024)
ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching
von: Zhu, Han, et al.
Veröffentlicht: (2025)
von: Zhu, Han, et al.
Veröffentlicht: (2025)
ZeroSyl: Simple Zero-Resource Syllable Tokenization for Spoken Language Modeling
von: Visser, Nicol, et al.
Veröffentlicht: (2026)
von: Visser, Nicol, et al.
Veröffentlicht: (2026)
URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models
von: Yan, Ruiqi, et al.
Veröffentlicht: (2025)
von: Yan, Ruiqi, et al.
Veröffentlicht: (2025)
SPAR-K: Scheduled Periodic Alternating Early Exit for Spoken Language Models
von: Huang, Hsiao-Ying, et al.
Veröffentlicht: (2026)
von: Huang, Hsiao-Ying, et al.
Veröffentlicht: (2026)
Align-SLM: Textless Spoken Language Models with Reinforcement Learning from AI Feedback
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2024)
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2024)
Spoken Stereoset: On Evaluating Social Bias Toward Speaker in Speech Large Language Models
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2024)
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2024)
Proactive for Uncertainty: Cause-Aware Error Diagnosis and Interactive Clarification for Spoken Dialogue Systems
von: Peng, Yizhou, et al.
Veröffentlicht: (2026)
von: Peng, Yizhou, et al.
Veröffentlicht: (2026)
Towards ASR Robust Spoken Language Understanding Through In-Context Learning With Word Confusion Networks
von: Everson, Kevin, et al.
Veröffentlicht: (2024)
von: Everson, Kevin, et al.
Veröffentlicht: (2024)
Advancing Large Language Models to Capture Varied Speaking Styles and Respond Properly in Spoken Conversations
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2024)
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2024)
UltraVoice: Scaling Fine-Grained Style-Controlled Speech Conversations for Spoken Dialogue Models
von: Tu, Wenming, et al.
Veröffentlicht: (2025)
von: Tu, Wenming, et al.
Veröffentlicht: (2025)
A Large Dataset of Spontaneous Speech with the Accent Spoken in São Paulo for Automatic Speech Recognition Evaluation
von: Lima, Rodrigo, et al.
Veröffentlicht: (2024)
von: Lima, Rodrigo, et al.
Veröffentlicht: (2024)
FD-Bench: A Full-Duplex Benchmarking Pipeline Designed for Full Duplex Spoken Dialogue Systems
von: Peng, Yizhou, et al.
Veröffentlicht: (2025)
von: Peng, Yizhou, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Pretraining End-to-End Keyword Search with Automatically Discovered Acoustic Units
von: Yusuf, Bolaji, et al.
Veröffentlicht: (2024) -
Joint Speech and Text Training for LLM-Based End-to-End Spoken Dialogue State Tracking
von: Vendrame, Katia, et al.
Veröffentlicht: (2025) -
Enhancing Few-shot Keyword Spotting Performance through Pre-Trained Self-supervised Speech Models
von: Gok, Alican, et al.
Veröffentlicht: (2025) -
BEST-STD: Bidirectional Mamba-Enhanced Speech Tokenization for Spoken Term Detection
von: Singh, Anup, et al.
Veröffentlicht: (2024) -
Spirit LM: Interleaved Spoken and Written Language Model
von: Nguyen, Tu Anh, et al.
Veröffentlicht: (2024)