On the Evaluation of Speech Foundation Models for Spoken Language Understanding
Fuente:
arXiv
Salvato in:
| Autori principali: | Arora, Siddhant, Pasad, Ankita, Chien, Chung-Ming, Han, Jionghao, Sharma, Roshan, Jung, Jee-weon, Dhamyal, Hira, Chen, William, Shon, Suwon, Lee, Hung-yi, Livescu, Karen, Watanabe, Shinji |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization?
di: Sharma, Roshan, et al.
Pubblicazione: (2024)
di: Sharma, Roshan, et al.
Pubblicazione: (2024)
UniverSLU: Universal Spoken Language Understanding for Diverse Tasks with Natural Language Instructions
di: Arora, Siddhant, et al.
Pubblicazione: (2023)
di: Arora, Siddhant, et al.
Pubblicazione: (2023)
DiscreteSLU: A Large Language Model with Self-Supervised Discrete Speech Units for Spoken Language Understanding
di: Shon, Suwon, et al.
Pubblicazione: (2024)
di: Shon, Suwon, et al.
Pubblicazione: (2024)
What Do Self-Supervised Speech Models Know About Words?
di: Pasad, Ankita, et al.
Pubblicazione: (2023)
di: Pasad, Ankita, et al.
Pubblicazione: (2023)
Self-Supervised Speech Representations are More Phonetic than Semantic
di: Choi, Kwanghee, et al.
Pubblicazione: (2024)
di: Choi, Kwanghee, et al.
Pubblicazione: (2024)
What do Speech Foundation Models Learn? Analysis and Applications
di: Pasad, Ankita
Pubblicazione: (2025)
di: Pasad, Ankita
Pubblicazione: (2025)
Chain-of-Thought Training for Open E2E Spoken Dialogue Systems
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
On The Landscape of Spoken Language Models: A Comprehensive Survey
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
AugSumm: towards generalizable speech summarization using synthetic labels from large language model
di: Jung, Jee-weon, et al.
Pubblicazione: (2024)
di: Jung, Jee-weon, et al.
Pubblicazione: (2024)
Improving Design of Input Condition Invariant Speech Enhancement
di: Zhang, Wangyou, et al.
Pubblicazione: (2024)
di: Zhang, Wangyou, et al.
Pubblicazione: (2024)
TMT: Tri-Modal Translation between Speech, Image, and Text by Processing Different Modalities as Different Languages
di: Kim, Minsu, et al.
Pubblicazione: (2024)
di: Kim, Minsu, et al.
Pubblicazione: (2024)
Beyond Performance Plateaus: A Comprehensive Study on Scalability in Speech Enhancement
di: Zhang, Wangyou, et al.
Pubblicazione: (2024)
di: Zhang, Wangyou, et al.
Pubblicazione: (2024)
Objective Measurements of Voice Quality
di: Dhamyal, Hira, et al.
Pubblicazione: (2024)
di: Dhamyal, Hira, et al.
Pubblicazione: (2024)
On the Effects of Heterogeneous Data Sources on Speech-to-Text Foundation Models
di: Tian, Jinchuan, et al.
Pubblicazione: (2024)
di: Tian, Jinchuan, et al.
Pubblicazione: (2024)
Finding Task-specific Subnetworks in Multi-task Spoken Language Understanding Model
di: Futami, Hayato, et al.
Pubblicazione: (2024)
di: Futami, Hayato, et al.
Pubblicazione: (2024)
Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
OWSM v3.1: Better and Faster Open Whisper-Style Speech Models based on E-Branchformer
di: Peng, Yifan, et al.
Pubblicazione: (2024)
di: Peng, Yifan, et al.
Pubblicazione: (2024)
Improving ASR Contextual Biasing with Guided Attention
di: Tang, Jiyang, et al.
Pubblicazione: (2024)
di: Tang, Jiyang, et al.
Pubblicazione: (2024)
AV2Wav: Diffusion-Based Re-synthesis from Continuous Self-supervised Features for Audio-Visual Speech Enhancement
di: Chou, Ju-Chieh, et al.
Pubblicazione: (2023)
di: Chou, Ju-Chieh, et al.
Pubblicazione: (2023)
Context-Driven Dynamic Pruning for Large Speech Foundation Models
di: Someki, Masao, et al.
Pubblicazione: (2025)
di: Someki, Masao, et al.
Pubblicazione: (2025)
Beyond Silence: Bias Analysis through Loss and Asymmetric Approach in Audio Anti-Spoofing
di: Shim, Hye-jin, et al.
Pubblicazione: (2024)
di: Shim, Hye-jin, et al.
Pubblicazione: (2024)
Decoder-only Architecture for Streaming End-to-end Speech Recognition
di: Tsunoo, Emiru, et al.
Pubblicazione: (2024)
di: Tsunoo, Emiru, et al.
Pubblicazione: (2024)
SELM: Enhancing Speech Emotion Recognition for Out-of-Domain Scenarios
di: Bukhari, Hazim, et al.
Pubblicazione: (2024)
di: Bukhari, Hazim, et al.
Pubblicazione: (2024)
Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks
di: Maiti, Soumi, et al.
Pubblicazione: (2023)
di: Maiti, Soumi, et al.
Pubblicazione: (2023)
Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition
di: Wang, Shih-heng, et al.
Pubblicazione: (2024)
di: Wang, Shih-heng, et al.
Pubblicazione: (2024)
Decoder-only Architecture for Speech Recognition with CTC Prompts and Text Data Augmentation
di: Tsunoo, Emiru, et al.
Pubblicazione: (2023)
di: Tsunoo, Emiru, et al.
Pubblicazione: (2023)
Hypothesis Clustering and Merging: Novel MultiTalker Speech Recognition with Speaker Tokens
di: Kashiwagi, Yosuke, et al.
Pubblicazione: (2024)
di: Kashiwagi, Yosuke, et al.
Pubblicazione: (2024)
Chain-of-Thought Reasoning in Streaming Full-Duplex End-to-End Spoken Dialogue Systems
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
Optimizing Conversational Quality in Spoken Dialogue Systems with Reinforcement Learning from AI Feedback
di: Arora, Siddhant, et al.
Pubblicazione: (2026)
di: Arora, Siddhant, et al.
Pubblicazione: (2026)
Spiralformer: Low Latency Encoder for Streaming Speech Recognition with Circular Layer Skipping and Early Exiting
di: Tsunoo, Emiru, et al.
Pubblicazione: (2025)
di: Tsunoo, Emiru, et al.
Pubblicazione: (2025)
Rapid Language Adaptation for Multilingual E2E Speech Recognition Using Encoder Prompting
di: Kashiwagi, Yosuke, et al.
Pubblicazione: (2024)
di: Kashiwagi, Yosuke, et al.
Pubblicazione: (2024)
To what extent can ASV systems naturally defend against spoofing attacks?
di: Jung, Jee-weon, et al.
Pubblicazione: (2024)
di: Jung, Jee-weon, et al.
Pubblicazione: (2024)
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
di: Futami, Hayato, et al.
Pubblicazione: (2025)
di: Futami, Hayato, et al.
Pubblicazione: (2025)
TASTE-Streaming: Towards Streamable Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
di: Tseng, Liang-Hsuan, et al.
Pubblicazione: (2026)
di: Tseng, Liang-Hsuan, et al.
Pubblicazione: (2026)
Dynamic-SUPERB: Towards A Dynamic, Collaborative, and Comprehensive Instruction-Tuning Benchmark for Speech
di: Huang, Chien-yu, et al.
Pubblicazione: (2023)
di: Huang, Chien-yu, et al.
Pubblicazione: (2023)
Semi-Autoregressive Streaming ASR With Label Context
di: Arora, Siddhant, et al.
Pubblicazione: (2023)
di: Arora, Siddhant, et al.
Pubblicazione: (2023)
Flow-SLM: Joint Learning of Linguistic and Acoustic Information for Spoken Language Modeling
di: Chou, Ju-Chieh, et al.
Pubblicazione: (2025)
di: Chou, Ju-Chieh, et al.
Pubblicazione: (2025)
Can you Remove the Downstream Model for Speaker Recognition with Self-Supervised Speech Features?
di: Aldeneh, Zakaria, et al.
Pubblicazione: (2024)
di: Aldeneh, Zakaria, et al.
Pubblicazione: (2024)
MoshiRAG: Asynchronous Knowledge Retrieval for Full-Duplex Speech Language Models
di: Chien, Chung-Ming, et al.
Pubblicazione: (2026)
di: Chien, Chung-Ming, et al.
Pubblicazione: (2026)
ML-SUPERB 2.0: Benchmarking Multilingual Speech Models Across Modeling Constraints, Languages, and Datasets
di: Shi, Jiatong, et al.
Pubblicazione: (2024)
di: Shi, Jiatong, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization?
di: Sharma, Roshan, et al.
Pubblicazione: (2024) -
UniverSLU: Universal Spoken Language Understanding for Diverse Tasks with Natural Language Instructions
di: Arora, Siddhant, et al.
Pubblicazione: (2023) -
DiscreteSLU: A Large Language Model with Self-Supervised Discrete Speech Units for Spoken Language Understanding
di: Shon, Suwon, et al.
Pubblicazione: (2024) -
What Do Self-Supervised Speech Models Know About Words?
di: Pasad, Ankita, et al.
Pubblicazione: (2023) -
Self-Supervised Speech Representations are More Phonetic than Semantic
di: Choi, Kwanghee, et al.
Pubblicazione: (2024)