SELMA: A Speech-Enabled Language Model for Virtual Assistant Interactions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wagner, Dominik, Churchill, Alexander, Sigtia, Siddharth, Marchi, Erik |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Multimodal Approach to Device-Directed Speech Detection with Large Language Models
von: Wagner, Dominik, et al.
Veröffentlicht: (2024)
von: Wagner, Dominik, et al.
Veröffentlicht: (2024)
Adapting Language Balance in Code-Switching Speech
von: Ugan, Enes Yavuz, et al.
Veröffentlicht: (2025)
von: Ugan, Enes Yavuz, et al.
Veröffentlicht: (2025)
Textually Pretrained Speech Language Models
von: Hassid, Michael, et al.
Veröffentlicht: (2023)
von: Hassid, Michael, et al.
Veröffentlicht: (2023)
SpeechX: Neural Codec Language Model as a Versatile Speech Transformer
von: Wang, Xiaofei, et al.
Veröffentlicht: (2023)
von: Wang, Xiaofei, et al.
Veröffentlicht: (2023)
Modeling Overlapped Speech with Shuffles
von: Wiesner, Matthew, et al.
Veröffentlicht: (2026)
von: Wiesner, Matthew, et al.
Veröffentlicht: (2026)
Large Language Models for Dysfluency Detection in Stuttered Speech
von: Wagner, Dominik, et al.
Veröffentlicht: (2024)
von: Wagner, Dominik, et al.
Veröffentlicht: (2024)
Energy-Based Models with Applications to Speech and Language Processing
von: Ou, Zhijian
Veröffentlicht: (2024)
von: Ou, Zhijian
Veröffentlicht: (2024)
DC-Spin: A Speaker-invariant Speech Tokenizer for Spoken Language Models
von: Chang, Heng-Jui, et al.
Veröffentlicht: (2024)
von: Chang, Heng-Jui, et al.
Veröffentlicht: (2024)
Exploring Fine-Tuning of Large Audio Language Models for Spoken Language Understanding under Limited Speech Data
von: Choi, Youngwon, et al.
Veröffentlicht: (2025)
von: Choi, Youngwon, et al.
Veröffentlicht: (2025)
Spontaneous Style Text-to-Speech Synthesis with Controllable Spontaneous Behaviors Based on Language Models
von: Li, Weiqin, et al.
Veröffentlicht: (2024)
von: Li, Weiqin, et al.
Veröffentlicht: (2024)
In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions
von: Fan, Xulin, et al.
Veröffentlicht: (2026)
von: Fan, Xulin, et al.
Veröffentlicht: (2026)
Mai Ho'omāuna i ka 'Ai: Language Models Improve Automatic Speech Recognition in Hawaiian
von: Chaparala, Kaavya, et al.
Veröffentlicht: (2024)
von: Chaparala, Kaavya, et al.
Veröffentlicht: (2024)
On the Problem of Text-To-Speech Model Selection for Synthetic Data Generation in Automatic Speech Recognition
von: Rossenbach, Nick, et al.
Veröffentlicht: (2024)
von: Rossenbach, Nick, et al.
Veröffentlicht: (2024)
Meta Learning Text-to-Speech Synthesis in over 7000 Languages
von: Lux, Florian, et al.
Veröffentlicht: (2024)
von: Lux, Florian, et al.
Veröffentlicht: (2024)
Unseen Speaker and Language Adaptation for Lightweight Text-To-Speech with Adapters
von: Falai, Alessio, et al.
Veröffentlicht: (2025)
von: Falai, Alessio, et al.
Veröffentlicht: (2025)
Prosodic ABX: A Language-Agnostic Method for Measuring Prosodic Contrast in Speech Representations
von: Sun, Haitong, et al.
Veröffentlicht: (2026)
von: Sun, Haitong, et al.
Veröffentlicht: (2026)
Speech Robust Bench: A Robustness Benchmark For Speech Recognition
von: Shah, Muhammad A., et al.
Veröffentlicht: (2024)
von: Shah, Muhammad A., et al.
Veröffentlicht: (2024)
Analyzing Speech Unit Selection for Textless Speech-to-Speech Translation
von: Duret, Jarod, et al.
Veröffentlicht: (2024)
von: Duret, Jarod, et al.
Veröffentlicht: (2024)
Generative Pre-training for Speech with Flow Matching
von: Liu, Alexander H., et al.
Veröffentlicht: (2023)
von: Liu, Alexander H., et al.
Veröffentlicht: (2023)
Coupling Speech Encoders with Downstream Text Models
von: Chelba, Ciprian, et al.
Veröffentlicht: (2024)
von: Chelba, Ciprian, et al.
Veröffentlicht: (2024)
Speculative End-Turn Detector for Efficient Speech Chatbot Assistant
von: Ok, Hyunjong, et al.
Veröffentlicht: (2025)
von: Ok, Hyunjong, et al.
Veröffentlicht: (2025)
How Redundant Is the Transformer Stack in Speech Representation Models?
von: Dorszewski, Teresa, et al.
Veröffentlicht: (2024)
von: Dorszewski, Teresa, et al.
Veröffentlicht: (2024)
Sonos Voice Control Bias Assessment Dataset: A Methodology for Demographic Bias Assessment in Voice Assistants
von: Sekkat, Chloé, et al.
Veröffentlicht: (2024)
von: Sekkat, Chloé, et al.
Veröffentlicht: (2024)
ProMode: A Speech Prosody Model Conditioned on Acoustic and Textual Inputs
von: Eren, Eray, et al.
Veröffentlicht: (2025)
von: Eren, Eray, et al.
Veröffentlicht: (2025)
Analyzing Multimodal Features of Spontaneous Voice Assistant Commands for Mild Cognitive Impairment Detection
von: Lin, Nana, et al.
Veröffentlicht: (2024)
von: Lin, Nana, et al.
Veröffentlicht: (2024)
SimulTron: On-Device Simultaneous Speech to Speech Translation
von: Agranovich, Alex, et al.
Veröffentlicht: (2024)
von: Agranovich, Alex, et al.
Veröffentlicht: (2024)
Translatotron 3: Speech to Speech Translation with Monolingual Data
von: Nachmani, Eliya, et al.
Veröffentlicht: (2023)
von: Nachmani, Eliya, et al.
Veröffentlicht: (2023)
OLMoASR: Open Models and Data for Training Robust Speech Recognition Models
von: Ngo, Huong, et al.
Veröffentlicht: (2025)
von: Ngo, Huong, et al.
Veröffentlicht: (2025)
Rasa: Building Expressive Speech Synthesis Systems for Indian Languages in Low-resource Settings
von: Varadhan, Praveen Srinivasa, et al.
Veröffentlicht: (2024)
von: Varadhan, Praveen Srinivasa, et al.
Veröffentlicht: (2024)
On the Semantic Latent Space of Diffusion-Based Text-to-Speech Models
von: Varshavsky-Hassid, Miri, et al.
Veröffentlicht: (2024)
von: Varshavsky-Hassid, Miri, et al.
Veröffentlicht: (2024)
Towards Early Prediction of Self-Supervised Speech Model Performance
von: Whetten, Ryan, et al.
Veröffentlicht: (2025)
von: Whetten, Ryan, et al.
Veröffentlicht: (2025)
Towards Robust FastSpeech 2 by Modelling Residual Multimodality
von: Kögel, Fabian, et al.
Veröffentlicht: (2023)
von: Kögel, Fabian, et al.
Veröffentlicht: (2023)
Investigating Disentanglement in a Phoneme-level Speech Codec for Prosody Modeling
von: Karapiperis, Sotirios, et al.
Veröffentlicht: (2024)
von: Karapiperis, Sotirios, et al.
Veröffentlicht: (2024)
Pre-Trained Foundation Model representations to uncover Breathing patterns in Speech
von: Mitra, Vikramjit, et al.
Veröffentlicht: (2024)
von: Mitra, Vikramjit, et al.
Veröffentlicht: (2024)
Universal Robust Speech Adaptation for Cross-Domain Speech Recognition and Enhancement
von: Wang, Chien-Chun, et al.
Veröffentlicht: (2026)
von: Wang, Chien-Chun, et al.
Veröffentlicht: (2026)
Speech Recognition With LLMs Adapted to Disordered Speech Using Reinforcement Learning
von: Nagpal, Chirag, et al.
Veröffentlicht: (2024)
von: Nagpal, Chirag, et al.
Veröffentlicht: (2024)
Efficient Training of Self-Supervised Speech Foundation Models on a Compute Budget
von: Liu, Andy T., et al.
Veröffentlicht: (2024)
von: Liu, Andy T., et al.
Veröffentlicht: (2024)
CoSTA: Code-Switched Speech Translation using Aligned Speech-Text Interleaving
von: Shankar, Bhavani, et al.
Veröffentlicht: (2024)
von: Shankar, Bhavani, et al.
Veröffentlicht: (2024)
The ParlaSpeech Collection of Automatically Generated Speech and Text Datasets from Parliamentary Proceedings
von: Ljubešić, Nikola, et al.
Veröffentlicht: (2024)
von: Ljubešić, Nikola, et al.
Veröffentlicht: (2024)
Phir Hera Fairy: An English Fairytaler is a Strong Faker of Fluent Speech in Low-Resource Indian Languages
von: Varadhan, Praveen Srinivasa, et al.
Veröffentlicht: (2025)
von: Varadhan, Praveen Srinivasa, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Multimodal Approach to Device-Directed Speech Detection with Large Language Models
von: Wagner, Dominik, et al.
Veröffentlicht: (2024) -
Adapting Language Balance in Code-Switching Speech
von: Ugan, Enes Yavuz, et al.
Veröffentlicht: (2025) -
Textually Pretrained Speech Language Models
von: Hassid, Michael, et al.
Veröffentlicht: (2023) -
SpeechX: Neural Codec Language Model as a Versatile Speech Transformer
von: Wang, Xiaofei, et al.
Veröffentlicht: (2023) -
Modeling Overlapped Speech with Shuffles
von: Wiesner, Matthew, et al.
Veröffentlicht: (2026)