Gespeichert in:
| Hauptverfasser: | Zhu, Yi, Falk, Tiago |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2406.18731 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
WavLLM: Towards Robust and Adaptive Speech Large Language Model
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling
von: Yang, Guanrou, et al.
Veröffentlicht: (2026)
von: Yang, Guanrou, et al.
Veröffentlicht: (2026)
Evaluating the Usefulness of Non-Diagnostic Speech Data for Developing Parkinson's Disease Classifiers
von: Zhong, Terry Yi, et al.
Veröffentlicht: (2025)
von: Zhong, Terry Yi, et al.
Veröffentlicht: (2025)
Continued Pretraining for Domain Adaptation of Wav2vec2.0 in Automatic Speech Recognition for Elementary Math Classroom Settings
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2024)
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2024)
Regularized Federated Learning for Privacy-Preserving Dysarthric and Elderly Speech Recognition
von: Zhong, Tao, et al.
Veröffentlicht: (2025)
von: Zhong, Tao, et al.
Veröffentlicht: (2025)
On the Impact of Voice Anonymization on Speech Diagnostic Applications: a Case Study on COVID-19 Detection
von: Zhu, Yi, et al.
Veröffentlicht: (2023)
von: Zhu, Yi, et al.
Veröffentlicht: (2023)
Exploring ASR-Based Wav2Vec2 for Automated Speech Disorder Assessment: Insights and Analysis
von: Nguyen, Tuan, et al.
Veröffentlicht: (2024)
von: Nguyen, Tuan, et al.
Veröffentlicht: (2024)
Human-like Linguistic Biases in Neural Speech Models: Phonetic Categorization and Phonotactic Constraints in Wav2Vec2.0
von: Kloots, Marianne de Heer, et al.
Veröffentlicht: (2024)
von: Kloots, Marianne de Heer, et al.
Veröffentlicht: (2024)
Unifying Model and Layer Fusion for Speech Foundation Models
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2025)
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2025)
TransVIP: Speech to Speech Translation System with Voice and Isochrony Preservation
von: Le, Chenyang, et al.
Veröffentlicht: (2024)
von: Le, Chenyang, et al.
Veröffentlicht: (2024)
Adaptability of ASR Models on Low-Resource Language: A Comparative Study of Whisper and Wav2Vec-BERT on Bangla
von: Ridoy, Md Sazzadul Islam, et al.
Veröffentlicht: (2025)
von: Ridoy, Md Sazzadul Islam, et al.
Veröffentlicht: (2025)
Self-supervised Speech Models for Word-Level Stuttered Speech Detection
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2024)
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2024)
MERaLiON-SpeechEncoder: Towards a Speech Foundation Model for Singapore and Beyond
von: Huzaifah, Muhammad, et al.
Veröffentlicht: (2024)
von: Huzaifah, Muhammad, et al.
Veröffentlicht: (2024)
A Benchmark for Early-stage Parkinson's Disease Detection from Speech
von: Zhong, Terry Yi, et al.
Veröffentlicht: (2026)
von: Zhong, Terry Yi, et al.
Veröffentlicht: (2026)
WavShape: Information-Theoretic Speech Representation Learning for Fair and Privacy-Aware Audio Processing
von: Baser, Oguzhan, et al.
Veröffentlicht: (2025)
von: Baser, Oguzhan, et al.
Veröffentlicht: (2025)
Transsion Multilingual Speech Recognition System for MLC-SLM 2025 Challenge
von: Li, Xiaoxiao, et al.
Veröffentlicht: (2025)
von: Li, Xiaoxiao, et al.
Veröffentlicht: (2025)
Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
Exploring In-Context Learning of Textless Speech Language Model for Speech Classification Tasks
von: Hsu, Ming-Hao, et al.
Veröffentlicht: (2023)
von: Hsu, Ming-Hao, et al.
Veröffentlicht: (2023)
Streaming Speech-to-Text Translation with a SpeechLLM
von: Parcollet, Titouan, et al.
Veröffentlicht: (2026)
von: Parcollet, Titouan, et al.
Veröffentlicht: (2026)
Style-Talker: Finetuning Audio Language Model and Style-Based Text-to-Speech Model for Fast Spoken Dialogue Generation
von: Li, Yinghao Aaron, et al.
Veröffentlicht: (2024)
von: Li, Yinghao Aaron, et al.
Veröffentlicht: (2024)
RECA-PD: A Robust Explainable Cross-Attention Method for Speech-based Parkinson's Disease Classification
von: Zhong, Terry Yi, et al.
Veröffentlicht: (2025)
von: Zhong, Terry Yi, et al.
Veröffentlicht: (2025)
MTR-DuplexBench: Towards a Comprehensive Evaluation of Multi-Round Conversations for Full-Duplex Speech Language Models
von: Zhang, He, et al.
Veröffentlicht: (2025)
von: Zhang, He, et al.
Veröffentlicht: (2025)
Can Speech LLMs Think while Listening?
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2025)
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2025)
HARNESS: Lightweight Distilled Arabic Speech Foundation Models
von: Sukhadia, Vrunda N., et al.
Veröffentlicht: (2026)
von: Sukhadia, Vrunda N., et al.
Veröffentlicht: (2026)
Phonology-Guided Speech-to-Speech Translation for African Languages
von: Ochieng, Peter, et al.
Veröffentlicht: (2024)
von: Ochieng, Peter, et al.
Veröffentlicht: (2024)
Incorporating Contextual Paralinguistic Understanding in Large Speech-Language Models
von: Wang, Qiongqiong, et al.
Veröffentlicht: (2025)
von: Wang, Qiongqiong, et al.
Veröffentlicht: (2025)
Speech Retrieval-Augmented Generation without Automatic Speech Recognition
von: Min, Do June, et al.
Veröffentlicht: (2024)
von: Min, Do June, et al.
Veröffentlicht: (2024)
What Makes a Good Speech Tokenizer for LLM-Centric Speech Generation? A Systematic Study
von: Fan, Xiaoran, et al.
Veröffentlicht: (2025)
von: Fan, Xiaoran, et al.
Veröffentlicht: (2025)
Privacy-Preserving End-to-End Full-Duplex Speech Dialogue Models
von: Kuzmin, Nikita, et al.
Veröffentlicht: (2026)
von: Kuzmin, Nikita, et al.
Veröffentlicht: (2026)
Conversational Speech Reveals Structural Robustness Failures in SpeechLLM Backbones
von: Teleki, Maria, et al.
Veröffentlicht: (2025)
von: Teleki, Maria, et al.
Veröffentlicht: (2025)
SyllableLM: Learning Coarse Semantic Units for Speech Language Models
von: Baade, Alan, et al.
Veröffentlicht: (2024)
von: Baade, Alan, et al.
Veröffentlicht: (2024)
Pretraining Large Brain Language Model for Active BCI: Silent Speech
von: Zhou, Jinzhao, et al.
Veröffentlicht: (2025)
von: Zhou, Jinzhao, et al.
Veröffentlicht: (2025)
End to end Hindi to English speech conversion using Bark, mBART and a finetuned XLSR Wav2Vec2
von: Tathe, Aniket, et al.
Veröffentlicht: (2024)
von: Tathe, Aniket, et al.
Veröffentlicht: (2024)
SeamlessExpressiveLM: Speech Language Model for Expressive Speech-to-Speech Translation with Chain-of-Thought
von: Gong, Hongyu, et al.
Veröffentlicht: (2024)
von: Gong, Hongyu, et al.
Veröffentlicht: (2024)
An Empirical Analysis of Discrete Unit Representations in Speech Language Modeling Pre-training
von: Labrak, Yanis, et al.
Veröffentlicht: (2025)
von: Labrak, Yanis, et al.
Veröffentlicht: (2025)
Lost in Phonation: Voice Quality Variation as an Evaluation Dimension for Speech Foundation Models
von: Lameris, Harm, et al.
Veröffentlicht: (2025)
von: Lameris, Harm, et al.
Veröffentlicht: (2025)
SPBA: Utilizing Speech Large Language Model for Backdoor Attacks on Speech Classification Models
von: Yao, Wenhan, et al.
Veröffentlicht: (2025)
von: Yao, Wenhan, et al.
Veröffentlicht: (2025)
KAME: Tandem Architecture for Enhancing Knowledge in Real-Time Speech-to-Speech Conversational AI
von: Kuroki, So, et al.
Veröffentlicht: (2025)
von: Kuroki, So, et al.
Veröffentlicht: (2025)
SimulU: Training-free Policy for Long-form Simultaneous Speech-to-Speech Translation
von: Djanibekov, Amirbek, et al.
Veröffentlicht: (2026)
von: Djanibekov, Amirbek, et al.
Veröffentlicht: (2026)
FastAdaSP: Multitask-Adapted Efficient Inference for Large Speech Language Model
von: Lu, Yichen, et al.
Veröffentlicht: (2024)
von: Lu, Yichen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
WavLLM: Towards Robust and Adaptive Speech Large Language Model
von: Hu, Shujie, et al.
Veröffentlicht: (2024) -
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling
von: Yang, Guanrou, et al.
Veröffentlicht: (2026) -
Evaluating the Usefulness of Non-Diagnostic Speech Data for Developing Parkinson's Disease Classifiers
von: Zhong, Terry Yi, et al.
Veröffentlicht: (2025) -
Continued Pretraining for Domain Adaptation of Wav2vec2.0 in Automatic Speech Recognition for Elementary Math Classroom Settings
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2024) -
Regularized Federated Learning for Privacy-Preserving Dysarthric and Elderly Speech Recognition
von: Zhong, Tao, et al.
Veröffentlicht: (2025)