Open-Source Conversational AI with SpeechBrain 1.0
Fuente:
arXiv
Salvato in:
| Autori principali: | Ravanelli, Mirco, Parcollet, Titouan, Moumen, Adel, de Langen, Sylvain, Subakan, Cem, Plantinga, Peter, Wang, Yingzhi, Mousavi, Pooneh, Della Libera, Luca, Ploujnikov, Artem, Paissan, Francesco, Borra, Davide, Zaiem, Salah, Zhao, Zeyu, Zhang, Shucong, Karakasidis, Georgios, Yeh, Sung-Lin, Champion, Pierre, Rouhe, Aku, Braun, Rudolf, Mai, Florian, Zuluaga-Gomez, Juan, Mousavi, Seyed Mahed, Nautsch, Andreas, Nguyen, Ha, Liu, Xuechen, Sagar, Sangeet, Duret, Jarod, Mdhaffar, Salima, Laperriere, Gaelle, Rouvier, Mickael, De Mori, Renato, Esteve, Yannick |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
How Should We Extract Discrete Audio Tokens from Self-Supervised Models?
di: Mousavi, Pooneh, et al.
Pubblicazione: (2024)
di: Mousavi, Pooneh, et al.
Pubblicazione: (2024)
DASB - Discrete Audio and Speech Benchmark
di: Mousavi, Pooneh, et al.
Pubblicazione: (2024)
di: Mousavi, Pooneh, et al.
Pubblicazione: (2024)
What Are They Doing? Joint Audio-Speech Co-Reasoning
di: Wang, Yingzhi, et al.
Pubblicazione: (2024)
di: Wang, Yingzhi, et al.
Pubblicazione: (2024)
Listen First, Then Answer: Timestamp-Grounded Speech Reasoning
di: Jeong, Jihoon, et al.
Pubblicazione: (2026)
di: Jeong, Jihoon, et al.
Pubblicazione: (2026)
Investigating Faithfulness in Large Audio Language Models
di: Mousavi, Pooneh, et al.
Pubblicazione: (2025)
di: Mousavi, Pooneh, et al.
Pubblicazione: (2025)
LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs
di: Mousavi, Pooneh, et al.
Pubblicazione: (2025)
di: Mousavi, Pooneh, et al.
Pubblicazione: (2025)
ALAS: Measuring Latent Speech-Text Alignment For Spoken Language Understanding In Multimodal LLMs
di: Mousavi, Pooneh, et al.
Pubblicazione: (2025)
di: Mousavi, Pooneh, et al.
Pubblicazione: (2025)
FocalCodec: Low-Bitrate Speech Coding via Focal Modulation Networks
di: Della Libera, Luca, et al.
Pubblicazione: (2025)
di: Della Libera, Luca, et al.
Pubblicazione: (2025)
Listenable Maps for Zero-Shot Audio Classifiers
di: Paissan, Francesco, et al.
Pubblicazione: (2024)
di: Paissan, Francesco, et al.
Pubblicazione: (2024)
Exploring Token-Space Manipulation in Latent Audio Tokenizers
di: Paissan, Francesco, et al.
Pubblicazione: (2026)
di: Paissan, Francesco, et al.
Pubblicazione: (2026)
Listenable Maps for Audio Classifiers
di: Paissan, Francesco, et al.
Pubblicazione: (2024)
di: Paissan, Francesco, et al.
Pubblicazione: (2024)
Audio Editing with Non-Rigid Text Prompts
di: Paissan, Francesco, et al.
Pubblicazione: (2023)
di: Paissan, Francesco, et al.
Pubblicazione: (2023)
Virtual Consistency for Audio Editing
di: Cervera, Matthieu, et al.
Pubblicazione: (2025)
di: Cervera, Matthieu, et al.
Pubblicazione: (2025)
LMAC-TD: Producing Time Domain Explanations for Audio Classifiers
di: Mancini, Eleonora, et al.
Pubblicazione: (2024)
di: Mancini, Eleonora, et al.
Pubblicazione: (2024)
Beyond Fixed Frames: Dynamic Character-Aligned Speech Tokenization
di: Della Libera, Luca, et al.
Pubblicazione: (2026)
di: Della Libera, Luca, et al.
Pubblicazione: (2026)
Autoregressive Speech Enhancement via Acoustic Tokens
di: Della Libera, Luca, et al.
Pubblicazione: (2025)
di: Della Libera, Luca, et al.
Pubblicazione: (2025)
FocalCodec-Stream: Streaming Low-Bitrate Speech Coding via Causal Distillation
di: Della Libera, Luca, et al.
Pubblicazione: (2025)
di: Della Libera, Luca, et al.
Pubblicazione: (2025)
WavSLM: Single-Stream Speech Language Modeling via WavLM Distillation
di: Della Libera, Luca, et al.
Pubblicazione: (2026)
di: Della Libera, Luca, et al.
Pubblicazione: (2026)
Focal Modulation Networks for Interpretable Sound Classification
di: Della Libera, Luca, et al.
Pubblicazione: (2024)
di: Della Libera, Luca, et al.
Pubblicazione: (2024)
Speech Self-Supervised Representations Benchmarking: a Case for Larger Probing Heads
di: Zaiem, Salah, et al.
Pubblicazione: (2023)
di: Zaiem, Salah, et al.
Pubblicazione: (2023)
Investigating the Effectiveness of Explainability Methods in Parkinson's Detection from Speech
di: Mancini, Eleonora, et al.
Pubblicazione: (2024)
di: Mancini, Eleonora, et al.
Pubblicazione: (2024)
LL-SDR: Low-Latency Speech enhancement through Discrete Representations
di: Li, Jingyi, et al.
Pubblicazione: (2026)
di: Li, Jingyi, et al.
Pubblicazione: (2026)
Are LLMs Robust for Spoken Dialogues?
di: Mousavi, Seyed Mahed, et al.
Pubblicazione: (2024)
di: Mousavi, Seyed Mahed, et al.
Pubblicazione: (2024)
Less Forgetting for Better Generalization: Exploring Continual-learning Fine-tuning Methods for Speech Self-supervised Representations
di: Zaiem, Salah, et al.
Pubblicazione: (2024)
di: Zaiem, Salah, et al.
Pubblicazione: (2024)
What Does Loss Optimization Actually Teach, If Anything? Knowledge Dynamics in Continual Pre-training of LLMs
di: Mousavi, Seyed Mahed, et al.
Pubblicazione: (2026)
di: Mousavi, Seyed Mahed, et al.
Pubblicazione: (2026)
LLMs as Repositories of Factual Knowledge: Limitations and Solutions
di: Mousavi, Seyed Mahed, et al.
Pubblicazione: (2025)
di: Mousavi, Seyed Mahed, et al.
Pubblicazione: (2025)
DyKnow: Dynamically Verifying Time-Sensitive Factual Knowledge in LLMs
di: Mousavi, Seyed Mahed, et al.
Pubblicazione: (2024)
di: Mousavi, Seyed Mahed, et al.
Pubblicazione: (2024)
Discrete Audio Tokens: More Than a Survey!
di: Mousavi, Pooneh, et al.
Pubblicazione: (2025)
di: Mousavi, Pooneh, et al.
Pubblicazione: (2025)
MSP-Podcast SER Challenge 2024: L'antenne du Ventoux Multimodal Self-Supervised Learning for Speech Emotion Recognition
di: Duret, Jarod, et al.
Pubblicazione: (2024)
di: Duret, Jarod, et al.
Pubblicazione: (2024)
Resource-Efficient Separation Transformer
di: Della Libera, Luca, et al.
Pubblicazione: (2022)
di: Della Libera, Luca, et al.
Pubblicazione: (2022)
Getting to the Point: Pointing Improves LVLMs at Counting
di: Alghisi, Simone, et al.
Pubblicazione: (2026)
di: Alghisi, Simone, et al.
Pubblicazione: (2026)
Garbage In, Reasoning Out? Why Benchmark Scores are Unreliable and What to Do About It
di: Mousavi, Seyed Mahed, et al.
Pubblicazione: (2025)
di: Mousavi, Seyed Mahed, et al.
Pubblicazione: (2025)
Analyzing Speech Unit Selection for Textless Speech-to-Speech Translation
di: Duret, Jarod, et al.
Pubblicazione: (2024)
di: Duret, Jarod, et al.
Pubblicazione: (2024)
ProGRes: Prompted Generative Rescoring on ASR n-Best
di: Tur, Ada Defne, et al.
Pubblicazione: (2024)
di: Tur, Ada Defne, et al.
Pubblicazione: (2024)
Zero-Shot End-To-End Spoken Question Answering In Medical Domain
di: Labrak, Yanis, et al.
Pubblicazione: (2024)
di: Labrak, Yanis, et al.
Pubblicazione: (2024)
In-domain SSL pre-training and streaming ASR
di: Duret, Jarod, et al.
Pubblicazione: (2025)
di: Duret, Jarod, et al.
Pubblicazione: (2025)
V-DyKnow: A Dynamic Benchmark for Time-Sensitive Knowledge in Vision Language Models
di: Mousavi, Seyed Mahed, et al.
Pubblicazione: (2026)
di: Mousavi, Seyed Mahed, et al.
Pubblicazione: (2026)
Should We Fine-Tune or RAG? Evaluating Different Techniques to Adapt LLMs for Dialogue
di: Alghisi, Simone, et al.
Pubblicazione: (2024)
di: Alghisi, Simone, et al.
Pubblicazione: (2024)
[De|Re]constructing VLMs' Reasoning in Counting
di: Alghisi, Simone, et al.
Pubblicazione: (2025)
di: Alghisi, Simone, et al.
Pubblicazione: (2025)
Phoneme Discretized Saliency Maps for Explainable Detection of AI-Generated Voice
di: Gupta, Shubham, et al.
Pubblicazione: (2024)
di: Gupta, Shubham, et al.
Pubblicazione: (2024)
Documenti analoghi
-
How Should We Extract Discrete Audio Tokens from Self-Supervised Models?
di: Mousavi, Pooneh, et al.
Pubblicazione: (2024) -
DASB - Discrete Audio and Speech Benchmark
di: Mousavi, Pooneh, et al.
Pubblicazione: (2024) -
What Are They Doing? Joint Audio-Speech Co-Reasoning
di: Wang, Yingzhi, et al.
Pubblicazione: (2024) -
Listen First, Then Answer: Timestamp-Grounded Speech Reasoning
di: Jeong, Jihoon, et al.
Pubblicazione: (2026) -
Investigating Faithfulness in Large Audio Language Models
di: Mousavi, Pooneh, et al.
Pubblicazione: (2025)