Continual Contrastive Spoken Language Understanding
Fuente:
arXiv
Guardado en:
| Autores principales: | Cappellazzo, Umberto, Fini, Enrico, Yang, Muqiao, Falavigna, Daniele, Brutti, Alessio, Raj, Bhiksha |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Evaluating and Improving Continual Learning in Spoken Language Understanding
por: Yang, Muqiao, et al.
Publicado: (2024)
por: Yang, Muqiao, et al.
Publicado: (2024)
Efficient Fine-tuning of Audio Spectrogram Transformers via Soft Mixture of Adapters
por: Cappellazzo, Umberto, et al.
Publicado: (2024)
por: Cappellazzo, Umberto, et al.
Publicado: (2024)
Parameter-Efficient Transfer Learning of Audio Spectrogram Transformers
por: Cappellazzo, Umberto, et al.
Publicado: (2023)
por: Cappellazzo, Umberto, et al.
Publicado: (2023)
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach
por: Cappellazzo, Umberto, et al.
Publicado: (2025)
por: Cappellazzo, Umberto, et al.
Publicado: (2025)
Training dynamic models using early exits for automatic speech recognition on resource-constrained devices
por: Wright, George August, et al.
Publicado: (2023)
por: Wright, George August, et al.
Publicado: (2023)
Large Language Models are Strong Audio-Visual Speech Recognition Learners
por: Cappellazzo, Umberto, et al.
Publicado: (2024)
por: Cappellazzo, Umberto, et al.
Publicado: (2024)
Input Conditioned Layer Dropping in Speech Foundation Models
por: Hannan, Abdul, et al.
Publicado: (2025)
por: Hannan, Abdul, et al.
Publicado: (2025)
AURA Score: A Metric For Holistic Audio Question Answering Evaluation
por: Dixit, Satvik, et al.
Publicado: (2025)
por: Dixit, Satvik, et al.
Publicado: (2025)
Speech LLMs in Low-Resource Scenarios: Data Volume Requirements and the Impact of Pretraining on High-Resource Languages
por: Fong, Seraphina, et al.
Publicado: (2025)
por: Fong, Seraphina, et al.
Publicado: (2025)
ADIFF: Explaining audio difference using natural language
por: Deshmukh, Soham, et al.
Publicado: (2025)
por: Deshmukh, Soham, et al.
Publicado: (2025)
Mellow: a small audio language model for reasoning
por: Deshmukh, Soham, et al.
Publicado: (2025)
por: Deshmukh, Soham, et al.
Publicado: (2025)
Domain Adaptation for Contrastive Audio-Language Models
por: Deshmukh, Soham, et al.
Publicado: (2024)
por: Deshmukh, Soham, et al.
Publicado: (2024)
Deciphering GunType Hierarchy through Acoustic Analysis of Gunshot Recordings
por: Shah, Ankit, et al.
Publicado: (2025)
por: Shah, Ankit, et al.
Publicado: (2025)
Splitformer: An improved early-exit architecture for automatic speech recognition on edge devices
por: Lasbordes, Maxence, et al.
Publicado: (2025)
por: Lasbordes, Maxence, et al.
Publicado: (2025)
Did You Hear That? Introducing AADG: A Framework for Generating Benchmark Data in Audio Anomaly Detection
por: Raghavan, Ksheeraja, et al.
Publicado: (2024)
por: Raghavan, Ksheeraja, et al.
Publicado: (2024)
HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling
por: Xue, Rongkun, et al.
Publicado: (2025)
por: Xue, Rongkun, et al.
Publicado: (2025)
Large Language Model Guided Decoding for Self-Supervised Speech Recognition
por: Cohen, Eyal, et al.
Publicado: (2025)
por: Cohen, Eyal, et al.
Publicado: (2025)
Efficient Autoregressive Audio Modeling via Next-Scale Prediction
por: Qiu, Kai, et al.
Publicado: (2024)
por: Qiu, Kai, et al.
Publicado: (2024)
SpokeN-100: A Cross-Lingual Benchmarking Dataset for The Classification of Spoken Numbers in Different Languages
por: Groh, René, et al.
Publicado: (2024)
por: Groh, René, et al.
Publicado: (2024)
Game-Time: Evaluating Temporal Dynamics in Spoken Language Models
por: Chang, Kai-Wei, et al.
Publicado: (2025)
por: Chang, Kai-Wei, et al.
Publicado: (2025)
Improving Speaker Representations Using Contrastive Losses on Multi-scale Features
por: Dixit, Satvik, et al.
Publicado: (2024)
por: Dixit, Satvik, et al.
Publicado: (2024)
Spoken question answering for visual queries
por: Shabtay, Nimrod, et al.
Publicado: (2025)
por: Shabtay, Nimrod, et al.
Publicado: (2025)
Human Voice is Unique
por: Singh, Rita, et al.
Publicado: (2025)
por: Singh, Rita, et al.
Publicado: (2025)
MOSEL: 950,000 Hours of Speech Data for Open-Source Speech Foundation Model Training on EU Languages
por: Gaido, Marco, et al.
Publicado: (2024)
por: Gaido, Marco, et al.
Publicado: (2024)
WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models
por: Chen, Yifu, et al.
Publicado: (2025)
por: Chen, Yifu, et al.
Publicado: (2025)
UniSLU: Unified Spoken Language Understanding from Heterogeneous Cross-Task Datasets
por: Sheng, Zhichao, et al.
Publicado: (2025)
por: Sheng, Zhichao, et al.
Publicado: (2025)
Data-Centric Improvements for Enhancing Multi-Modal Understanding in Spoken Conversation Modeling
por: Chen, Maximillian, et al.
Publicado: (2024)
por: Chen, Maximillian, et al.
Publicado: (2024)
Towards Unified Neural Decoding of Perceived, Spoken and Imagined Speech from EEG Signals
por: Lee, Jung-Sun, et al.
Publicado: (2024)
por: Lee, Jung-Sun, et al.
Publicado: (2024)
Evaluating Hallucinations in Audio-Visual Multimodal LLMs with Spoken Queries under Diverse Acoustic Conditions
por: Park, Hansol, et al.
Publicado: (2025)
por: Park, Hansol, et al.
Publicado: (2025)
Enhancing Neural Spoken Language Recognition: An Exploration with Multilingual Datasets
por: Anidjar, Or Haim, et al.
Publicado: (2025)
por: Anidjar, Or Haim, et al.
Publicado: (2025)
TiCo: Time-Controllable Spoken Dialogue Model
por: Chang, Kai-Wei, et al.
Publicado: (2026)
por: Chang, Kai-Wei, et al.
Publicado: (2026)
SPUR: A Plug-and-Play Framework for Integrating Spatial Audio Understanding and Reasoning into Large Audio-Language Models
por: Sakshi, S, et al.
Publicado: (2025)
por: Sakshi, S, et al.
Publicado: (2025)
Style-Talker: Finetuning Audio Language Model and Style-Based Text-to-Speech Model for Fast Spoken Dialogue Generation
por: Li, Yinghao Aaron, et al.
Publicado: (2024)
por: Li, Yinghao Aaron, et al.
Publicado: (2024)
GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness
por: Chen, Hongjie, et al.
Publicado: (2025)
por: Chen, Hongjie, et al.
Publicado: (2025)
Gibberish is All You Need for Membership Inference Detection in Contrastive Language-Audio Pretraining
por: Cheng, Ruoxi, et al.
Publicado: (2024)
por: Cheng, Ruoxi, et al.
Publicado: (2024)
Privacy-oriented manipulation of speaker representations
por: Teixeira, Francisco, et al.
Publicado: (2023)
por: Teixeira, Francisco, et al.
Publicado: (2023)
CoLLAP: Contrastive Long-form Language-Audio Pretraining with Musical Temporal Structure Augmentation
por: Wu, Junda, et al.
Publicado: (2024)
por: Wu, Junda, et al.
Publicado: (2024)
MACE: Leveraging Audio for Evaluating Audio Captioning Systems
por: Dixit, Satvik, et al.
Publicado: (2024)
por: Dixit, Satvik, et al.
Publicado: (2024)
TELEVAL: A Dynamic Benchmark Designed for Spoken Language Models in Chinese Interactive Scenarios
por: Li, Zehan, et al.
Publicado: (2025)
por: Li, Zehan, et al.
Publicado: (2025)
Data Augmentation for Spoken Grammatical Error Correction
por: Karanasou, Penny, et al.
Publicado: (2025)
por: Karanasou, Penny, et al.
Publicado: (2025)
Ejemplares similares
-
Evaluating and Improving Continual Learning in Spoken Language Understanding
por: Yang, Muqiao, et al.
Publicado: (2024) -
Efficient Fine-tuning of Audio Spectrogram Transformers via Soft Mixture of Adapters
por: Cappellazzo, Umberto, et al.
Publicado: (2024) -
Parameter-Efficient Transfer Learning of Audio Spectrogram Transformers
por: Cappellazzo, Umberto, et al.
Publicado: (2023) -
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach
por: Cappellazzo, Umberto, et al.
Publicado: (2025) -
Training dynamic models using early exits for automatic speech recognition on resource-constrained devices
por: Wright, George August, et al.
Publicado: (2023)