BiosERC: Integrating Biography Speakers Supported by LLMs for ERC Tasks
Fuente:
arXiv
Guardado en:
| Autores principales: | Xue, Jieying, Nguyen, Minh Phuong, Matheny, Blake, Nguyen, Le Minh |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
VoxVietnam: a Large-Scale Multi-Genre Dataset for Vietnamese Speaker Recognition
por: Vu, Hoang Long, et al.
Publicado: (2024)
por: Vu, Hoang Long, et al.
Publicado: (2024)
MultiMed-ST: Large-scale Many-to-many Multilingual Medical Speech Translation
por: Le-Duc, Khai, et al.
Publicado: (2025)
por: Le-Duc, Khai, et al.
Publicado: (2025)
Label-Context-Dependent Internal Language Model Estimation for CTC
por: Yang, Zijian, et al.
Publicado: (2025)
por: Yang, Zijian, et al.
Publicado: (2025)
Medical Spoken Named Entity Recognition
por: Le-Duc, Khai, et al.
Publicado: (2024)
por: Le-Duc, Khai, et al.
Publicado: (2024)
Speech Rhythm-Based Speaker Embeddings Extraction from Phonemes and Phoneme Duration for Multi-Speaker Speech Synthesis
por: Fujita, Kenichi, et al.
Publicado: (2024)
por: Fujita, Kenichi, et al.
Publicado: (2024)
mSTEB: Massively Multilingual Evaluation of LLMs on Speech and Text Tasks
por: Beyene, Luel Hagos, et al.
Publicado: (2025)
por: Beyene, Luel Hagos, et al.
Publicado: (2025)
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
por: Nguyen, Thai-Binh, et al.
Publicado: (2024)
por: Nguyen, Thai-Binh, et al.
Publicado: (2024)
Analyzing and Improving Speaker Similarity Assessment for Speech Synthesis
por: Carbonneau, Marc-André, et al.
Publicado: (2025)
por: Carbonneau, Marc-André, et al.
Publicado: (2025)
A conversational gesture synthesis system based on emotions and semantics
por: Hoang-Minh, Thanh
Publicado: (2025)
por: Hoang-Minh, Thanh
Publicado: (2025)
Towards Robust FastSpeech 2 by Modelling Residual Multimodality
por: Kögel, Fabian, et al.
Publicado: (2023)
por: Kögel, Fabian, et al.
Publicado: (2023)
Unseen Speaker and Language Adaptation for Lightweight Text-To-Speech with Adapters
por: Falai, Alessio, et al.
Publicado: (2025)
por: Falai, Alessio, et al.
Publicado: (2025)
DC-Spin: A Speaker-invariant Speech Tokenizer for Spoken Language Models
por: Chang, Heng-Jui, et al.
Publicado: (2024)
por: Chang, Heng-Jui, et al.
Publicado: (2024)
Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems
por: Park, Taejin, et al.
Publicado: (2024)
por: Park, Taejin, et al.
Publicado: (2024)
SKILL: Similarity-aware Knowledge distILLation for Speech Self-Supervised Learning
por: Zampierin, Luca, et al.
Publicado: (2024)
por: Zampierin, Luca, et al.
Publicado: (2024)
Exploring Pathological Speech Quality Assessment with ASR-Powered Wav2Vec2 in Data-Scarce Context
por: Nguyen, Tuan, et al.
Publicado: (2024)
por: Nguyen, Tuan, et al.
Publicado: (2024)
LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning
por: Kawamura, Masaya, et al.
Publicado: (2024)
por: Kawamura, Masaya, et al.
Publicado: (2024)
Exploring Fine-Tuning of Large Audio Language Models for Spoken Language Understanding under Limited Speech Data
por: Choi, Youngwon, et al.
Publicado: (2025)
por: Choi, Youngwon, et al.
Publicado: (2025)
Real-time Speech Summarization for Medical Conversations
por: Le-Duc, Khai, et al.
Publicado: (2024)
por: Le-Duc, Khai, et al.
Publicado: (2024)
Sentiment Reasoning for Healthcare
por: Nguyen, Khai-Nguyen, et al.
Publicado: (2024)
por: Nguyen, Khai-Nguyen, et al.
Publicado: (2024)
What You Read Isn't What You Hear: Linguistic Sensitivity in Deepfake Speech Detection
por: Nguyen, Binh, et al.
Publicado: (2025)
por: Nguyen, Binh, et al.
Publicado: (2025)
In-Context Learning Boosts Speech Recognition via Human-like Adaptation to Speakers and Language Varieties
por: Roll, Nathan, et al.
Publicado: (2025)
por: Roll, Nathan, et al.
Publicado: (2025)
Sonos Voice Control Bias Assessment Dataset: A Methodology for Demographic Bias Assessment in Voice Assistants
por: Sekkat, Chloé, et al.
Publicado: (2024)
por: Sekkat, Chloé, et al.
Publicado: (2024)
MultiMed: Multilingual Medical Speech Recognition via Attention Encoder Decoder
por: Le-Duc, Khai, et al.
Publicado: (2024)
por: Le-Duc, Khai, et al.
Publicado: (2024)
Remastering Divide and Remaster: A Cinematic Audio Source Separation Dataset with Multilingual Support
por: Watcharasupat, Karn N., et al.
Publicado: (2024)
por: Watcharasupat, Karn N., et al.
Publicado: (2024)
Task Oriented Dialogue as a Catalyst for Self-Supervised Automatic Speech Recognition
por: Chan, David M., et al.
Publicado: (2024)
por: Chan, David M., et al.
Publicado: (2024)
Evaluating the Effectiveness of Transformer Layers in Wav2Vec 2.0, XLS-R, and Whisper for Speaker Identification Tasks
por: Stuhlmann, Linus, et al.
Publicado: (2025)
por: Stuhlmann, Linus, et al.
Publicado: (2025)
Speech Recognition With LLMs Adapted to Disordered Speech Using Reinforcement Learning
por: Nagpal, Chirag, et al.
Publicado: (2024)
por: Nagpal, Chirag, et al.
Publicado: (2024)
Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents
por: Veluri, Bandhav, et al.
Publicado: (2024)
por: Veluri, Bandhav, et al.
Publicado: (2024)
VoxKnesset: A Large-Scale Longitudinal Hebrew Speech Dataset for Aging Speaker Modeling
por: Marmor, Yanir, et al.
Publicado: (2026)
por: Marmor, Yanir, et al.
Publicado: (2026)
TSPC: A Two-Stage Phoneme-Centric Architecture for code-switching Vietnamese-English Speech Recognition
por: Anh, Tran Nguyen, et al.
Publicado: (2025)
por: Anh, Tran Nguyen, et al.
Publicado: (2025)
Textually Pretrained Speech Language Models
por: Hassid, Michael, et al.
Publicado: (2023)
por: Hassid, Michael, et al.
Publicado: (2023)
Conformer-1: Robust ASR via Large-Scale Semisupervised Bootstrapping
por: Zhang, Kevin, et al.
Publicado: (2024)
por: Zhang, Kevin, et al.
Publicado: (2024)
A Semi-Automatic Approach to Create Large Gender- and Age-Balanced Speaker Corpora: Usefulness of Speaker Diarization & Identification
por: Uro, Rémi, et al.
Publicado: (2024)
por: Uro, Rémi, et al.
Publicado: (2024)
Investigation of Speaker Representation for Target-Speaker Speech Processing
por: Ashihara, Takanori, et al.
Publicado: (2024)
por: Ashihara, Takanori, et al.
Publicado: (2024)
DiariST: Streaming Speech Translation with Speaker Diarization
por: Yang, Mu, et al.
Publicado: (2023)
por: Yang, Mu, et al.
Publicado: (2023)
Qwen vs. Gemma Integration with Whisper: A Comparative Study in Multilingual SpeechLLM Systems
por: Nguyen, Tuan, et al.
Publicado: (2025)
por: Nguyen, Tuan, et al.
Publicado: (2025)
Generative Pre-training for Speech with Flow Matching
por: Liu, Alexander H., et al.
Publicado: (2023)
por: Liu, Alexander H., et al.
Publicado: (2023)
Zero-Shot Text-to-Speech for Vietnamese
por: Vu, Thi, et al.
Publicado: (2025)
por: Vu, Thi, et al.
Publicado: (2025)
Adaptive Noise Resilient Keyword Spotting Using One-Shot Learning
por: Martinez-Rau, Luciano Sebastian, et al.
Publicado: (2025)
por: Martinez-Rau, Luciano Sebastian, et al.
Publicado: (2025)
Gammatonegram Representation for End-to-End Dysarthric Speech Processing Tasks: Speech Recognition, Speaker Identification, and Intelligibility Assessment
por: Farhadipour, Aref, et al.
Publicado: (2023)
por: Farhadipour, Aref, et al.
Publicado: (2023)
Ejemplares similares
-
VoxVietnam: a Large-Scale Multi-Genre Dataset for Vietnamese Speaker Recognition
por: Vu, Hoang Long, et al.
Publicado: (2024) -
MultiMed-ST: Large-scale Many-to-many Multilingual Medical Speech Translation
por: Le-Duc, Khai, et al.
Publicado: (2025) -
Label-Context-Dependent Internal Language Model Estimation for CTC
por: Yang, Zijian, et al.
Publicado: (2025) -
Medical Spoken Named Entity Recognition
por: Le-Duc, Khai, et al.
Publicado: (2024) -
Speech Rhythm-Based Speaker Embeddings Extraction from Phonemes and Phoneme Duration for Multi-Speaker Speech Synthesis
por: Fujita, Kenichi, et al.
Publicado: (2024)