Speakers Unembedded: Embedding-free Approach to Long-form Neural Diarization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Xiang, Govindan, Vivek, Paturi, Rohit, Srinivasan, Sundararajan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AG-LSEC: Audio Grounded Lexical Speaker Error Correction
von: Paturi, Rohit, et al.
Veröffentlicht: (2024)
von: Paturi, Rohit, et al.
Veröffentlicht: (2024)
SEAL: Speaker Error Correction using Acoustic-conditioned Large Language Models
von: Kumar, Anurag, et al.
Veröffentlicht: (2025)
von: Kumar, Anurag, et al.
Veröffentlicht: (2025)
ASR-Synchronized Speaker-Role Diarization
von: Ghosh, Arindam, et al.
Veröffentlicht: (2025)
von: Ghosh, Arindam, et al.
Veröffentlicht: (2025)
Leveraging Speaker Embeddings in End-to-End Neural Diarization for Two-Speaker Scenarios
von: Alvarez-Trejos, Juan Ignacio, et al.
Veröffentlicht: (2024)
von: Alvarez-Trejos, Juan Ignacio, et al.
Veröffentlicht: (2024)
Exploring Speech Foundation Models for Speaker Diarization in Child-Adult Dyadic Interactions
von: Xu, Anfeng, et al.
Veröffentlicht: (2024)
von: Xu, Anfeng, et al.
Veröffentlicht: (2024)
Iterative LLM-based improvement for French Clinical Interview Transcription and Speaker Diarization
von: Marie, Ambre, et al.
Veröffentlicht: (2026)
von: Marie, Ambre, et al.
Veröffentlicht: (2026)
Make It Hard to Hear, Easy to Learn: Long-Form Bengali ASR and Speaker Diarization via Extreme Augmentation and Perfect Alignment
von: Hasan, Sanjid, et al.
Veröffentlicht: (2026)
von: Hasan, Sanjid, et al.
Veröffentlicht: (2026)
TIMIT Speaker Profiling: A Comparison of Multi-task learning and Single-task learning Approaches
von: Wang, Rong, et al.
Veröffentlicht: (2024)
von: Wang, Rong, et al.
Veröffentlicht: (2024)
DiarizationLM: Speaker Diarization Post-Processing with Large Language Models
von: Wang, Quan, et al.
Veröffentlicht: (2024)
von: Wang, Quan, et al.
Veröffentlicht: (2024)
From Modular to End-to-End Speaker Diarization
von: Landini, Federico
Veröffentlicht: (2024)
von: Landini, Federico
Veröffentlicht: (2024)
Personalized Speech Enhancement Without a Separate Speaker Embedding Model
von: Pärnamaa, Tanel, et al.
Veröffentlicht: (2024)
von: Pärnamaa, Tanel, et al.
Veröffentlicht: (2024)
Improving Neural Diarization through Speaker Attribute Attractors and Local Dependency Modeling
von: Palzer, David, et al.
Veröffentlicht: (2025)
von: Palzer, David, et al.
Veröffentlicht: (2025)
SDBench: A Comprehensive Benchmark Suite for Speaker Diarization
von: Pacheco, Eduardo, et al.
Veröffentlicht: (2025)
von: Pacheco, Eduardo, et al.
Veröffentlicht: (2025)
Investigating Confidence Estimation Measures for Speaker Diarization
von: Chowdhury, Anurag, et al.
Veröffentlicht: (2024)
von: Chowdhury, Anurag, et al.
Veröffentlicht: (2024)
Multi-Stage Speaker Diarization for Noisy Classrooms
von: Khan, Ali Sartaz, et al.
Veröffentlicht: (2025)
von: Khan, Ali Sartaz, et al.
Veröffentlicht: (2025)
Language Modelling for Speaker Diarization in Telephonic Interviews
von: India, Miquel, et al.
Veröffentlicht: (2025)
von: India, Miquel, et al.
Veröffentlicht: (2025)
Speaker- and Text-Independent Estimation of Articulatory Movements and Phoneme Alignments from Speech
von: Weise, Tobias, et al.
Veröffentlicht: (2024)
von: Weise, Tobias, et al.
Veröffentlicht: (2024)
A Semi-Automatic Approach to Create Large Gender- and Age-Balanced Speaker Corpora: Usefulness of Speaker Diarization & Identification
von: Uro, Rémi, et al.
Veröffentlicht: (2024)
von: Uro, Rémi, et al.
Veröffentlicht: (2024)
End-to-End Supervised Hierarchical Graph Clustering for Speaker Diarization
von: Singh, Prachi, et al.
Veröffentlicht: (2024)
von: Singh, Prachi, et al.
Veröffentlicht: (2024)
SimulU: Training-free Policy for Long-form Simultaneous Speech-to-Speech Translation
von: Djanibekov, Amirbek, et al.
Veröffentlicht: (2026)
von: Djanibekov, Amirbek, et al.
Veröffentlicht: (2026)
Learning to Rewrite Prompts for Bootstrapping LLMs on Downstream Tasks
von: Zhou, Qinhao, et al.
Veröffentlicht: (2025)
von: Zhou, Qinhao, et al.
Veröffentlicht: (2025)
Mitigating Intra-Speaker Variability in Diarization with Style-Controllable Speech Augmentation
von: Kim, Miseul, et al.
Veröffentlicht: (2025)
von: Kim, Miseul, et al.
Veröffentlicht: (2025)
Large Language Model Based Generative Error Correction: A Challenge and Baselines for Speech Recognition, Speaker Tagging, and Emotion Recognition
von: Yang, Chao-Han Huck, et al.
Veröffentlicht: (2024)
von: Yang, Chao-Han Huck, et al.
Veröffentlicht: (2024)
Evaluating the Effectiveness of Transformer Layers in Wav2Vec 2.0, XLS-R, and Whisper for Speaker Identification Tasks
von: Stuhlmann, Linus, et al.
Veröffentlicht: (2025)
von: Stuhlmann, Linus, et al.
Veröffentlicht: (2025)
ASoBO: Attentive Beamformer Selection for Distant Speaker Diarization in Meetings
von: Mariotte, Theo, et al.
Veröffentlicht: (2024)
von: Mariotte, Theo, et al.
Veröffentlicht: (2024)
Self-Tuning Spectral Clustering for Speaker Diarization
von: Raghav, Nikhil, et al.
Veröffentlicht: (2024)
von: Raghav, Nikhil, et al.
Veröffentlicht: (2024)
GenTSE: Enhancing Target Speaker Extraction via a Coarse-to-Fine Generative Language Model
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
Latent Speech-Text Transformer
von: Lu, Yen-Ju, et al.
Veröffentlicht: (2025)
von: Lu, Yen-Ju, et al.
Veröffentlicht: (2025)
Speaker Diarization with Overlapping Community Detection Using Graph Attention Networks and Label Propagation Algorithm
von: Li, Zhaoyang, et al.
Veröffentlicht: (2025)
von: Li, Zhaoyang, et al.
Veröffentlicht: (2025)
Speech Enhancement Using Continuous Embeddings of Neural Audio Codec
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
Bangla-WhisperDiar: Fine-Tuning Whisper and PyAnnote for Bangla Long-Form Speech Recognition and Speaker Diarization
von: Bhuiyan, Mohammed Aman, et al.
Veröffentlicht: (2026)
von: Bhuiyan, Mohammed Aman, et al.
Veröffentlicht: (2026)
Towards Speaker Identification with Minimal Dataset and Constrained Resources using 1D-Convolution Neural Network
von: Shahan, Irfan Nafiz, et al.
Veröffentlicht: (2024)
von: Shahan, Irfan Nafiz, et al.
Veröffentlicht: (2024)
The Speech-LLM Takes It All: A Truly Fully End-to-End Spoken Dialogue State Tracking Approach
von: Ghazal, Nizar El, et al.
Veröffentlicht: (2025)
von: Ghazal, Nizar El, et al.
Veröffentlicht: (2025)
Unsupervised Speaker Diarization in Distributed IoT Networks Using Federated Learning
von: Bhuyan, Amit Kumar, et al.
Veröffentlicht: (2024)
von: Bhuyan, Amit Kumar, et al.
Veröffentlicht: (2024)
SpeakerLLM: A Speaker-Specialized Audio-LLM for Speaker Understanding and Verification Reasoning
von: Nam, KiHyun, et al.
Veröffentlicht: (2026)
von: Nam, KiHyun, et al.
Veröffentlicht: (2026)
Neural Blind Source Separation and Diarization for Distant Speech Recognition
von: Bando, Yoshiaki, et al.
Veröffentlicht: (2024)
von: Bando, Yoshiaki, et al.
Veröffentlicht: (2024)
COMET: Concept Space Dissection of the Modality Gap in Audio-Text Multimodal Contrastive Embeddings
von: Zhu, Yonggang, et al.
Veröffentlicht: (2026)
von: Zhu, Yonggang, et al.
Veröffentlicht: (2026)
LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2025)
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2025)
A Unified Speech LLM for Diarization and Speech Recognition in Multilingual Conversations
von: Saengthong, Phurich, et al.
Veröffentlicht: (2025)
von: Saengthong, Phurich, et al.
Veröffentlicht: (2025)
TICL: Text-Embedding KNN For Speech In-Context Learning Unlocks Speech Recognition Abilities of Large Multimodal Models
von: Zheng, Haolong, et al.
Veröffentlicht: (2025)
von: Zheng, Haolong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
AG-LSEC: Audio Grounded Lexical Speaker Error Correction
von: Paturi, Rohit, et al.
Veröffentlicht: (2024) -
SEAL: Speaker Error Correction using Acoustic-conditioned Large Language Models
von: Kumar, Anurag, et al.
Veröffentlicht: (2025) -
ASR-Synchronized Speaker-Role Diarization
von: Ghosh, Arindam, et al.
Veröffentlicht: (2025) -
Leveraging Speaker Embeddings in End-to-End Neural Diarization for Two-Speaker Scenarios
von: Alvarez-Trejos, Juan Ignacio, et al.
Veröffentlicht: (2024) -
Exploring Speech Foundation Models for Speaker Diarization in Child-Adult Dyadic Interactions
von: Xu, Anfeng, et al.
Veröffentlicht: (2024)