Can you Remove the Downstream Model for Speaker Recognition with Self-Supervised Speech Features?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Aldeneh, Zakaria, Higuchi, Takuya, Jung, Jee-weon, Seto, Skyler, Likhomanenko, Tatiana, Shum, Stephen, Abdelaziz, Ahmed Hussen, Watanabe, Shinji, Theobald, Barry-John |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Speaker-IPL: Unsupervised Learning of Speaker Characteristics with i-Vector based Pseudo-Labels
von: Aldeneh, Zakaria, et al.
Veröffentlicht: (2024)
von: Aldeneh, Zakaria, et al.
Veröffentlicht: (2024)
Which Data Matter? Embedding-Based Data Selection for Speech Recognition
von: Aldeneh, Zakaria, et al.
Veröffentlicht: (2026)
von: Aldeneh, Zakaria, et al.
Veröffentlicht: (2026)
Towards Automatic Assessment of Self-Supervised Speech Models using Rank
von: Aldeneh, Zakaria, et al.
Veröffentlicht: (2024)
von: Aldeneh, Zakaria, et al.
Veröffentlicht: (2024)
ESPnet-SPK: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models
von: Jung, Jee-weon, et al.
Veröffentlicht: (2024)
von: Jung, Jee-weon, et al.
Veröffentlicht: (2024)
Exploring Prediction Targets in Masked Pre-Training for Speech Foundation Models
von: Chen, Li-Wei, et al.
Veröffentlicht: (2024)
von: Chen, Li-Wei, et al.
Veröffentlicht: (2024)
DiceHuBERT: Distilling HuBERT with a Self-Supervised Learning Objective
von: Chi, Hyung Gun, et al.
Veröffentlicht: (2025)
von: Chi, Hyung Gun, et al.
Veröffentlicht: (2025)
A Variational Framework for Improving Naturalness in Generative Spoken Language Models
von: Chen, Li-Wei, et al.
Veröffentlicht: (2025)
von: Chen, Li-Wei, et al.
Veröffentlicht: (2025)
Closing the Gap Between Text and Speech Understanding in LLMs
von: Cuervo, Santiago, et al.
Veröffentlicht: (2025)
von: Cuervo, Santiago, et al.
Veröffentlicht: (2025)
Learning Spatially-Aware Language and Audio Embeddings
von: Devnani, Bhavika, et al.
Veröffentlicht: (2024)
von: Devnani, Bhavika, et al.
Veröffentlicht: (2024)
dMel: Speech Tokenization made Simple
von: Bai, Richard He, et al.
Veröffentlicht: (2024)
von: Bai, Richard He, et al.
Veröffentlicht: (2024)
Improving Design of Input Condition Invariant Speech Enhancement
von: Zhang, Wangyou, et al.
Veröffentlicht: (2024)
von: Zhang, Wangyou, et al.
Veröffentlicht: (2024)
Mechanisms of Multimodal Synchronization: Insights from Decoder-Based Video-Text-to-Speech Synthesis
von: Gupta, Akshita, et al.
Veröffentlicht: (2024)
von: Gupta, Akshita, et al.
Veröffentlicht: (2024)
Beyond Performance Plateaus: A Comprehensive Study on Scalability in Speech Enhancement
von: Zhang, Wangyou, et al.
Veröffentlicht: (2024)
von: Zhang, Wangyou, et al.
Veröffentlicht: (2024)
Disentangled Representation Learning for Environment-agnostic Speaker Recognition
von: Nam, KiHyun, et al.
Veröffentlicht: (2024)
von: Nam, KiHyun, et al.
Veröffentlicht: (2024)
Omni-Router: Sharing Routing Decisions in Sparse Mixture-of-Experts for Speech Recognition
von: Gu, Zijin, et al.
Veröffentlicht: (2025)
von: Gu, Zijin, et al.
Veröffentlicht: (2025)
Speech Foundation Models Generalize to Time Series Tasks from Wearable Sensor Data
von: Narain, Jaya, et al.
Veröffentlicht: (2025)
von: Narain, Jaya, et al.
Veröffentlicht: (2025)
End-to-End Speech Recognition with Pre-trained Masked Language Model
von: Higuchi, Yosuke, et al.
Veröffentlicht: (2024)
von: Higuchi, Yosuke, et al.
Veröffentlicht: (2024)
Leveraging Audio-Visual Data to Reduce the Multilingual Gap in Self-Supervised Speech Models
von: Blandón, María Andrea Cruz, et al.
Veröffentlicht: (2025)
von: Blandón, María Andrea Cruz, et al.
Veröffentlicht: (2025)
AugSumm: towards generalizable speech summarization using synthetic labels from large language model
von: Jung, Jee-weon, et al.
Veröffentlicht: (2024)
von: Jung, Jee-weon, et al.
Veröffentlicht: (2024)
ExpertLens: Activation steering features are highly interpretable
von: Fedzechkina, Masha, et al.
Veröffentlicht: (2025)
von: Fedzechkina, Masha, et al.
Veröffentlicht: (2025)
Steering into New Embedding Spaces: Analyzing Cross-Lingual Alignment Induced by Model Interventions in Multilingual Language Models
von: Sundar, Anirudh, et al.
Veröffentlicht: (2025)
von: Sundar, Anirudh, et al.
Veröffentlicht: (2025)
The VoxCeleb Speaker Recognition Challenge: A Retrospective
von: Huh, Jaesung, et al.
Veröffentlicht: (2024)
von: Huh, Jaesung, et al.
Veröffentlicht: (2024)
Hypothesis Clustering and Merging: Novel MultiTalker Speech Recognition with Speaker Tokens
von: Kashiwagi, Yosuke, et al.
Veröffentlicht: (2024)
von: Kashiwagi, Yosuke, et al.
Veröffentlicht: (2024)
Beyond Silence: Bias Analysis through Loss and Asymmetric Approach in Audio Anti-Spoofing
von: Shim, Hye-jin, et al.
Veröffentlicht: (2024)
von: Shim, Hye-jin, et al.
Veröffentlicht: (2024)
Improving Audio Captioning Models with Fine-grained Audio Features, Text Embedding Supervision, and LLM Mix-up Augmentation
von: Wu, Shih-Lun, et al.
Veröffentlicht: (2023)
von: Wu, Shih-Lun, et al.
Veröffentlicht: (2023)
Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks
von: Maiti, Soumi, et al.
Veröffentlicht: (2023)
von: Maiti, Soumi, et al.
Veröffentlicht: (2023)
Aggregate-and-Adapt Natural Language Prompts for Downstream Generalization of CLIP
von: Huang, Chen, et al.
Veröffentlicht: (2024)
von: Huang, Chen, et al.
Veröffentlicht: (2024)
TMT: Tri-Modal Translation between Speech, Image, and Text by Processing Different Modalities as Different Languages
von: Kim, Minsu, et al.
Veröffentlicht: (2024)
von: Kim, Minsu, et al.
Veröffentlicht: (2024)
Does Single-channel Speech Enhancement Improve Keyword Spotting Accuracy? A Case Study
von: Brueggeman, Avamarie, et al.
Veröffentlicht: (2023)
von: Brueggeman, Avamarie, et al.
Veröffentlicht: (2023)
ChipChat: Low-Latency Cascaded Conversational Agent in MLX
von: Likhomanenko, Tatiana, et al.
Veröffentlicht: (2025)
von: Likhomanenko, Tatiana, et al.
Veröffentlicht: (2025)
Context-Driven Dynamic Pruning for Large Speech Foundation Models
von: Someki, Masao, et al.
Veröffentlicht: (2025)
von: Someki, Masao, et al.
Veröffentlicht: (2025)
Two-level adiabatic transition probability for small avoided crossings generated by tangential intersections
von: Higuchi, Kenta, et al.
Veröffentlicht: (2024)
von: Higuchi, Kenta, et al.
Veröffentlicht: (2024)
Neural Blind Source Separation and Diarization for Distant Speech Recognition
von: Bando, Yoshiaki, et al.
Veröffentlicht: (2024)
von: Bando, Yoshiaki, et al.
Veröffentlicht: (2024)
SpeakStream: Streaming Text-to-Speech with Interleaved Data
von: Bai, Richard He, et al.
Veröffentlicht: (2025)
von: Bai, Richard He, et al.
Veröffentlicht: (2025)
Multichannel Voice Trigger Detection Based on Transform-average-concatenate
von: Higuchi, Takuya, et al.
Veröffentlicht: (2023)
von: Higuchi, Takuya, et al.
Veröffentlicht: (2023)
Generating Data with Text-to-Speech and Large-Language Models for Conversational Speech Recognition
von: Cornell, Samuele, et al.
Veröffentlicht: (2024)
von: Cornell, Samuele, et al.
Veröffentlicht: (2024)
Useful Blunders: Can Automated Speech Recognition Errors Improve Downstream Dementia Classification?
von: Li, Changye, et al.
Veröffentlicht: (2024)
von: Li, Changye, et al.
Veröffentlicht: (2024)
SEED: Speaker Embedding Enhancement Diffusion Model
von: Nam, KiHyun, et al.
Veröffentlicht: (2025)
von: Nam, KiHyun, et al.
Veröffentlicht: (2025)
Proxy-FDA: Proxy-based Feature Distribution Alignment for Fine-tuning Vision Foundation Models without Forgetting
von: Huang, Chen, et al.
Veröffentlicht: (2025)
von: Huang, Chen, et al.
Veröffentlicht: (2025)
Chain-of-Thought Training for Open E2E Spoken Dialogue Systems
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Speaker-IPL: Unsupervised Learning of Speaker Characteristics with i-Vector based Pseudo-Labels
von: Aldeneh, Zakaria, et al.
Veröffentlicht: (2024) -
Which Data Matter? Embedding-Based Data Selection for Speech Recognition
von: Aldeneh, Zakaria, et al.
Veröffentlicht: (2026) -
Towards Automatic Assessment of Self-Supervised Speech Models using Rank
von: Aldeneh, Zakaria, et al.
Veröffentlicht: (2024) -
ESPnet-SPK: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models
von: Jung, Jee-weon, et al.
Veröffentlicht: (2024) -
Exploring Prediction Targets in Masked Pre-Training for Speech Foundation Models
von: Chen, Li-Wei, et al.
Veröffentlicht: (2024)