ESPnet-SPK: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models
Fuente:
arXiv
Saved in:
| Main Authors: | Jung, Jee-weon, Zhang, Wangyou, Shi, Jiatong, Aldeneh, Zakaria, Higuchi, Takuya, Theobald, Barry-John, Abdelaziz, Ahmed Hussen, Watanabe, Shinji |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can you Remove the Downstream Model for Speaker Recognition with Self-Supervised Speech Features?
by: Aldeneh, Zakaria, et al.
Published: (2024)
by: Aldeneh, Zakaria, et al.
Published: (2024)
Speaker-IPL: Unsupervised Learning of Speaker Characteristics with i-Vector based Pseudo-Labels
by: Aldeneh, Zakaria, et al.
Published: (2024)
by: Aldeneh, Zakaria, et al.
Published: (2024)
Exploring Prediction Targets in Masked Pre-Training for Speech Foundation Models
by: Chen, Li-Wei, et al.
Published: (2024)
by: Chen, Li-Wei, et al.
Published: (2024)
A Variational Framework for Improving Naturalness in Generative Spoken Language Models
by: Chen, Li-Wei, et al.
Published: (2025)
by: Chen, Li-Wei, et al.
Published: (2025)
Towards Automatic Assessment of Self-Supervised Speech Models using Rank
by: Aldeneh, Zakaria, et al.
Published: (2024)
by: Aldeneh, Zakaria, et al.
Published: (2024)
Which Data Matter? Embedding-Based Data Selection for Speech Recognition
by: Aldeneh, Zakaria, et al.
Published: (2026)
by: Aldeneh, Zakaria, et al.
Published: (2026)
Improving Design of Input Condition Invariant Speech Enhancement
by: Zhang, Wangyou, et al.
Published: (2024)
by: Zhang, Wangyou, et al.
Published: (2024)
DiceHuBERT: Distilling HuBERT with a Self-Supervised Learning Objective
by: Chi, Hyung Gun, et al.
Published: (2025)
by: Chi, Hyung Gun, et al.
Published: (2025)
ESPnet-EZ: Python-only ESPnet for Easy Fine-tuning and Integration
by: Someki, Masao, et al.
Published: (2024)
by: Someki, Masao, et al.
Published: (2024)
Curriculum learning for self-supervised speaker verification
by: Heo, Hee-Soo, et al.
Published: (2022)
by: Heo, Hee-Soo, et al.
Published: (2022)
Beyond Performance Plateaus: A Comprehensive Study on Scalability in Speech Enhancement
by: Zhang, Wangyou, et al.
Published: (2024)
by: Zhang, Wangyou, et al.
Published: (2024)
ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech
by: Shi, Jiatong, et al.
Published: (2024)
by: Shi, Jiatong, et al.
Published: (2024)
ESPnet-SDS: Unified Toolkit and Demo for Spoken Dialogue Systems
by: Arora, Siddhant, et al.
Published: (2025)
by: Arora, Siddhant, et al.
Published: (2025)
Muskits-ESPnet: A Comprehensive Toolkit for Singing Voice Synthesis in New Paradigm
by: Wu, Yuning, et al.
Published: (2024)
by: Wu, Yuning, et al.
Published: (2024)
Learning Spatially-Aware Language and Audio Embeddings
by: Devnani, Bhavika, et al.
Published: (2024)
by: Devnani, Bhavika, et al.
Published: (2024)
ESPnet-SpeechLM: An Open Speech Language Model Toolkit
by: Tian, Jinchuan, et al.
Published: (2025)
by: Tian, Jinchuan, et al.
Published: (2025)
a-DCF: an architecture agnostic metric with application to spoofing-robust speaker verification
by: Shim, Hye-jin, et al.
Published: (2024)
by: Shim, Hye-jin, et al.
Published: (2024)
Speech Foundation Models Generalize to Time Series Tasks from Wearable Sensor Data
by: Narain, Jaya, et al.
Published: (2025)
by: Narain, Jaya, et al.
Published: (2025)
Chain-of-Thought Training for Open E2E Spoken Dialogue Systems
by: Arora, Siddhant, et al.
Published: (2025)
by: Arora, Siddhant, et al.
Published: (2025)
AugSumm: towards generalizable speech summarization using synthetic labels from large language model
by: Jung, Jee-weon, et al.
Published: (2024)
by: Jung, Jee-weon, et al.
Published: (2024)
Improving Speech Enhancement with Multi-Metric Supervision from Learned Quality Assessment
by: Wang, Wei, et al.
Published: (2025)
by: Wang, Wei, et al.
Published: (2025)
Ants3 toolkit: front-end for Geant4 with interactive GUI and Python scripting
by: Morozov, A., et al.
Published: (2025)
by: Morozov, A., et al.
Published: (2025)
Beyond Silence: Bias Analysis through Loss and Asymmetric Approach in Audio Anti-Spoofing
by: Shim, Hye-jin, et al.
Published: (2024)
by: Shim, Hye-jin, et al.
Published: (2024)
Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks
by: Maiti, Soumi, et al.
Published: (2023)
by: Maiti, Soumi, et al.
Published: (2023)
PURE Codec: Progressive Unfolding of Residual Entropy for Speech Codec Learning
by: Shi, Jiatong, et al.
Published: (2025)
by: Shi, Jiatong, et al.
Published: (2025)
Text adaptation for speaker verification with speaker-text factorized embeddings
by: Yang, Yexin, et al.
Published: (2025)
by: Yang, Yexin, et al.
Published: (2025)
Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet
by: Ying, Anyu, et al.
Published: (2025)
by: Ying, Anyu, et al.
Published: (2025)
An approach to optimize inference of the DIART speaker diarization pipeline
by: Aperdannier, Roman, et al.
Published: (2024)
by: Aperdannier, Roman, et al.
Published: (2024)
Hubble tension environment correction (reproducible pipeline)
by: Baiyi Wang
Published: (2026)
by: Baiyi Wang
Published: (2026)
Hubble tension environment correction (reproducible pipeline)
by: Baiyi Wang
Published: (2026)
by: Baiyi Wang
Published: (2026)
Leveraging Audio-Visual Data to Reduce the Multilingual Gap in Self-Supervised Speech Models
by: Blandón, María Andrea Cruz, et al.
Published: (2025)
by: Blandón, María Andrea Cruz, et al.
Published: (2025)
MMM: Multi-Layer Multi-Residual Multi-Stream Discrete Speech Representation from Self-supervised Learning Model
by: Shi, Jiatong, et al.
Published: (2024)
by: Shi, Jiatong, et al.
Published: (2024)
OWSM v3.1: Better and Faster Open Whisper-Style Speech Models based on E-Branchformer
by: Peng, Yifan, et al.
Published: (2024)
by: Peng, Yifan, et al.
Published: (2024)
Improving Audio Captioning Models with Fine-grained Audio Features, Text Embedding Supervision, and LLM Mix-up Augmentation
by: Wu, Shih-Lun, et al.
Published: (2023)
by: Wu, Shih-Lun, et al.
Published: (2023)
Comfortability of quantum walks on embedded graphs on surfaces
by: Higuchi, Yusuke, et al.
Published: (2025)
by: Higuchi, Yusuke, et al.
Published: (2025)
Quantum walks on graphs embedded in orientable surfaces
by: Higuchi, Yusuke, et al.
Published: (2024)
by: Higuchi, Yusuke, et al.
Published: (2024)
Investigation of perception inconsistency in speaker embedding for asynchronous voice anonymization
by: Wang, Rui, et al.
Published: (2025)
by: Wang, Rui, et al.
Published: (2025)
A reproducible pipeline for activity-based travel demand generation in England
by: Anonymous, Anonymous
Published: (2025)
by: Anonymous, Anonymous
Published: (2025)
A reproducible pipeline for extracting representative signals from wire cuts
by: Lin, Yuhang, et al.
Published: (2024)
by: Lin, Yuhang, et al.
Published: (2024)
TMT: Tri-Modal Translation between Speech, Image, and Text by Processing Different Modalities as Different Languages
by: Kim, Minsu, et al.
Published: (2024)
by: Kim, Minsu, et al.
Published: (2024)
Similar Items
-
Can you Remove the Downstream Model for Speaker Recognition with Self-Supervised Speech Features?
by: Aldeneh, Zakaria, et al.
Published: (2024) -
Speaker-IPL: Unsupervised Learning of Speaker Characteristics with i-Vector based Pseudo-Labels
by: Aldeneh, Zakaria, et al.
Published: (2024) -
Exploring Prediction Targets in Masked Pre-Training for Speech Foundation Models
by: Chen, Li-Wei, et al.
Published: (2024) -
A Variational Framework for Improving Naturalness in Generative Spoken Language Models
by: Chen, Li-Wei, et al.
Published: (2025) -
Towards Automatic Assessment of Self-Supervised Speech Models using Rank
by: Aldeneh, Zakaria, et al.
Published: (2024)