ESPnet-EZ: Python-only ESPnet for Easy Fine-tuning and Integration
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Someki, Masao, Choi, Kwanghee, Arora, Siddhant, Chen, William, Cornell, Samuele, Han, Jionghao, Peng, Yifan, Shi, Jiatong, Srivastav, Vaibhav, Watanabe, Shinji |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ESPnet-SDS: Unified Toolkit and Demo for Spoken Dialogue Systems
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
ESPnet-SpeechLM: An Open Speech Language Model Toolkit
von: Tian, Jinchuan, et al.
Veröffentlicht: (2025)
von: Tian, Jinchuan, et al.
Veröffentlicht: (2025)
Muskits-ESPnet: A Comprehensive Toolkit for Singing Voice Synthesis in New Paradigm
von: Wu, Yuning, et al.
Veröffentlicht: (2024)
von: Wu, Yuning, et al.
Veröffentlicht: (2024)
On-device Streaming Discrete Speech Units
von: Choi, Kwanghee, et al.
Veröffentlicht: (2025)
von: Choi, Kwanghee, et al.
Veröffentlicht: (2025)
ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
ESPnet-SPK: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models
von: Jung, Jee-weon, et al.
Veröffentlicht: (2024)
von: Jung, Jee-weon, et al.
Veröffentlicht: (2024)
Segment-Level Vectorized Beam Search Based on Partially Autoregressive Inference
von: Someki, Masao, et al.
Veröffentlicht: (2023)
von: Someki, Masao, et al.
Veröffentlicht: (2023)
OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder
von: Bharadwaj, Shikhar, et al.
Veröffentlicht: (2025)
von: Bharadwaj, Shikhar, et al.
Veröffentlicht: (2025)
The CMU-AIST submission for the ICME 2025 Audio Encoder Challenge
von: Bharadwaj, Shikhar, et al.
Veröffentlicht: (2026)
von: Bharadwaj, Shikhar, et al.
Veröffentlicht: (2026)
MAPSS: Manifold-based Assessment of Perceptual Source Separation
von: Ivry, Amir, et al.
Veröffentlicht: (2025)
von: Ivry, Amir, et al.
Veröffentlicht: (2025)
ARECHO: Autoregressive Evaluation via Chain-Based Hypothesis Optimization for Speech Multi-Metric Estimation
von: Shi, Jiatong, et al.
Veröffentlicht: (2025)
von: Shi, Jiatong, et al.
Veröffentlicht: (2025)
Decoder-only Architecture for Speech Recognition with CTC Prompts and Text Data Augmentation
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2023)
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2023)
Diffusion-based Generative Modeling with Discriminative Guidance for Streamable Speech Enhancement
von: Li, Chenda, et al.
Veröffentlicht: (2024)
von: Li, Chenda, et al.
Veröffentlicht: (2024)
Generating Data with Text-to-Speech and Large-Language Models for Conversational Speech Recognition
von: Cornell, Samuele, et al.
Veröffentlicht: (2024)
von: Cornell, Samuele, et al.
Veröffentlicht: (2024)
Chain-of-Thought Reasoning in Streaming Full-Duplex End-to-End Spoken Dialogue Systems
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
On the Effects of Heterogeneous Data Sources on Speech-to-Text Foundation Models
von: Tian, Jinchuan, et al.
Veröffentlicht: (2024)
von: Tian, Jinchuan, et al.
Veröffentlicht: (2024)
Semi-Autoregressive Streaming ASR With Label Context
von: Arora, Siddhant, et al.
Veröffentlicht: (2023)
von: Arora, Siddhant, et al.
Veröffentlicht: (2023)
Uni-VERSA: Versatile Speech Assessment with a Unified Network
von: Shi, Jiatong, et al.
Veröffentlicht: (2025)
von: Shi, Jiatong, et al.
Veröffentlicht: (2025)
Cross-Talk Speech Reduction, by Separation, for Separation
von: Wang, Zhong-Qiu, et al.
Veröffentlicht: (2026)
von: Wang, Zhong-Qiu, et al.
Veröffentlicht: (2026)
Improving Multilingual Speech Models on ML-SUPERB 2.0: Fine-tuning with Data Augmentation and LID-Aware CTC
von: Wang, Qingzheng, et al.
Veröffentlicht: (2025)
von: Wang, Qingzheng, et al.
Veröffentlicht: (2025)
EFFUSE: Efficient Self-Supervised Feature Fusion for E2E ASR in Low Resource and Multilingual Scenarios
von: Srivastava, Tejes, et al.
Veröffentlicht: (2023)
von: Srivastava, Tejes, et al.
Veröffentlicht: (2023)
EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion
von: Joglekar, Advait, et al.
Veröffentlicht: (2025)
von: Joglekar, Advait, et al.
Veröffentlicht: (2025)
Spiralformer: Low Latency Encoder for Streaming Speech Recognition with Circular Layer Skipping and Early Exiting
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2025)
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2025)
Rapid Language Adaptation for Multilingual E2E Speech Recognition Using Encoder Prompting
von: Kashiwagi, Yosuke, et al.
Veröffentlicht: (2024)
von: Kashiwagi, Yosuke, et al.
Veröffentlicht: (2024)
Hypothesis Clustering and Merging: Novel MultiTalker Speech Recognition with Speaker Tokens
von: Kashiwagi, Yosuke, et al.
Veröffentlicht: (2024)
von: Kashiwagi, Yosuke, et al.
Veröffentlicht: (2024)
Chain-of-Thought Training for Open E2E Spoken Dialogue Systems
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
The CHiME-8 DASR Challenge for Generalizable and Array Agnostic Distant Automatic Speech Recognition and Diarization
von: Cornell, Samuele, et al.
Veröffentlicht: (2024)
von: Cornell, Samuele, et al.
Veröffentlicht: (2024)
Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
MMM: Multi-Layer Multi-Residual Multi-Stream Discrete Speech Representation from Self-supervised Learning Model
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
VISinger2+: End-to-End Singing Voice Synthesis Augmented by Self-Supervised Learning Representation
von: Yu, Yifeng, et al.
Veröffentlicht: (2024)
von: Yu, Yifeng, et al.
Veröffentlicht: (2024)
Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks
von: Maiti, Soumi, et al.
Veröffentlicht: (2023)
von: Maiti, Soumi, et al.
Veröffentlicht: (2023)
OpusLM: A Family of Open Unified Speech Language Models
von: Tian, Jinchuan, et al.
Veröffentlicht: (2025)
von: Tian, Jinchuan, et al.
Veröffentlicht: (2025)
Improving Speech Enhancement with Multi-Metric Supervision from Learned Quality Assessment
von: Wang, Wei, et al.
Veröffentlicht: (2025)
von: Wang, Wei, et al.
Veröffentlicht: (2025)
ICASSP 2026 URGENT Speech Enhancement Challenge
von: Li, Chenda, et al.
Veröffentlicht: (2026)
von: Li, Chenda, et al.
Veröffentlicht: (2026)
Speech-DRAME: A Framework for Human-Aligned Benchmarks in Speech Role-Play
von: Shi, Jiatong, et al.
Veröffentlicht: (2025)
von: Shi, Jiatong, et al.
Veröffentlicht: (2025)
Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition
von: Wang, Shih-heng, et al.
Veröffentlicht: (2024)
von: Wang, Shih-heng, et al.
Veröffentlicht: (2024)
VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
von: Futami, Hayato, et al.
Veröffentlicht: (2025)
von: Futami, Hayato, et al.
Veröffentlicht: (2025)
URGENT Challenge: Universality, Robustness, and Generalizability For Speech Enhancement
von: Zhang, Wangyou, et al.
Veröffentlicht: (2024)
von: Zhang, Wangyou, et al.
Veröffentlicht: (2024)
PURE Codec: Progressive Unfolding of Residual Entropy for Speech Codec Learning
von: Shi, Jiatong, et al.
Veröffentlicht: (2025)
von: Shi, Jiatong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ESPnet-SDS: Unified Toolkit and Demo for Spoken Dialogue Systems
von: Arora, Siddhant, et al.
Veröffentlicht: (2025) -
ESPnet-SpeechLM: An Open Speech Language Model Toolkit
von: Tian, Jinchuan, et al.
Veröffentlicht: (2025) -
Muskits-ESPnet: A Comprehensive Toolkit for Singing Voice Synthesis in New Paradigm
von: Wu, Yuning, et al.
Veröffentlicht: (2024) -
On-device Streaming Discrete Speech Units
von: Choi, Kwanghee, et al.
Veröffentlicht: (2025) -
ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)