Release of Pre-Trained Models for the Japanese Language
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sawada, Kei, Zhao, Tianyu, Shing, Makoto, Mitsui, Kentaro, Kaga, Akio, Hono, Yukiya, Wakatsuki, Toshiaki, Mitsuda, Koh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Integrating Pre-Trained Speech and Language Models for End-to-End Speech Recognition
von: Hono, Yukiya, et al.
Veröffentlicht: (2023)
von: Hono, Yukiya, et al.
Veröffentlicht: (2023)
PSLM: Parallel Generation of Text and Speech with LLMs for Low-Latency Spoken Dialogue Systems
von: Mitsui, Kentaro, et al.
Veröffentlicht: (2024)
von: Mitsui, Kentaro, et al.
Veröffentlicht: (2024)
PeriodGrad: Towards Pitch-Controllable Neural Vocoder Based on a Diffusion Probabilistic Model
von: Hono, Yukiya, et al.
Veröffentlicht: (2024)
von: Hono, Yukiya, et al.
Veröffentlicht: (2024)
J-CHAT: Japanese Large-scale Spoken Dialogue Corpus for Spoken Dialogue Language Modeling
von: Nakata, Wataru, et al.
Veröffentlicht: (2024)
von: Nakata, Wataru, et al.
Veröffentlicht: (2024)
Findings of the 2023 ML-SUPERB Challenge: Pre-Training and Evaluation over More Languages and Beyond
von: Shi, Jiatong, et al.
Veröffentlicht: (2023)
von: Shi, Jiatong, et al.
Veröffentlicht: (2023)
From Statistical Methods to Pre-Trained Models; A Survey on Automatic Speech Recognition for Resource Scarce Urdu Language
von: Sharif, Muhammad, et al.
Veröffentlicht: (2024)
von: Sharif, Muhammad, et al.
Veröffentlicht: (2024)
Towards a Japanese Full-duplex Spoken Dialogue System
von: Ohashi, Atsumoto, et al.
Veröffentlicht: (2025)
von: Ohashi, Atsumoto, et al.
Veröffentlicht: (2025)
Self-Noise Reduction for Capacitive Sensors via Photoelectric DC Servo: Application to Condenser Microphones
von: Obo, Hirotaka, et al.
Veröffentlicht: (2026)
von: Obo, Hirotaka, et al.
Veröffentlicht: (2026)
Multilingual Zero Resource Speech Recognition Base on Self-Supervise Pre-Trained Acoustic Models
von: Wang, Haoyu, et al.
Veröffentlicht: (2022)
von: Wang, Haoyu, et al.
Veröffentlicht: (2022)
MTLM: Incorporating Bidirectional Text Information to Enhance Language Model Training in Speech Recognition Systems
von: Meng, Qingliang, et al.
Veröffentlicht: (2025)
von: Meng, Qingliang, et al.
Veröffentlicht: (2025)
MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
von: Cheng, Ho Kei, et al.
Veröffentlicht: (2024)
von: Cheng, Ho Kei, et al.
Veröffentlicht: (2024)
Enhancing Few-shot Keyword Spotting Performance through Pre-Trained Self-supervised Speech Models
von: Gok, Alican, et al.
Veröffentlicht: (2025)
von: Gok, Alican, et al.
Veröffentlicht: (2025)
Generative Pre-trained Speech Language Model with Efficient Hierarchical Transformer
von: Zhu, Yongxin, et al.
Veröffentlicht: (2024)
von: Zhu, Yongxin, et al.
Veröffentlicht: (2024)
BLSP-KD: Bootstrapping Language-Speech Pre-training via Knowledge Distillation
von: Wang, Chen, et al.
Veröffentlicht: (2024)
von: Wang, Chen, et al.
Veröffentlicht: (2024)
Building Tailored Speech Recognizers for Japanese Speaking Assessment
von: Kubo, Yotaro, et al.
Veröffentlicht: (2025)
von: Kubo, Yotaro, et al.
Veröffentlicht: (2025)
GenDistiller: Distilling Pre-trained Language Models based on an Autoregressive Generative Model
von: Gao, Yingying, et al.
Veröffentlicht: (2024)
von: Gao, Yingying, et al.
Veröffentlicht: (2024)
BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing
von: Wang, Chen, et al.
Veröffentlicht: (2023)
von: Wang, Chen, et al.
Veröffentlicht: (2023)
CMT-LLM: Contextual Multi-Talker ASR Utilizing Large Language Models
von: He, Jiajun, et al.
Veröffentlicht: (2025)
von: He, Jiajun, et al.
Veröffentlicht: (2025)
Discovering and Causally Validating Emotion-Sensitive Neurons in Large Audio-Language Models
von: Zhao, Xiutian, et al.
Veröffentlicht: (2026)
von: Zhao, Xiutian, et al.
Veröffentlicht: (2026)
Memory-Efficient Training for Text-Dependent SV with Independent Pre-trained Models
von: Farokh, Seyed Ali, et al.
Veröffentlicht: (2024)
von: Farokh, Seyed Ali, et al.
Veröffentlicht: (2024)
NEST-RQ: Next Token Prediction for Speech Self-Supervised Pre-Training
von: Han, Minglun, et al.
Veröffentlicht: (2024)
von: Han, Minglun, et al.
Veröffentlicht: (2024)
Transcript-Prompted Whisper with Dictionary-Enhanced Decoding for Japanese Speech Annotation
von: Hu, Rui, et al.
Veröffentlicht: (2025)
von: Hu, Rui, et al.
Veröffentlicht: (2025)
Benchmarking Prosody Encoding in Discrete Speech Tokens
von: Onda, Kentaro, et al.
Veröffentlicht: (2025)
von: Onda, Kentaro, et al.
Veröffentlicht: (2025)
A Pilot Study of GSLM-based Simulation of Foreign Accentuation Only Using Native Speech Corpora
von: Onda, Kentaro, et al.
Veröffentlicht: (2024)
von: Onda, Kentaro, et al.
Veröffentlicht: (2024)
Speechless: Speech Instruction Training Without Speech for Low Resource Languages
von: Dao, Alan, et al.
Veröffentlicht: (2025)
von: Dao, Alan, et al.
Veröffentlicht: (2025)
How to Estimate Model Transferability of Pre-Trained Speech Models?
von: Chen, Zih-Ching, et al.
Veröffentlicht: (2023)
von: Chen, Zih-Ching, et al.
Veröffentlicht: (2023)
Neuron-Level Emotion Control in Speech-Generative Large Audio-Language Models
von: Zhao, Xiutian, et al.
Veröffentlicht: (2026)
von: Zhao, Xiutian, et al.
Veröffentlicht: (2026)
Training and Inference Efficiency of Encoder-Decoder Speech Models
von: Żelasko, Piotr, et al.
Veröffentlicht: (2025)
von: Żelasko, Piotr, et al.
Veröffentlicht: (2025)
Low-Resourced Speech Recognition for Iu Mien Language via Weakly-Supervised Phoneme-based Multilingual Pre-training
von: Dong, Lukuan, et al.
Veröffentlicht: (2024)
von: Dong, Lukuan, et al.
Veröffentlicht: (2024)
Comparative Evaluation of Expressive Japanese Character Text-to-Speech with VITS and Style-BERT-VITS2
von: Rackauckas, Zackary, et al.
Veröffentlicht: (2025)
von: Rackauckas, Zackary, et al.
Veröffentlicht: (2025)
ProLAP: Probabilistic Language-Audio Pre-Training
von: Manabe, Toranosuke, et al.
Veröffentlicht: (2025)
von: Manabe, Toranosuke, et al.
Veröffentlicht: (2025)
Training-Free Deepfake Voice Recognition by Leveraging Large-Scale Pre-Trained Models
von: Pianese, Alessandro, et al.
Veröffentlicht: (2024)
von: Pianese, Alessandro, et al.
Veröffentlicht: (2024)
Toward Corpus Size Requirements for Training and Evaluating Depression Risk Models Using Spoken Language
von: Rutowski, Tomek, et al.
Veröffentlicht: (2024)
von: Rutowski, Tomek, et al.
Veröffentlicht: (2024)
Longer is (Not Necessarily) Stronger: Punctuated Long-Sequence Training for Enhanced Speech Recognition and Translation
von: Koluguri, Nithin Rao, et al.
Veröffentlicht: (2024)
von: Koluguri, Nithin Rao, et al.
Veröffentlicht: (2024)
Benchmarking Japanese Speech Recognition on ASR-LLM Setups with Multi-Pass Augmented Generative Error Correction
von: Ko, Yuka, et al.
Veröffentlicht: (2024)
von: Ko, Yuka, et al.
Veröffentlicht: (2024)
Fine-grained Speech Sentiment Analysis in Chinese Psychological Support Hotlines Based on Large-scale Pre-trained Model
von: Chen, Zhonglong, et al.
Veröffentlicht: (2024)
von: Chen, Zhonglong, et al.
Veröffentlicht: (2024)
Teaching a Multilingual Large Language Model to Understand Multilingual Speech via Multi-Instructional Training
von: Denisov, Pavel, et al.
Veröffentlicht: (2024)
von: Denisov, Pavel, et al.
Veröffentlicht: (2024)
Enhancing Multilingual ASR for Unseen Languages via Language Embedding Modeling
von: Huang, Shao-Syuan, et al.
Veröffentlicht: (2024)
von: Huang, Shao-Syuan, et al.
Veröffentlicht: (2024)
MINT: Boosting Audio-Language Model via Multi-Target Pre-Training and Instruction Tuning
von: Zhao, Hang, et al.
Veröffentlicht: (2024)
von: Zhao, Hang, et al.
Veröffentlicht: (2024)
An Empirical Analysis of Discrete Unit Representations in Speech Language Modeling Pre-training
von: Labrak, Yanis, et al.
Veröffentlicht: (2025)
von: Labrak, Yanis, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Integrating Pre-Trained Speech and Language Models for End-to-End Speech Recognition
von: Hono, Yukiya, et al.
Veröffentlicht: (2023) -
PSLM: Parallel Generation of Text and Speech with LLMs for Low-Latency Spoken Dialogue Systems
von: Mitsui, Kentaro, et al.
Veröffentlicht: (2024) -
PeriodGrad: Towards Pitch-Controllable Neural Vocoder Based on a Diffusion Probabilistic Model
von: Hono, Yukiya, et al.
Veröffentlicht: (2024) -
J-CHAT: Japanese Large-scale Spoken Dialogue Corpus for Spoken Dialogue Language Modeling
von: Nakata, Wataru, et al.
Veröffentlicht: (2024) -
Findings of the 2023 ML-SUPERB Challenge: Pre-Training and Evaluation over More Languages and Beyond
von: Shi, Jiatong, et al.
Veröffentlicht: (2023)