J-CHAT: Japanese Large-scale Spoken Dialogue Corpus for Spoken Dialogue Language Modeling
Fuente:
arXiv
Guardado en:
| Autores principales: | Nakata, Wataru, Seki, Kentaro, Yanaka, Hitomi, Saito, Yuki, Takamichi, Shinnosuke, Saruwatari, Hiroshi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
DialogueSidon: Recovering Full-Duplex Dialogue Tracks from In-the-Wild Dialogue Audio
por: Nakata, Wataru, et al.
Publicado: (2026)
por: Nakata, Wataru, et al.
Publicado: (2026)
SRC4VC: Smartphone-Recorded Corpus for Voice Conversion Benchmark
por: Saito, Yuki, et al.
Publicado: (2024)
por: Saito, Yuki, et al.
Publicado: (2024)
JVNV: A Corpus of Japanese Emotional Speech with Verbal Content and Nonverbal Expressions
por: Xin, Detai, et al.
Publicado: (2023)
por: Xin, Detai, et al.
Publicado: (2023)
Building speech corpus with diverse voice characteristics for its prompt-based representation
por: Watanabe, Aya, et al.
Publicado: (2024)
por: Watanabe, Aya, et al.
Publicado: (2024)
Active Learning for Text-to-Speech Synthesis with Informative Sample Collection
por: Seki, Kentaro, et al.
Publicado: (2025)
por: Seki, Kentaro, et al.
Publicado: (2025)
Spatial Voice Conversion: Voice Conversion Preserving Spatial Information and Non-target Signals
por: Seki, Kentaro, et al.
Publicado: (2024)
por: Seki, Kentaro, et al.
Publicado: (2024)
Noise-Robust Voice Conversion by Conditional Denoising Training Using Latent Variables of Recording Quality and Environment
por: Igarashi, Takuto, et al.
Publicado: (2024)
por: Igarashi, Takuto, et al.
Publicado: (2024)
Sidon: Fast and Robust Open-Source Multilingual Speech Restoration for Large-scale Dataset Cleansing
por: Nakata, Wataru, et al.
Publicado: (2025)
por: Nakata, Wataru, et al.
Publicado: (2025)
SaSLaW: Dialogue Speech Corpus with Audio-visual Egocentric Information Toward Environment-adaptive Dialogue Speech Synthesis
por: Take, Osamu, et al.
Publicado: (2024)
por: Take, Osamu, et al.
Publicado: (2024)
JaCappella Corpus: A Japanese a Cappella Vocal Ensemble Corpus
por: Nakamura, Tomohiko, et al.
Publicado: (2022)
por: Nakamura, Tomohiko, et al.
Publicado: (2022)
Geneses: Unified Generative Speech Enhancement and Separation
por: Asai, Kohei, et al.
Publicado: (2026)
por: Asai, Kohei, et al.
Publicado: (2026)
Causal Speech Enhancement with Predicting Semantics based on Quantized Self-supervised Learning Features
por: Tsunoo, Emiru, et al.
Publicado: (2024)
por: Tsunoo, Emiru, et al.
Publicado: (2024)
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio
por: Kanamori, Yusuke, et al.
Publicado: (2025)
por: Kanamori, Yusuke, et al.
Publicado: (2025)
BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
por: Xin, Detai, et al.
Publicado: (2024)
por: Xin, Detai, et al.
Publicado: (2024)
E-chat: Emotion-sensitive Spoken Dialogue System with Large Language Models
por: Xue, Hongfei, et al.
Publicado: (2023)
por: Xue, Hongfei, et al.
Publicado: (2023)
Multi-Sampling-Frequency Naturalness MOS Prediction Using Self-Supervised Learning Model with Sampling-Frequency-Independent Layer
por: Nishikawa, Go, et al.
Publicado: (2025)
por: Nishikawa, Go, et al.
Publicado: (2025)
The T05 System for The VoiceMOS Challenge 2024: Transfer Learning from Deep Image Classifier to Naturalness MOS Prediction of High-Quality Synthetic Speech
por: Baba, Kaito, et al.
Publicado: (2024)
por: Baba, Kaito, et al.
Publicado: (2024)
DNN-based ensemble singing voice synthesis with interactions between singers
por: Hyodo, Hiroaki, et al.
Publicado: (2024)
por: Hyodo, Hiroaki, et al.
Publicado: (2024)
SpeechBERTScore: Reference-Aware Automatic Evaluation of Speech Generation Leveraging NLP Evaluation Metrics
por: Saeki, Takaaki, et al.
Publicado: (2024)
por: Saeki, Takaaki, et al.
Publicado: (2024)
Drum-to-Vocal Percussion Sound Conversion and Its Evaluation Methodology
por: Nobukawa, Rinka, et al.
Publicado: (2025)
por: Nobukawa, Rinka, et al.
Publicado: (2025)
LALM-as-a-Judge: Benchmarking Large Audio-Language Models for Safety Evaluation in Multi-Turn Spoken Dialogues
por: Ivry, Amir, et al.
Publicado: (2026)
por: Ivry, Amir, et al.
Publicado: (2026)
Semantic-Aware Interruption Detection in Spoken Dialogue Systems: Benchmark, Metric, and Model
por: Xia, Kangxiang, et al.
Publicado: (2026)
por: Xia, Kangxiang, et al.
Publicado: (2026)
DeepDialogue: A Multi-Turn Emotionally-Rich Spoken Dialogue Dataset
por: Koudounas, Alkis, et al.
Publicado: (2025)
por: Koudounas, Alkis, et al.
Publicado: (2025)
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training
por: Udupa, Sathvik, et al.
Publicado: (2025)
por: Udupa, Sathvik, et al.
Publicado: (2025)
ESPnet-SDS: Unified Toolkit and Demo for Spoken Dialogue Systems
por: Arora, Siddhant, et al.
Publicado: (2025)
por: Arora, Siddhant, et al.
Publicado: (2025)
SLIDE: Integrating Speech Language Model with LLM for Spontaneous Spoken Dialogue Generation
por: Lu, Haitian, et al.
Publicado: (2025)
por: Lu, Haitian, et al.
Publicado: (2025)
Cross-Dialect Text-To-Speech in Pitch-Accent Language Incorporating Multi-Dialect Phoneme-Level BERT
por: Yamauchi, Kazuki, et al.
Publicado: (2024)
por: Yamauchi, Kazuki, et al.
Publicado: (2024)
From Turn-Taking to Synchronous Dialogue: A Survey of Full-Duplex Spoken Language Models
por: Chen, Yuxuan, et al.
Publicado: (2025)
por: Chen, Yuxuan, et al.
Publicado: (2025)
Chain-of-Thought Training for Open E2E Spoken Dialogue Systems
por: Arora, Siddhant, et al.
Publicado: (2025)
por: Arora, Siddhant, et al.
Publicado: (2025)
WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models
por: Chen, Yifu, et al.
Publicado: (2025)
por: Chen, Yifu, et al.
Publicado: (2025)
SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words
por: Ao, Junyi, et al.
Publicado: (2024)
por: Ao, Junyi, et al.
Publicado: (2024)
Stream RAG: Instant and Accurate Spoken Dialogue Systems with Streaming Tool Usage
por: Arora, Siddhant, et al.
Publicado: (2025)
por: Arora, Siddhant, et al.
Publicado: (2025)
Aligning Spoken Dialogue Models from User Interactions
por: Wu, Anne, et al.
Publicado: (2025)
por: Wu, Anne, et al.
Publicado: (2025)
Who Finds This Voice Attractive? A Large-Scale Experiment Using In-the-Wild Data
por: Suda, Hitoshi, et al.
Publicado: (2024)
por: Suda, Hitoshi, et al.
Publicado: (2024)
Chain-of-Thought Reasoning in Streaming Full-Duplex End-to-End Spoken Dialogue Systems
por: Arora, Siddhant, et al.
Publicado: (2025)
por: Arora, Siddhant, et al.
Publicado: (2025)
Human-CLAP: Human-perception-based contrastive language-audio pretraining
por: Takano, Taisei, et al.
Publicado: (2025)
por: Takano, Taisei, et al.
Publicado: (2025)
UTDUSS: UTokyo-SaruLab System for Interspeech2024 Speech Processing Using Discrete Speech Unit Challenge
por: Nakata, Wataru, et al.
Publicado: (2024)
por: Nakata, Wataru, et al.
Publicado: (2024)
Joint Speech and Text Training for LLM-Based End-to-End Spoken Dialogue State Tracking
por: Vendrame, Katia, et al.
Publicado: (2025)
por: Vendrame, Katia, et al.
Publicado: (2025)
Leveraging Chain of Thought towards Empathetic Spoken Dialogue without Corresponding Question-Answering Data
por: Xie, Jingran, et al.
Publicado: (2025)
por: Xie, Jingran, et al.
Publicado: (2025)
WavChat: A Survey of Spoken Dialogue Models
por: Ji, Shengpeng, et al.
Publicado: (2024)
por: Ji, Shengpeng, et al.
Publicado: (2024)
Ejemplares similares
-
DialogueSidon: Recovering Full-Duplex Dialogue Tracks from In-the-Wild Dialogue Audio
por: Nakata, Wataru, et al.
Publicado: (2026) -
SRC4VC: Smartphone-Recorded Corpus for Voice Conversion Benchmark
por: Saito, Yuki, et al.
Publicado: (2024) -
JVNV: A Corpus of Japanese Emotional Speech with Verbal Content and Nonverbal Expressions
por: Xin, Detai, et al.
Publicado: (2023) -
Building speech corpus with diverse voice characteristics for its prompt-based representation
por: Watanabe, Aya, et al.
Publicado: (2024) -
Active Learning for Text-to-Speech Synthesis with Informative Sample Collection
por: Seki, Kentaro, et al.
Publicado: (2025)