SaSLaW: Dialogue Speech Corpus with Audio-visual Egocentric Information Toward Environment-adaptive Dialogue Speech Synthesis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Take, Osamu, Takamichi, Shinnosuke, Seki, Kentaro, Bando, Yoshiaki, Saruwatari, Hiroshi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Active Learning for Text-to-Speech Synthesis with Informative Sample Collection
von: Seki, Kentaro, et al.
Veröffentlicht: (2025)
von: Seki, Kentaro, et al.
Veröffentlicht: (2025)
J-CHAT: Japanese Large-scale Spoken Dialogue Corpus for Spoken Dialogue Language Modeling
von: Nakata, Wataru, et al.
Veröffentlicht: (2024)
von: Nakata, Wataru, et al.
Veröffentlicht: (2024)
SRC4VC: Smartphone-Recorded Corpus for Voice Conversion Benchmark
von: Saito, Yuki, et al.
Veröffentlicht: (2024)
von: Saito, Yuki, et al.
Veröffentlicht: (2024)
JVNV: A Corpus of Japanese Emotional Speech with Verbal Content and Nonverbal Expressions
von: Xin, Detai, et al.
Veröffentlicht: (2023)
von: Xin, Detai, et al.
Veröffentlicht: (2023)
BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
von: Xin, Detai, et al.
Veröffentlicht: (2024)
von: Xin, Detai, et al.
Veröffentlicht: (2024)
Noise-Robust Voice Conversion by Conditional Denoising Training Using Latent Variables of Recording Quality and Environment
von: Igarashi, Takuto, et al.
Veröffentlicht: (2024)
von: Igarashi, Takuto, et al.
Veröffentlicht: (2024)
SpeechBERTScore: Reference-Aware Automatic Evaluation of Speech Generation Leveraging NLP Evaluation Metrics
von: Saeki, Takaaki, et al.
Veröffentlicht: (2024)
von: Saeki, Takaaki, et al.
Veröffentlicht: (2024)
Spatial Voice Conversion: Voice Conversion Preserving Spatial Information and Non-target Signals
von: Seki, Kentaro, et al.
Veröffentlicht: (2024)
von: Seki, Kentaro, et al.
Veröffentlicht: (2024)
JaCappella Corpus: A Japanese a Cappella Vocal Ensemble Corpus
von: Nakamura, Tomohiko, et al.
Veröffentlicht: (2022)
von: Nakamura, Tomohiko, et al.
Veröffentlicht: (2022)
DialogueSidon: Recovering Full-Duplex Dialogue Tracks from In-the-Wild Dialogue Audio
von: Nakata, Wataru, et al.
Veröffentlicht: (2026)
von: Nakata, Wataru, et al.
Veröffentlicht: (2026)
Voice Conversion for Likability Control via Automated Rating of Speech Synthesis Corpora
von: Suda, Hitoshi, et al.
Veröffentlicht: (2025)
von: Suda, Hitoshi, et al.
Veröffentlicht: (2025)
Analysing the Language of Neural Audio Codecs
von: Park, Joonyong, et al.
Veröffentlicht: (2025)
von: Park, Joonyong, et al.
Veröffentlicht: (2025)
Annotation-Free MIDI-to-Audio Synthesis via Concatenative Synthesis and Generative Refinement
von: Take, Osamu, et al.
Veröffentlicht: (2024)
von: Take, Osamu, et al.
Veröffentlicht: (2024)
DNN-based ensemble singing voice synthesis with interactions between singers
von: Hyodo, Hiroaki, et al.
Veröffentlicht: (2024)
von: Hyodo, Hiroaki, et al.
Veröffentlicht: (2024)
Drum-to-Vocal Percussion Sound Conversion and Its Evaluation Methodology
von: Nobukawa, Rinka, et al.
Veröffentlicht: (2025)
von: Nobukawa, Rinka, et al.
Veröffentlicht: (2025)
Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models
von: Kando, Shunsuke, et al.
Veröffentlicht: (2025)
von: Kando, Shunsuke, et al.
Veröffentlicht: (2025)
YODAS: Youtube-Oriented Dataset for Audio and Speech
von: Li, Xinjian, et al.
Veröffentlicht: (2024)
von: Li, Xinjian, et al.
Veröffentlicht: (2024)
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio
von: Kanamori, Yusuke, et al.
Veröffentlicht: (2025)
von: Kanamori, Yusuke, et al.
Veröffentlicht: (2025)
Building speech corpus with diverse voice characteristics for its prompt-based representation
von: Watanabe, Aya, et al.
Veröffentlicht: (2024)
von: Watanabe, Aya, et al.
Veröffentlicht: (2024)
Neural Blind Source Separation and Diarization for Distant Speech Recognition
von: Bando, Yoshiaki, et al.
Veröffentlicht: (2024)
von: Bando, Yoshiaki, et al.
Veröffentlicht: (2024)
AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences
von: Kishi, Minoru, et al.
Veröffentlicht: (2025)
von: Kishi, Minoru, et al.
Veröffentlicht: (2025)
Geneses: Unified Generative Speech Enhancement and Separation
von: Asai, Kohei, et al.
Veröffentlicht: (2026)
von: Asai, Kohei, et al.
Veröffentlicht: (2026)
Shallow Flow Matching for Coarse-to-Fine Text-to-Speech Synthesis
von: Yang, Dong, et al.
Veröffentlicht: (2025)
von: Yang, Dong, et al.
Veröffentlicht: (2025)
SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue
von: Li, Ruiqi, et al.
Veröffentlicht: (2026)
von: Li, Ruiqi, et al.
Veröffentlicht: (2026)
Causal Speech Enhancement with Predicting Semantics based on Quantized Self-supervised Learning Features
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2024)
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2024)
RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis
von: Xin, Detai, et al.
Veröffentlicht: (2024)
von: Xin, Detai, et al.
Veröffentlicht: (2024)
Sidon: Fast and Robust Open-Source Multilingual Speech Restoration for Large-scale Dataset Cleansing
von: Nakata, Wataru, et al.
Veröffentlicht: (2025)
von: Nakata, Wataru, et al.
Veröffentlicht: (2025)
Input-Adaptive Spectral Feature Compression by Sequence Modeling for Source Separation
von: Saijo, Kohei, et al.
Veröffentlicht: (2026)
von: Saijo, Kohei, et al.
Veröffentlicht: (2026)
Subspace Track-before-Detect for Passive Multi-Target Tracking with Unknown Emitted Signals
von: Ito, Nobutaka, et al.
Veröffentlicht: (2026)
von: Ito, Nobutaka, et al.
Veröffentlicht: (2026)
Is MixIT Really Unsuitable for Correlated Sources? Exploring MixIT for Unsupervised Pre-training in Music Source Separation
von: Saijo, Kohei, et al.
Veröffentlicht: (2025)
von: Saijo, Kohei, et al.
Veröffentlicht: (2025)
Hyperbolic Embeddings for Order-Aware Classification of Audio Effect Chains
von: Wada, Aogu, et al.
Veröffentlicht: (2025)
von: Wada, Aogu, et al.
Veröffentlicht: (2025)
Cross-Dialect Text-To-Speech in Pitch-Accent Language Incorporating Multi-Dialect Phoneme-Level BERT
von: Yamauchi, Kazuki, et al.
Veröffentlicht: (2024)
von: Yamauchi, Kazuki, et al.
Veröffentlicht: (2024)
Text-To-Speech Synthesis In The Wild
von: Jung, Jee-weon, et al.
Veröffentlicht: (2024)
von: Jung, Jee-weon, et al.
Veröffentlicht: (2024)
DialoSpeech: Dual-Speaker Dialogue Generation with LLM and Flow Matching
von: Xie, Hanke, et al.
Veröffentlicht: (2025)
von: Xie, Hanke, et al.
Veröffentlicht: (2025)
DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
Who Finds This Voice Attractive? A Large-Scale Experiment Using In-the-Wild Data
von: Suda, Hitoshi, et al.
Veröffentlicht: (2024)
von: Suda, Hitoshi, et al.
Veröffentlicht: (2024)
Speaker-Conditioned Phrase Break Prediction for Text-to-Speech with Phoneme-Level Pre-trained Language Model
von: Yang, Dong, et al.
Veröffentlicht: (2025)
von: Yang, Dong, et al.
Veröffentlicht: (2025)
Audio Dialogues: Dialogues dataset for audio and music understanding
von: Goel, Arushi, et al.
Veröffentlicht: (2024)
von: Goel, Arushi, et al.
Veröffentlicht: (2024)
Real-time Speech Extraction Using Spatially Regularized Independent Low-rank Matrix Analysis and Rank-constrained Spatial Covariance Matrix Estimation
von: Ishikawa, Yuto, et al.
Veröffentlicht: (2024)
von: Ishikawa, Yuto, et al.
Veröffentlicht: (2024)
Localizing Acoustic Energy in Sound Field Synthesis by Directionally Weighted Exterior Radiation Suppression
von: Tomita, Yoshihide, et al.
Veröffentlicht: (2024)
von: Tomita, Yoshihide, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Active Learning for Text-to-Speech Synthesis with Informative Sample Collection
von: Seki, Kentaro, et al.
Veröffentlicht: (2025) -
J-CHAT: Japanese Large-scale Spoken Dialogue Corpus for Spoken Dialogue Language Modeling
von: Nakata, Wataru, et al.
Veröffentlicht: (2024) -
SRC4VC: Smartphone-Recorded Corpus for Voice Conversion Benchmark
von: Saito, Yuki, et al.
Veröffentlicht: (2024) -
JVNV: A Corpus of Japanese Emotional Speech with Verbal Content and Nonverbal Expressions
von: Xin, Detai, et al.
Veröffentlicht: (2023) -
BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
von: Xin, Detai, et al.
Veröffentlicht: (2024)