Construction and Analysis of Impression Caption Dataset for Environmental Sounds
Fuente:
arXiv
Guardado en:
| Autores principales: | Okamoto, Yuki, Nagase, Ryotaro, Okamoto, Minami, Saito, Yuki, Imoto, Keisuke, Fukumori, Takahiro, Yamashita, Yoichi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Human-CLAP: Human-perception-based contrastive language-audio pretraining
por: Takano, Taisei, et al.
Publicado: (2025)
por: Takano, Taisei, et al.
Publicado: (2025)
RISC: A Corpus for Shout Type Classification and Shout Intensity Prediction
por: Fukumori, Takahiro, et al.
Publicado: (2023)
por: Fukumori, Takahiro, et al.
Publicado: (2023)
Correlation of Fréchet Audio Distance With Human Perception of Environmental Audio Is Embedding Dependant
por: Tailleur, Modan, et al.
Publicado: (2024)
por: Tailleur, Modan, et al.
Publicado: (2024)
Color-based Emotion Representation for Speech Emotion Recognition
por: Nagase, Ryotaro, et al.
Publicado: (2026)
por: Nagase, Ryotaro, et al.
Publicado: (2026)
Joint Analysis of Acoustic Scenes and Sound Events Based on Semi-Supervised Training of Sound Events With Partial Labels
por: Imoto, Keisuke
Publicado: (2025)
por: Imoto, Keisuke
Publicado: (2025)
AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences
por: Kishi, Minoru, et al.
Publicado: (2025)
por: Kishi, Minoru, et al.
Publicado: (2025)
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio
por: Kanamori, Yusuke, et al.
Publicado: (2025)
por: Kanamori, Yusuke, et al.
Publicado: (2025)
Sound Scene Synthesis at the DCASE 2024 Challenge
por: Lagrange, Mathieu, et al.
Publicado: (2025)
por: Lagrange, Mathieu, et al.
Publicado: (2025)
LEAD Dataset: How Can Labels for Sound Event Detection Vary Depending on Annotators?
por: Koga, Naoki, et al.
Publicado: (2024)
por: Koga, Naoki, et al.
Publicado: (2024)
Trainingless Adaptation of Pretrained Models for Environmental Sound Classification
por: Tonami, Noriyuki, et al.
Publicado: (2024)
por: Tonami, Noriyuki, et al.
Publicado: (2024)
How Much Does Machine Identity Matter in Anomalous Sound Detection at Test Time?
por: Wilkinghoff, Kevin, et al.
Publicado: (2026)
por: Wilkinghoff, Kevin, et al.
Publicado: (2026)
Context-Aware Query Refinement for Target Sound Extraction: Handling Partially Matched Queries
por: Sato, Ryo, et al.
Publicado: (2025)
por: Sato, Ryo, et al.
Publicado: (2025)
Challenge on Sound Scene Synthesis: Evaluating Text-to-Audio Generation
por: Lee, Junwon, et al.
Publicado: (2024)
por: Lee, Junwon, et al.
Publicado: (2024)
Discrete Tokens Exhibit Interlanguage Speech Intelligibility Benefit: an Analytical Study Towards Accent-robust ASR Only with Native Speech Data
por: Onda, Kentaro, et al.
Publicado: (2025)
por: Onda, Kentaro, et al.
Publicado: (2025)
Prosodically Enhanced Foreign Accent Simulation by Discrete Token-based Resynthesis Only with Native Speech Corpora
por: Onda, Kentaro, et al.
Publicado: (2025)
por: Onda, Kentaro, et al.
Publicado: (2025)
LibriTTS-VI: A Public Corpus and Novel Methods for Efficient Voice Impression Control
por: Ohmura, Junki, et al.
Publicado: (2025)
por: Ohmura, Junki, et al.
Publicado: (2025)
Sidon: Fast and Robust Open-Source Multilingual Speech Restoration for Large-scale Dataset Cleansing
por: Nakata, Wataru, et al.
Publicado: (2025)
por: Nakata, Wataru, et al.
Publicado: (2025)
Handling Domain Shifts for Anomalous Sound Detection: A Review of DCASE-Related Work
por: Wilkinghoff, Kevin, et al.
Publicado: (2025)
por: Wilkinghoff, Kevin, et al.
Publicado: (2025)
Frame-Wise Breath Detection with Self-Training: An Exploration of Enhancing Breath Naturalness in Text-to-Speech
por: Yang, Dong, et al.
Publicado: (2024)
por: Yang, Dong, et al.
Publicado: (2024)
Layer-wise Analysis for Quality of Multilingual Synthesized Speech
por: Cooper, Erica, et al.
Publicado: (2025)
por: Cooper, Erica, et al.
Publicado: (2025)
Causal Speech Enhancement with Predicting Semantics based on Quantized Self-supervised Learning Features
por: Tsunoo, Emiru, et al.
Publicado: (2024)
por: Tsunoo, Emiru, et al.
Publicado: (2024)
Description and Discussion on DCASE 2026 Challenge Task 2: Noise-aware Unsupervised Anomalous Sound Detection for Machine Condition Monitoring
por: Nishida, Tomoya, et al.
Publicado: (2026)
por: Nishida, Tomoya, et al.
Publicado: (2026)
Geneses: Unified Generative Speech Enhancement and Separation
por: Asai, Kohei, et al.
Publicado: (2026)
por: Asai, Kohei, et al.
Publicado: (2026)
ACES: Evaluating Automated Audio Captioning Models on the Semantics of Sounds
por: Wijngaard, Gijs, et al.
Publicado: (2024)
por: Wijngaard, Gijs, et al.
Publicado: (2024)
Enhance Temporal Relations in Audio Captioning with Sound Event Detection
por: Xie, Zeyu, et al.
Publicado: (2023)
por: Xie, Zeyu, et al.
Publicado: (2023)
UTDUSS: UTokyo-SaruLab System for Interspeech2024 Speech Processing Using Discrete Speech Unit Challenge
por: Nakata, Wataru, et al.
Publicado: (2024)
por: Nakata, Wataru, et al.
Publicado: (2024)
DialogueSidon: Recovering Full-Duplex Dialogue Tracks from In-the-Wild Dialogue Audio
por: Nakata, Wataru, et al.
Publicado: (2026)
por: Nakata, Wataru, et al.
Publicado: (2026)
Zero- and Few-shot Sound Event Localization and Detection
por: Shimada, Kazuki, et al.
Publicado: (2023)
por: Shimada, Kazuki, et al.
Publicado: (2023)
Refining Knowledge Transfer on Audio-Image Temporal Agreement for Audio-Text Cross Retrieval
por: Tsubaki, Shunsuke, et al.
Publicado: (2024)
por: Tsubaki, Shunsuke, et al.
Publicado: (2024)
SONAR: Self-Distilled Continual Pre-training for Domain Adaptive Audio Representation
por: Zhang, Yizhou, et al.
Publicado: (2025)
por: Zhang, Yizhou, et al.
Publicado: (2025)
Description and Discussion on DCASE 2025 Challenge Task 2: First-shot Unsupervised Anomalous Sound Detection for Machine Condition Monitoring
por: Nishida, Tomoya, et al.
Publicado: (2025)
por: Nishida, Tomoya, et al.
Publicado: (2025)
Emotional Text-To-Speech Based on Mutual-Information-Guided Emotion-Timbre Disentanglement
por: Yang, Jianing, et al.
Publicado: (2025)
por: Yang, Jianing, et al.
Publicado: (2025)
SoundCTM: Unifying Score-based and Consistency Models for Full-band Text-to-Sound Generation
por: Saito, Koichi, et al.
Publicado: (2024)
por: Saito, Koichi, et al.
Publicado: (2024)
Spatial Voice Conversion: Voice Conversion Preserving Spatial Information and Non-target Signals
por: Seki, Kentaro, et al.
Publicado: (2024)
por: Seki, Kentaro, et al.
Publicado: (2024)
Building speech corpus with diverse voice characteristics for its prompt-based representation
por: Watanabe, Aya, et al.
Publicado: (2024)
por: Watanabe, Aya, et al.
Publicado: (2024)
Schrödinger Bridge Consistency Trajectory Models for Speech Enhancement
por: Nishigori, Shuichiro, et al.
Publicado: (2025)
por: Nishigori, Shuichiro, et al.
Publicado: (2025)
JVNV: A Corpus of Japanese Emotional Speech with Verbal Content and Nonverbal Expressions
por: Xin, Detai, et al.
Publicado: (2023)
por: Xin, Detai, et al.
Publicado: (2023)
Multi-Sampling-Frequency Naturalness MOS Prediction Using Self-Supervised Learning Model with Sampling-Frequency-Independent Layer
por: Nishikawa, Go, et al.
Publicado: (2025)
por: Nishikawa, Go, et al.
Publicado: (2025)
Retrieval-Augmented Approach for Unsupervised Anomalous Sound Detection and Captioning without Model Training
por: Ogura, Ryoya, et al.
Publicado: (2024)
por: Ogura, Ryoya, et al.
Publicado: (2024)
Retaining Mixture Representations for Domain Generalized Anomalous Sound Detection
por: Saengthong, Phurich, et al.
Publicado: (2025)
por: Saengthong, Phurich, et al.
Publicado: (2025)
Ejemplares similares
-
Human-CLAP: Human-perception-based contrastive language-audio pretraining
por: Takano, Taisei, et al.
Publicado: (2025) -
RISC: A Corpus for Shout Type Classification and Shout Intensity Prediction
por: Fukumori, Takahiro, et al.
Publicado: (2023) -
Correlation of Fréchet Audio Distance With Human Perception of Environmental Audio Is Embedding Dependant
por: Tailleur, Modan, et al.
Publicado: (2024) -
Color-based Emotion Representation for Speech Emotion Recognition
por: Nagase, Ryotaro, et al.
Publicado: (2026) -
Joint Analysis of Acoustic Scenes and Sound Events Based on Semi-Supervised Training of Sound Events With Partial Labels
por: Imoto, Keisuke
Publicado: (2025)