AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kishi, Minoru, Sakai, Ryosuke, Takamichi, Shinnosuke, Kanamori, Yusuke, Okamoto, Yuki |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio
von: Kanamori, Yusuke, et al.
Veröffentlicht: (2025)
von: Kanamori, Yusuke, et al.
Veröffentlicht: (2025)
SpeechBERTScore: Reference-Aware Automatic Evaluation of Speech Generation Leveraging NLP Evaluation Metrics
von: Saeki, Takaaki, et al.
Veröffentlicht: (2024)
von: Saeki, Takaaki, et al.
Veröffentlicht: (2024)
Correlation of Fréchet Audio Distance With Human Perception of Environmental Audio Is Embedding Dependant
von: Tailleur, Modan, et al.
Veröffentlicht: (2024)
von: Tailleur, Modan, et al.
Veröffentlicht: (2024)
Drum-to-Vocal Percussion Sound Conversion and Its Evaluation Methodology
von: Nobukawa, Rinka, et al.
Veröffentlicht: (2025)
von: Nobukawa, Rinka, et al.
Veröffentlicht: (2025)
Voice Conversion for Likability Control via Automated Rating of Speech Synthesis Corpora
von: Suda, Hitoshi, et al.
Veröffentlicht: (2025)
von: Suda, Hitoshi, et al.
Veröffentlicht: (2025)
Construction and Analysis of Impression Caption Dataset for Environmental Sounds
von: Okamoto, Yuki, et al.
Veröffentlicht: (2024)
von: Okamoto, Yuki, et al.
Veröffentlicht: (2024)
YODAS: Youtube-Oriented Dataset for Audio and Speech
von: Li, Xinjian, et al.
Veröffentlicht: (2024)
von: Li, Xinjian, et al.
Veröffentlicht: (2024)
Human-CLAP: Human-perception-based contrastive language-audio pretraining
von: Takano, Taisei, et al.
Veröffentlicht: (2025)
von: Takano, Taisei, et al.
Veröffentlicht: (2025)
Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models
von: Kando, Shunsuke, et al.
Veröffentlicht: (2025)
von: Kando, Shunsuke, et al.
Veröffentlicht: (2025)
Active Learning for Text-to-Speech Synthesis with Informative Sample Collection
von: Seki, Kentaro, et al.
Veröffentlicht: (2025)
von: Seki, Kentaro, et al.
Veröffentlicht: (2025)
BINAQUAL: A Full-Reference Objective Localization Similarity Metric for Binaural Audio
von: Panah, Davoud Shariat, et al.
Veröffentlicht: (2025)
von: Panah, Davoud Shariat, et al.
Veröffentlicht: (2025)
Challenge on Sound Scene Synthesis: Evaluating Text-to-Audio Generation
von: Lee, Junwon, et al.
Veröffentlicht: (2024)
von: Lee, Junwon, et al.
Veröffentlicht: (2024)
Who Finds This Voice Attractive? A Large-Scale Experiment Using In-the-Wild Data
von: Suda, Hitoshi, et al.
Veröffentlicht: (2024)
von: Suda, Hitoshi, et al.
Veröffentlicht: (2024)
DnR-nonverbal: Cinematic Audio Source Separation Dataset Containing Non-Verbal Sounds
von: Hasumi, Takuya, et al.
Veröffentlicht: (2025)
von: Hasumi, Takuya, et al.
Veröffentlicht: (2025)
Spatial Voice Conversion: Voice Conversion Preserving Spatial Information and Non-target Signals
von: Seki, Kentaro, et al.
Veröffentlicht: (2024)
von: Seki, Kentaro, et al.
Veröffentlicht: (2024)
Building speech corpus with diverse voice characteristics for its prompt-based representation
von: Watanabe, Aya, et al.
Veröffentlicht: (2024)
von: Watanabe, Aya, et al.
Veröffentlicht: (2024)
JVNV: A Corpus of Japanese Emotional Speech with Verbal Content and Nonverbal Expressions
von: Xin, Detai, et al.
Veröffentlicht: (2023)
von: Xin, Detai, et al.
Veröffentlicht: (2023)
ACES: Evaluating Automated Audio Captioning Models on the Semantics of Sounds
von: Wijngaard, Gijs, et al.
Veröffentlicht: (2024)
von: Wijngaard, Gijs, et al.
Veröffentlicht: (2024)
Online Single-Channel Audio-Based Sound Speed Estimation for Robust Multi-Channel Audio Control
von: Fuglsig, Andreas Jonas, et al.
Veröffentlicht: (2026)
von: Fuglsig, Andreas Jonas, et al.
Veröffentlicht: (2026)
SpecMaskGIT: Masked Generative Modeling of Audio Spectrograms for Efficient Audio Synthesis and Beyond
von: Comunità, Marco, et al.
Veröffentlicht: (2024)
von: Comunità, Marco, et al.
Veröffentlicht: (2024)
Detection of Deepfake Environmental Audio
von: Ouajdi, Hafsa, et al.
Veröffentlicht: (2024)
von: Ouajdi, Hafsa, et al.
Veröffentlicht: (2024)
SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer
von: Wang, Helin, et al.
Veröffentlicht: (2024)
von: Wang, Helin, et al.
Veröffentlicht: (2024)
BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
von: Xin, Detai, et al.
Veröffentlicht: (2024)
von: Xin, Detai, et al.
Veröffentlicht: (2024)
Audio Fingerprinting with Holographic Reduced Representations
von: Fujita, Yusuke, et al.
Veröffentlicht: (2024)
von: Fujita, Yusuke, et al.
Veröffentlicht: (2024)
AudioSpa: Spatializing Sound Events with Text
von: Feng, Linfeng, et al.
Veröffentlicht: (2025)
von: Feng, Linfeng, et al.
Veröffentlicht: (2025)
Region-Specific Audio Tagging for Spatial Sound
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2025)
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2025)
Assessing the Alignment of Audio Representations with Timbre Similarity Ratings
von: Tian, Haokun, et al.
Veröffentlicht: (2025)
von: Tian, Haokun, et al.
Veröffentlicht: (2025)
Masked Audio Modeling with CLAP and Multi-Objective Learning
von: Xin, Yifei, et al.
Veröffentlicht: (2024)
von: Xin, Yifei, et al.
Veröffentlicht: (2024)
MACE: Leveraging Audio for Evaluating Audio Captioning Systems
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
Noise-Robust Voice Conversion by Conditional Denoising Training Using Latent Variables of Recording Quality and Environment
von: Igarashi, Takuto, et al.
Veröffentlicht: (2024)
von: Igarashi, Takuto, et al.
Veröffentlicht: (2024)
SRC4VC: Smartphone-Recorded Corpus for Voice Conversion Benchmark
von: Saito, Yuki, et al.
Veröffentlicht: (2024)
von: Saito, Yuki, et al.
Veröffentlicht: (2024)
Enhancing Neural Audio Fingerprint Robustness to Audio Degradation for Music Identification
von: Araz, R. Oguz, et al.
Veröffentlicht: (2025)
von: Araz, R. Oguz, et al.
Veröffentlicht: (2025)
Evaluating CNN with Stacked Feature Representations and Audio Spectrogram Transformer Models for Sound Classification
von: Dehaghania, Parinaz Binandeh, et al.
Veröffentlicht: (2026)
von: Dehaghania, Parinaz Binandeh, et al.
Veröffentlicht: (2026)
DNN-based ensemble singing voice synthesis with interactions between singers
von: Hyodo, Hiroaki, et al.
Veröffentlicht: (2024)
von: Hyodo, Hiroaki, et al.
Veröffentlicht: (2024)
SoundCollage: Automated Discovery of New Classes in Audio Datasets
von: Choi, Ryuhaerang, et al.
Veröffentlicht: (2024)
von: Choi, Ryuhaerang, et al.
Veröffentlicht: (2024)
Effective Pre-Training of Audio Transformers for Sound Event Detection
von: Schmid, Florian, et al.
Veröffentlicht: (2024)
von: Schmid, Florian, et al.
Veröffentlicht: (2024)
Enhance Temporal Relations in Audio Captioning with Sound Event Detection
von: Xie, Zeyu, et al.
Veröffentlicht: (2023)
von: Xie, Zeyu, et al.
Veröffentlicht: (2023)
Beyond Video-to-SFX: Video to Audio Synthesis with Environmentally Aware Speech
von: Niu, Xinlei, et al.
Veröffentlicht: (2025)
von: Niu, Xinlei, et al.
Veröffentlicht: (2025)
Evaluating Sound Similarity Metrics for Differentiable, Iterative Sound-Matching
von: Salimi, Amir, et al.
Veröffentlicht: (2025)
von: Salimi, Amir, et al.
Veröffentlicht: (2025)
SoundBeam meets M2D: Target Sound Extraction with Audio Foundation Model
von: Hernandez-Olivan, Carlos, et al.
Veröffentlicht: (2024)
von: Hernandez-Olivan, Carlos, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio
von: Kanamori, Yusuke, et al.
Veröffentlicht: (2025) -
SpeechBERTScore: Reference-Aware Automatic Evaluation of Speech Generation Leveraging NLP Evaluation Metrics
von: Saeki, Takaaki, et al.
Veröffentlicht: (2024) -
Correlation of Fréchet Audio Distance With Human Perception of Environmental Audio Is Embedding Dependant
von: Tailleur, Modan, et al.
Veröffentlicht: (2024) -
Drum-to-Vocal Percussion Sound Conversion and Its Evaluation Methodology
von: Nobukawa, Rinka, et al.
Veröffentlicht: (2025) -
Voice Conversion for Likability Control via Automated Rating of Speech Synthesis Corpora
von: Suda, Hitoshi, et al.
Veröffentlicht: (2025)