An Empirical Analysis of Task-Induced Encoder Bias in Fréchet Audio Distance
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Jeong, Wonwoo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Optimal Transport Audio Distance with Learned Riemannian Ground Metrics
von: Jeong, Wonwoo
Veröffentlicht: (2026)
von: Jeong, Wonwoo
Veröffentlicht: (2026)
Correlation of Fréchet Audio Distance With Human Perception of Environmental Audio Is Embedding Dependant
von: Tailleur, Modan, et al.
Veröffentlicht: (2024)
von: Tailleur, Modan, et al.
Veröffentlicht: (2024)
Addressing Emotion Bias in Music Emotion Recognition and Generation with Frechet Audio Distance
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
The ICME 2025 Audio Encoder Capability Challenge
von: Zhang, Junbo, et al.
Veröffentlicht: (2025)
von: Zhang, Junbo, et al.
Veröffentlicht: (2025)
The Interspeech 2026 Audio Encoder Capability Challenge for Large Audio Language Models
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2026)
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2026)
Efficient Audio Captioning with Encoder-Level Knowledge Distillation
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
The CMU-AIST submission for the ICME 2025 Audio Encoder Challenge
von: Bharadwaj, Shikhar, et al.
Veröffentlicht: (2026)
von: Bharadwaj, Shikhar, et al.
Veröffentlicht: (2026)
Pengi: An Audio Language Model for Audio Tasks
von: Deshmukh, Soham, et al.
Veröffentlicht: (2023)
von: Deshmukh, Soham, et al.
Veröffentlicht: (2023)
Angular Distance Distribution Loss for Audio Classification
von: Almudévar, Antonio, et al.
Veröffentlicht: (2024)
von: Almudévar, Antonio, et al.
Veröffentlicht: (2024)
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance
von: Zhang, Junbo, et al.
Veröffentlicht: (2025)
von: Zhang, Junbo, et al.
Veröffentlicht: (2025)
OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder
von: Bharadwaj, Shikhar, et al.
Veröffentlicht: (2025)
von: Bharadwaj, Shikhar, et al.
Veröffentlicht: (2025)
Fx-Encoder++: Extracting Instrument-Wise Audio Effects Representations from Mixtures
von: Yeh, Yen-Tung, et al.
Veröffentlicht: (2025)
von: Yeh, Yen-Tung, et al.
Veröffentlicht: (2025)
Enhancing Speech Large Language Models with Prompt-Aware Mixture of Audio Encoders
von: Shan, Weiqiao, et al.
Veröffentlicht: (2025)
von: Shan, Weiqiao, et al.
Veröffentlicht: (2025)
Genuine-Focused Learning using Mask AutoEncoder for Generalized Fake Audio Detection
von: Wang, Xiaopeng, et al.
Veröffentlicht: (2024)
von: Wang, Xiaopeng, et al.
Veröffentlicht: (2024)
Speaker Distance Estimation in Enclosures from Single-Channel Audio
von: Neri, Michael, et al.
Veröffentlicht: (2024)
von: Neri, Michael, et al.
Veröffentlicht: (2024)
Measuring Audio Prompt Adherence with Distribution-based Embedding Distances
von: Grachten, Maarten
Veröffentlicht: (2024)
von: Grachten, Maarten
Veröffentlicht: (2024)
MT2KD: Towards A General-Purpose Encoder for Speech, Speaker, and Audio Events
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2024)
UniAudio 1.5: Large Language Model-driven Audio Codec is A Few-shot Audio Task Learner
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
CatchPhrase: EXPrompt-Guided Encoder Adaptation for Audio-to-Image Generation
von: Oh, Hyunwoo, et al.
Veröffentlicht: (2025)
von: Oh, Hyunwoo, et al.
Veröffentlicht: (2025)
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey
von: Su, Yi, et al.
Veröffentlicht: (2025)
von: Su, Yi, et al.
Veröffentlicht: (2025)
Jointly Recognizing Speech and Singing Voices Based on Multi-Task Audio Source Separation
von: Bai, Ye, et al.
Veröffentlicht: (2024)
von: Bai, Ye, et al.
Veröffentlicht: (2024)
Frechet Music Distance: A Metric For Generative Symbolic Music Evaluation
von: Retkowski, Jan, et al.
Veröffentlicht: (2024)
von: Retkowski, Jan, et al.
Veröffentlicht: (2024)
Do Foundational Audio Encoders Understand Music Structure?
von: Toyama, Keisuke, et al.
Veröffentlicht: (2025)
von: Toyama, Keisuke, et al.
Veröffentlicht: (2025)
Can Audio Reveal Music Performance Difficulty? Insights from the Piano Syllabus Dataset
von: Ramoneda, Pedro, et al.
Veröffentlicht: (2024)
von: Ramoneda, Pedro, et al.
Veröffentlicht: (2024)
AVFSNet: Audio-Visual Speech Separation for Flexible Number of Speakers with Multi-Scale and Multi-Task Learning
von: Zhang, Daning, et al.
Veröffentlicht: (2025)
von: Zhang, Daning, et al.
Veröffentlicht: (2025)
Residual Learning for Neural Ambisonics Encoders
von: Deppisch, Thomas, et al.
Veröffentlicht: (2026)
von: Deppisch, Thomas, et al.
Veröffentlicht: (2026)
UniAudio: An Audio Foundation Model Toward Universal Audio Generation
von: Yang, Dongchao, et al.
Veröffentlicht: (2023)
von: Yang, Dongchao, et al.
Veröffentlicht: (2023)
A Practical Guide to Spectrogram Analysis for Audio Signal Processing
von: Khodzhaev, Zulfidin
Veröffentlicht: (2024)
von: Khodzhaev, Zulfidin
Veröffentlicht: (2024)
Analysis of ABC Frontend Audio Systems for the NIST-SRE24
von: Barahona, Sara, et al.
Veröffentlicht: (2025)
von: Barahona, Sara, et al.
Veröffentlicht: (2025)
ADD 2023: Towards Audio Deepfake Detection and Analysis in the Wild
von: Yi, Jiangyan, et al.
Veröffentlicht: (2024)
von: Yi, Jiangyan, et al.
Veröffentlicht: (2024)
MOSS-Audio-Tokenizer: Scaling Audio Tokenizers for Future Audio Foundation Models
von: Gong, Yitian, et al.
Veröffentlicht: (2026)
von: Gong, Yitian, et al.
Veröffentlicht: (2026)
Streaming Audio Transformers for Online Audio Tagging
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2023)
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2023)
Discrete Audio Representations for Automated Audio Captioning
von: Tian, Jingguang, et al.
Veröffentlicht: (2025)
von: Tian, Jingguang, et al.
Veröffentlicht: (2025)
Comprehensive Layer-wise Analysis of SSL Models for Audio Deepfake Detection
von: Kheir, Yassine El, et al.
Veröffentlicht: (2025)
von: Kheir, Yassine El, et al.
Veröffentlicht: (2025)
Audio-Mind: An Auditable Agentic Framework for Audio Understanding
von: Wang, Yucheng, et al.
Veröffentlicht: (2026)
von: Wang, Yucheng, et al.
Veröffentlicht: (2026)
SemanticAudio: Audio Generation and Editing in Semantic Space
von: Dai, Zheqi, et al.
Veröffentlicht: (2026)
von: Dai, Zheqi, et al.
Veröffentlicht: (2026)
MACE: Leveraging Audio for Evaluating Audio Captioning Systems
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
Audio Entailment: Assessing Deductive Reasoning for Audio Understanding
von: Deshmukh, Soham, et al.
Veröffentlicht: (2024)
von: Deshmukh, Soham, et al.
Veröffentlicht: (2024)
A Comparative Analysis of Poetry Reading Audio: Singing, Narrating, or Somewhere In Between?
von: Choi, Kahyun, et al.
Veröffentlicht: (2024)
von: Choi, Kahyun, et al.
Veröffentlicht: (2024)
Comparative Analysis of Finite Difference and Finite Element Method for Audio Waveform Simulation
von: Florin, Juliette
Veröffentlicht: (2025)
von: Florin, Juliette
Veröffentlicht: (2025)
Ähnliche Einträge
-
Optimal Transport Audio Distance with Learned Riemannian Ground Metrics
von: Jeong, Wonwoo
Veröffentlicht: (2026) -
Correlation of Fréchet Audio Distance With Human Perception of Environmental Audio Is Embedding Dependant
von: Tailleur, Modan, et al.
Veröffentlicht: (2024) -
Addressing Emotion Bias in Music Emotion Recognition and Generation with Frechet Audio Distance
von: Li, Yuanchao, et al.
Veröffentlicht: (2024) -
The ICME 2025 Audio Encoder Capability Challenge
von: Zhang, Junbo, et al.
Veröffentlicht: (2025) -
The Interspeech 2026 Audio Encoder Capability Challenge for Large Audio Language Models
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2026)