Audio Explanation Synthesis with Generative Foundation Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Akman, Alican, Sun, Qiyang, Schuller, Björn W. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Cross-Dialect Bird Species Recognition with Dialect-Calibrated Augmentation
von: Ding, Jiani, et al.
Veröffentlicht: (2025)
von: Ding, Jiani, et al.
Veröffentlicht: (2025)
Audio-based Kinship Verification Using Age Domain Conversion
von: Sun, Qiyang, et al.
Veröffentlicht: (2024)
von: Sun, Qiyang, et al.
Veröffentlicht: (2024)
Audio Enhancement for Computer Audition -- An Iterative Training Paradigm Using Sample Importance
von: Milling, Manuel, et al.
Veröffentlicht: (2024)
von: Milling, Manuel, et al.
Veröffentlicht: (2024)
Raw Audio Classification with Cosine Convolutional Neural Network (CosCovNN)
von: Haque, Kazi Nazmul, et al.
Veröffentlicht: (2024)
von: Haque, Kazi Nazmul, et al.
Veröffentlicht: (2024)
Explainable Detection of Machine Generated Music and Early Systematic Evaluation
von: Li, Yupei, et al.
Veröffentlicht: (2024)
von: Li, Yupei, et al.
Veröffentlicht: (2024)
MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge
von: Jing, Xin, et al.
Veröffentlicht: (2025)
von: Jing, Xin, et al.
Veröffentlicht: (2025)
UltraEval-Audio: A Unified Framework for Comprehensive Evaluation of Audio Foundation Models
von: Shi, Qundong, et al.
Veröffentlicht: (2026)
von: Shi, Qundong, et al.
Veröffentlicht: (2026)
From Audio Deepfake Detection to AI-Generated Music Detection -- A Pathway and Overview
von: Li, Yupei, et al.
Veröffentlicht: (2024)
von: Li, Yupei, et al.
Veröffentlicht: (2024)
Are you sure? Analysing Uncertainty Quantification Approaches for Real-world Speech Emotion Recognition
von: Schrüfer, Oliver, et al.
Veröffentlicht: (2024)
von: Schrüfer, Oliver, et al.
Veröffentlicht: (2024)
FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
von: An, Keyu, et al.
Veröffentlicht: (2024)
von: An, Keyu, et al.
Veröffentlicht: (2024)
DreamAudio: Customized Text-to-Audio Generation with Diffusion Models
von: Yuan, Yi, et al.
Veröffentlicht: (2025)
von: Yuan, Yi, et al.
Veröffentlicht: (2025)
Breaking Resource Barriers in Speech Emotion Recognition via Data Distillation
von: Chang, Yi, et al.
Veröffentlicht: (2024)
von: Chang, Yi, et al.
Veröffentlicht: (2024)
LJ-Spoof: A Generatively Varied Corpus for Audio Anti-Spoofing and Synthesis Source Tracing
von: Subramani, Surya, et al.
Veröffentlicht: (2026)
von: Subramani, Surya, et al.
Veröffentlicht: (2026)
Audio Mamba: Pretrained Audio State Space Model For Audio Tagging
von: Lin, Jiaju, et al.
Veröffentlicht: (2024)
von: Lin, Jiaju, et al.
Veröffentlicht: (2024)
AudioTurbo: Fast Text-to-Audio Generation with Rectified Diffusion
von: Zhao, Junqi, et al.
Veröffentlicht: (2025)
von: Zhao, Junqi, et al.
Veröffentlicht: (2025)
autrainer: A Modular and Extensible Deep Learning Toolkit for Computer Audition Tasks
von: Rampp, Simon, et al.
Veröffentlicht: (2024)
von: Rampp, Simon, et al.
Veröffentlicht: (2024)
Audio-based Step-count Estimation for Running -- Windowing and Neural Network Baselines
von: Wagner, Philipp, et al.
Veröffentlicht: (2024)
von: Wagner, Philipp, et al.
Veröffentlicht: (2024)
Audio Codec Augmentation for Robust Collaborative Watermarking of Speech Synthesis
von: Juvela, Lauri, et al.
Veröffentlicht: (2024)
von: Juvela, Lauri, et al.
Veröffentlicht: (2024)
LiteFocus: Accelerated Diffusion Inference for Long Audio Synthesis
von: Tan, Zhenxiong, et al.
Veröffentlicht: (2024)
von: Tan, Zhenxiong, et al.
Veröffentlicht: (2024)
SLIM: Style-Linguistics Mismatch Model for Generalized Audio Deepfake Detection
von: Zhu, Yi, et al.
Veröffentlicht: (2024)
von: Zhu, Yi, et al.
Veröffentlicht: (2024)
EGSTalker: Real-Time Audio-Driven Talking Head Generation with Efficient Gaussian Deformation
von: Zhu, Tianheng, et al.
Veröffentlicht: (2025)
von: Zhu, Tianheng, et al.
Veröffentlicht: (2025)
Unify Variables in Neural Scaling Laws for General Audio Representations via Embedding Effective Rank
von: Deng, Xuyao, et al.
Veröffentlicht: (2025)
von: Deng, Xuyao, et al.
Veröffentlicht: (2025)
A Framework for Synthetic Audio Conversations Generation using Large Language Models
von: Kyaw, Kaung Myat, et al.
Veröffentlicht: (2024)
von: Kyaw, Kaung Myat, et al.
Veröffentlicht: (2024)
WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models
von: Chen, Yifu, et al.
Veröffentlicht: (2025)
von: Chen, Yifu, et al.
Veröffentlicht: (2025)
Audio Mamba: Bidirectional State Space Model for Audio Representation Learning
von: Erol, Mehmet Hamza, et al.
Veröffentlicht: (2024)
von: Erol, Mehmet Hamza, et al.
Veröffentlicht: (2024)
Efficient Autoregressive Audio Modeling via Next-Scale Prediction
von: Qiu, Kai, et al.
Veröffentlicht: (2024)
von: Qiu, Kai, et al.
Veröffentlicht: (2024)
Affect Decoding in Phonated and Silent Speech Production from Surface EMG
von: Pistrosch, Simon, et al.
Veröffentlicht: (2026)
von: Pistrosch, Simon, et al.
Veröffentlicht: (2026)
ViSAGe: Video-to-Spatial Audio Generation
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2025)
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2025)
Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models
von: Song, Zirui, et al.
Veröffentlicht: (2025)
von: Song, Zirui, et al.
Veröffentlicht: (2025)
AFSS: Artifact-Focused Self-Synthesis for Mitigating Bias in Audio Deepfake Detection
von: Nguyen-Le, Hai-Son, et al.
Veröffentlicht: (2026)
von: Nguyen-Le, Hai-Son, et al.
Veröffentlicht: (2026)
Articulatory Feature Prediction from Surface EMG during Speech Production
von: Lee, Jihwan, et al.
Veröffentlicht: (2025)
von: Lee, Jihwan, et al.
Veröffentlicht: (2025)
Harder or Different? Understanding Generalization of Audio Deepfake Detection
von: Müller, Nicolas M., et al.
Veröffentlicht: (2024)
von: Müller, Nicolas M., et al.
Veröffentlicht: (2024)
Who Can Withstand Chat-Audio Attacks? An Evaluation Benchmark for Large Audio-Language Models
von: Yang, Wanqi, et al.
Veröffentlicht: (2024)
von: Yang, Wanqi, et al.
Veröffentlicht: (2024)
Does Current Deepfake Audio Detection Model Effectively Detect ALM-based Deepfake Audio?
von: Xie, Yuankun, et al.
Veröffentlicht: (2024)
von: Xie, Yuankun, et al.
Veröffentlicht: (2024)
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion
von: Chen, Shunian, et al.
Veröffentlicht: (2025)
von: Chen, Shunian, et al.
Veröffentlicht: (2025)
Exploring Musical Roots: Applying Audio Embeddings to Empower Influence Attribution for a Generative Music Model
von: Barnett, Julia, et al.
Veröffentlicht: (2024)
von: Barnett, Julia, et al.
Veröffentlicht: (2024)
TTMBA: Towards Text To Multiple Sources Binaural Audio Generation
von: He, Yuxuan, et al.
Veröffentlicht: (2025)
von: He, Yuxuan, et al.
Veröffentlicht: (2025)
SpectroStream: A Versatile Neural Codec for General Audio
von: Li, Yunpeng, et al.
Veröffentlicht: (2025)
von: Li, Yunpeng, et al.
Veröffentlicht: (2025)
Example-Based Framework for Perceptually Guided Audio Texture Generation
von: Kamath, Purnima, et al.
Veröffentlicht: (2023)
von: Kamath, Purnima, et al.
Veröffentlicht: (2023)
Computer Audition: From Task-Specific Machine Learning to Foundation Models
von: Triantafyllopoulos, Andreas, et al.
Veröffentlicht: (2024)
von: Triantafyllopoulos, Andreas, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Cross-Dialect Bird Species Recognition with Dialect-Calibrated Augmentation
von: Ding, Jiani, et al.
Veröffentlicht: (2025) -
Audio-based Kinship Verification Using Age Domain Conversion
von: Sun, Qiyang, et al.
Veröffentlicht: (2024) -
Audio Enhancement for Computer Audition -- An Iterative Training Paradigm Using Sample Importance
von: Milling, Manuel, et al.
Veröffentlicht: (2024) -
Raw Audio Classification with Cosine Convolutional Neural Network (CosCovNN)
von: Haque, Kazi Nazmul, et al.
Veröffentlicht: (2024) -
Explainable Detection of Machine Generated Music and Early Systematic Evaluation
von: Li, Yupei, et al.
Veröffentlicht: (2024)