Towards High-Fidelity and Controllable Bioacoustic Generation via Enhanced Diffusion Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Song, Tianyu, Ta, Ton Viet |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DiTSE: High-Fidelity Generative Speech Enhancement via Latent Diffusion Transformers
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2025)
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2025)
Towards Deep Active Learning in Avian Bioacoustics
von: Rauch, Lukas, et al.
Veröffentlicht: (2024)
von: Rauch, Lukas, et al.
Veröffentlicht: (2024)
Adaptive Learning via a Negative Selection Strategy for Few-Shot Bioacoustic Event Detection
von: Chen, Yaxiong, et al.
Veröffentlicht: (2024)
von: Chen, Yaxiong, et al.
Veröffentlicht: (2024)
Robust Bioacoustic Detection via Richly Labelled Synthetic Soundscape Augmentation
von: Soltero, Kaspar, et al.
Veröffentlicht: (2025)
von: Soltero, Kaspar, et al.
Veröffentlicht: (2025)
Few-Shot Bioacoustic Event Detection with Frame-Level Embedding Learning System
von: Zhao, PengYuan, et al.
Veröffentlicht: (2024)
von: Zhao, PengYuan, et al.
Veröffentlicht: (2024)
Learning Domain-Robust Bioacoustic Representations for Mosquito Species Classification with Contrastive Learning and Distribution Alignment
von: Hou, Yuanbo, et al.
Veröffentlicht: (2025)
von: Hou, Yuanbo, et al.
Veröffentlicht: (2025)
InstructSing: High-Fidelity Singing Voice Generation via Instructing Yourself
von: Zeng, Chang, et al.
Veröffentlicht: (2024)
von: Zeng, Chang, et al.
Veröffentlicht: (2024)
Distilling Spectrograms into Tokens: Fast and Lightweight Bioacoustic Classification for BirdCLEF+ 2025
von: Miyaguchi, Anthony, et al.
Veröffentlicht: (2025)
von: Miyaguchi, Anthony, et al.
Veröffentlicht: (2025)
Large Language Models and Non-Negative Matrix Factorization for Bioacoustic Signal Decomposition
von: Torabi, Yasaman, et al.
Veröffentlicht: (2025)
von: Torabi, Yasaman, et al.
Veröffentlicht: (2025)
BioME: A Resource-Efficient Bioacoustic Foundational Model for IoT Applications
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2026)
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2026)
High-Fidelity Generative Audio Compression at 0.275kbps
von: Ma, Hao, et al.
Veröffentlicht: (2026)
von: Ma, Hao, et al.
Veröffentlicht: (2026)
LongCat-AudioDiT: High-Fidelity Diffusion Text-to-Speech in the Waveform Latent Space
von: Xin, Detai, et al.
Veröffentlicht: (2026)
von: Xin, Detai, et al.
Veröffentlicht: (2026)
FlashAudio: Rectified Flows for Fast and High-Fidelity Text-to-Audio Generation
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
METEOR: Melody-aware Texture-controllable Symbolic Orchestral Music Generation via Transformer VAE
von: Le, Dinh-Viet-Toan, et al.
Veröffentlicht: (2024)
von: Le, Dinh-Viet-Toan, et al.
Veröffentlicht: (2024)
High-Fidelity Neural Phonetic Posteriorgrams
von: Churchwell, Cameron, et al.
Veröffentlicht: (2024)
von: Churchwell, Cameron, et al.
Veröffentlicht: (2024)
SonicRAG : High Fidelity Sound Effects Synthesis Based on Retrival Augmented Generation
von: Guo, Yu-Ren, et al.
Veröffentlicht: (2025)
von: Guo, Yu-Ren, et al.
Veröffentlicht: (2025)
STSR: High-Fidelity Speech Super-Resolution via Spectral-Transient Context Modeling
von: Yuan, Jiajun, et al.
Veröffentlicht: (2025)
von: Yuan, Jiajun, et al.
Veröffentlicht: (2025)
High Fidelity Text-Guided Music Editing via Single-Stage Flow Matching
von: Lan, Gael Le, et al.
Veröffentlicht: (2024)
von: Lan, Gael Le, et al.
Veröffentlicht: (2024)
HILCodec: High-Fidelity and Lightweight Neural Audio Codec
von: Ahn, Sunghwan, et al.
Veröffentlicht: (2024)
von: Ahn, Sunghwan, et al.
Veröffentlicht: (2024)
AudioGAN: A Compact and Efficient Framework for Real-Time High-Fidelity Text-to-Audio Generation
von: Chung, HaeChun
Veröffentlicht: (2025)
von: Chung, HaeChun
Veröffentlicht: (2025)
Perch 2.0: The Bittern Lesson for Bioacoustics
von: van Merriënboer, Bart, et al.
Veröffentlicht: (2025)
von: van Merriënboer, Bart, et al.
Veröffentlicht: (2025)
AccentBox: Towards High-Fidelity Zero-Shot Accent Generation
von: Zhong, Jinzuomu, et al.
Veröffentlicht: (2024)
von: Zhong, Jinzuomu, et al.
Veröffentlicht: (2024)
Enhancing Emotional Text-to-Speech Controllability with Natural Language Guidance through Contrastive Learning and Diffusion Models
von: Jing, Xin, et al.
Veröffentlicht: (2024)
von: Jing, Xin, et al.
Veröffentlicht: (2024)
Aggregation Strategies for Efficient Annotation of Bioacoustic Sound Events Using Active Learning
von: Lindholm, Richard, et al.
Veröffentlicht: (2025)
von: Lindholm, Richard, et al.
Veröffentlicht: (2025)
HiFi-SR: A Unified Generative Transformer-Convolutional Adversarial Network for High-Fidelity Speech Super-Resolution
von: Zhao, Shengkui, et al.
Veröffentlicht: (2025)
von: Zhao, Shengkui, et al.
Veröffentlicht: (2025)
STFTCodec: High-Fidelity Audio Compression through Time-Frequency Domain Representation
von: Feng, Tao, et al.
Veröffentlicht: (2025)
von: Feng, Tao, et al.
Veröffentlicht: (2025)
SwitchCodec: A High-Fidelity Nerual Audio Codec With Sparse Quantization
von: Wang, Jin, et al.
Veröffentlicht: (2025)
von: Wang, Jin, et al.
Veröffentlicht: (2025)
EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer
von: Hai, Jiarui, et al.
Veröffentlicht: (2024)
von: Hai, Jiarui, et al.
Veröffentlicht: (2024)
GD-Retriever: Controllable Generative Text-Music Retrieval with Diffusion Models
von: Guinot, Julien, et al.
Veröffentlicht: (2025)
von: Guinot, Julien, et al.
Veröffentlicht: (2025)
ConSinger: Efficient High-Fidelity Singing Voice Generation with Minimal Steps
von: Song, Yulin, et al.
Veröffentlicht: (2024)
von: Song, Yulin, et al.
Veröffentlicht: (2024)
VQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature
von: Du, Chenpeng, et al.
Veröffentlicht: (2022)
von: Du, Chenpeng, et al.
Veröffentlicht: (2022)
HiFi-Glot: High-Fidelity Neural Formant Synthesis with Differentiable Resonant Filters
von: Gu, Yicheng, et al.
Veröffentlicht: (2024)
von: Gu, Yicheng, et al.
Veröffentlicht: (2024)
DualSpeech: Enhancing Speaker-Fidelity and Text-Intelligibility Through Dual Classifier-Free Guidance
von: Yang, Jinhyeok, et al.
Veröffentlicht: (2024)
von: Yang, Jinhyeok, et al.
Veröffentlicht: (2024)
Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer
von: Lei, Ke, et al.
Veröffentlicht: (2026)
von: Lei, Ke, et al.
Veröffentlicht: (2026)
MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
Temporal Feature Learning in Weakly Labelled Bioacoustic Cetacean Datasets via a Variational Autoencoder and Temporal Convolutional Network: An Interdisciplinary Approach
von: Fonollosa, Laia Garrobé, et al.
Veröffentlicht: (2024)
von: Fonollosa, Laia Garrobé, et al.
Veröffentlicht: (2024)
Regularized Contrastive Pre-training for Few-shot Bioacoustic Sound Detection
von: Moummad, Ilyass, et al.
Veröffentlicht: (2023)
von: Moummad, Ilyass, et al.
Veröffentlicht: (2023)
APCodec+: A Spectrum-Coding-Based High-Fidelity and High-Compression-Rate Neural Audio Codec with Staged Training Paradigm
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024)
von: Du, Hui-Peng, et al.
Veröffentlicht: (2024)
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation
von: Kumar, Sonal, et al.
Veröffentlicht: (2024)
von: Kumar, Sonal, et al.
Veröffentlicht: (2024)
Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments
von: Yoneyama, Reo, et al.
Veröffentlicht: (2025)
von: Yoneyama, Reo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
DiTSE: High-Fidelity Generative Speech Enhancement via Latent Diffusion Transformers
von: Guimarães, Heitor R., et al.
Veröffentlicht: (2025) -
Towards Deep Active Learning in Avian Bioacoustics
von: Rauch, Lukas, et al.
Veröffentlicht: (2024) -
Adaptive Learning via a Negative Selection Strategy for Few-Shot Bioacoustic Event Detection
von: Chen, Yaxiong, et al.
Veröffentlicht: (2024) -
Robust Bioacoustic Detection via Richly Labelled Synthetic Soundscape Augmentation
von: Soltero, Kaspar, et al.
Veröffentlicht: (2025) -
Few-Shot Bioacoustic Event Detection with Frame-Level Embedding Learning System
von: Zhao, PengYuan, et al.
Veröffentlicht: (2024)