A Generative-First Neural Audio Autoencoder
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Casebeer, Jonah, Zhu, Ge, Wang, Zhepei, Bryan, Nicholas J. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Re-Bottleneck: Latent Re-Structuring for Neural Audio Autoencoders
von: Bralios, Dimitrios, et al.
Veröffentlicht: (2025)
von: Bralios, Dimitrios, et al.
Veröffentlicht: (2025)
Scaling Up Adaptive Filter Optimizers
von: Casebeer, Jonah, et al.
Veröffentlicht: (2024)
von: Casebeer, Jonah, et al.
Veröffentlicht: (2024)
Learning to Upsample and Upmix Audio in the Latent Domain
von: Bralios, Dimitrios, et al.
Veröffentlicht: (2025)
von: Bralios, Dimitrios, et al.
Veröffentlicht: (2025)
Presto! Distilling Steps and Layers for Accelerating Music Generation
von: Novack, Zachary, et al.
Veröffentlicht: (2024)
von: Novack, Zachary, et al.
Veröffentlicht: (2024)
Audio Generation Through Score-Based Generative Modeling: Design Principles and Implementation
von: Zhu, Ge, et al.
Veröffentlicht: (2025)
von: Zhu, Ge, et al.
Veröffentlicht: (2025)
On Class Separability Pitfalls In Audio-Text Contrastive Zero-Shot Learning
von: Tavares, Tiago, et al.
Veröffentlicht: (2024)
von: Tavares, Tiago, et al.
Veröffentlicht: (2024)
Code Drift: Towards Idempotent Neural Audio Codecs
von: O'Reilly, Patrick, et al.
Veröffentlicht: (2024)
von: O'Reilly, Patrick, et al.
Veröffentlicht: (2024)
Network Modulation Synthesis: New Algorithms for Generating Musical Audio Using Autoencoder Networks
von: Hyrkas, Jeremy
Veröffentlicht: (2021)
von: Hyrkas, Jeremy
Veröffentlicht: (2021)
Cacophony: An Improved Contrastive Audio-Text Model
von: Zhu, Ge, et al.
Veröffentlicht: (2024)
von: Zhu, Ge, et al.
Veröffentlicht: (2024)
CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions
von: Zhu, Xinfa, et al.
Veröffentlicht: (2025)
von: Zhu, Xinfa, et al.
Veröffentlicht: (2025)
First-Shot Unsupervised Anomalous Sound Detection With Unknown Anomalies Estimated by Metadata-Assisted Audio Generation
von: Zhang, Hejing, et al.
Veröffentlicht: (2023)
von: Zhang, Hejing, et al.
Veröffentlicht: (2023)
Audio Editing with Non-Rigid Text Prompts
von: Paissan, Francesco, et al.
Veröffentlicht: (2023)
von: Paissan, Francesco, et al.
Veröffentlicht: (2023)
MusicHiFi: Fast High-Fidelity Stereo Vocoding
von: Zhu, Ge, et al.
Veröffentlicht: (2024)
von: Zhu, Ge, et al.
Veröffentlicht: (2024)
Investigating Neural Audio Codecs for Speech Language Model-Based Speech Generation
von: Li, Jiaqi, et al.
Veröffentlicht: (2024)
von: Li, Jiaqi, et al.
Veröffentlicht: (2024)
DualCodec: A Low-Frame-Rate, Semantically-Enhanced Neural Audio Codec for Speech Generation
von: Li, Jiaqi, et al.
Veröffentlicht: (2025)
von: Li, Jiaqi, et al.
Veröffentlicht: (2025)
The Rhythm In Anything: Audio-Prompted Drums Generation with Masked Language Modeling
von: O'Reilly, Patrick, et al.
Veröffentlicht: (2025)
von: O'Reilly, Patrick, et al.
Veröffentlicht: (2025)
Sketch2Sound: Controllable Audio Generation via Time-Varying Signals and Sonic Imitations
von: García, Hugo Flores, et al.
Veröffentlicht: (2024)
von: García, Hugo Flores, et al.
Veröffentlicht: (2024)
AudioLCM: Text-to-Audio Generation with Latent Consistency Models
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
Enhancing Neural Audio Fingerprint Robustness to Audio Degradation for Music Identification
von: Araz, R. Oguz, et al.
Veröffentlicht: (2025)
von: Araz, R. Oguz, et al.
Veröffentlicht: (2025)
Combining Audio and Non-Audio Inputs in Evolved Neural Networks for Ovenbird
von: Hernandez, Sergio Poo, et al.
Veröffentlicht: (2025)
von: Hernandez, Sergio Poo, et al.
Veröffentlicht: (2025)
UniAudio: An Audio Foundation Model Toward Universal Audio Generation
von: Yang, Dongchao, et al.
Veröffentlicht: (2023)
von: Yang, Dongchao, et al.
Veröffentlicht: (2023)
Enhancing Generalization in Audio Deepfake Detection: A Neural Collapse based Sampling and Training Approach
von: Yousif, Mohammed, et al.
Veröffentlicht: (2024)
von: Yousif, Mohammed, et al.
Veröffentlicht: (2024)
EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer
von: Hai, Jiarui, et al.
Veröffentlicht: (2024)
von: Hai, Jiarui, et al.
Veröffentlicht: (2024)
AudioComposer: Towards Fine-grained Audio Generation with Natural Language Descriptions
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2024)
MiDashengLM: Efficient Audio Understanding with General Audio Captions
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2025)
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2025)
SemanticAudio: Audio Generation and Editing in Semantic Space
von: Dai, Zheqi, et al.
Veröffentlicht: (2026)
von: Dai, Zheqi, et al.
Veröffentlicht: (2026)
PyNeuralFx: A Python Package for Neural Audio Effect Modeling
von: Yeh, Yen-Tung, et al.
Veröffentlicht: (2024)
von: Yeh, Yen-Tung, et al.
Veröffentlicht: (2024)
FlashAudio: Rectified Flows for Fast and High-Fidelity Text-to-Audio Generation
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
Open-Set Source Tracing of Audio Deepfake Systems
von: Klein, Nicholas, et al.
Veröffentlicht: (2025)
von: Klein, Nicholas, et al.
Veröffentlicht: (2025)
SRC-gAudio: Sampling-Rate-Controlled Audio Generation
von: Li, Chenxing, et al.
Veröffentlicht: (2024)
von: Li, Chenxing, et al.
Veröffentlicht: (2024)
Towards Neural Audio Codec Source Parsing
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
LAFMA: A Latent Flow Matching Model for Text-to-Audio Generation
von: Guan, Wenhao, et al.
Veröffentlicht: (2024)
von: Guan, Wenhao, et al.
Veröffentlicht: (2024)
Deep Audio Watermarks are Shallow: Limitations of Post-Hoc Watermarking Techniques for Speech
von: O'Reilly, Patrick, et al.
Veröffentlicht: (2025)
von: O'Reilly, Patrick, et al.
Veröffentlicht: (2025)
Text2FX: Harnessing CLAP Embeddings for Text-Guided Audio Effects
von: Chu, Annie, et al.
Veröffentlicht: (2024)
von: Chu, Annie, et al.
Veröffentlicht: (2024)
AudioEval: Automatic Dual-Perspective and Multi-Dimensional Evaluation of Text-to-Audio-Generation
von: Wang, Hui, et al.
Veröffentlicht: (2025)
von: Wang, Hui, et al.
Veröffentlicht: (2025)
Source Tracing of Audio Deepfake Systems
von: Klein, Nicholas, et al.
Veröffentlicht: (2024)
von: Klein, Nicholas, et al.
Veröffentlicht: (2024)
A Reference-free Metric for Language-Queried Audio Source Separation using Contrastive Language-Audio Pretraining
von: Xiao, Feiyang, et al.
Veröffentlicht: (2024)
von: Xiao, Feiyang, et al.
Veröffentlicht: (2024)
Natural Language Supervision for General-Purpose Audio Representations
von: Elizalde, Benjamin, et al.
Veröffentlicht: (2023)
von: Elizalde, Benjamin, et al.
Veröffentlicht: (2023)
Streaming Audio Transformers for Online Audio Tagging
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2023)
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2023)
Complex Image-Generative Diffusion Transformer for Audio Denoising
von: Li, Junhui, et al.
Veröffentlicht: (2024)
von: Li, Junhui, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Re-Bottleneck: Latent Re-Structuring for Neural Audio Autoencoders
von: Bralios, Dimitrios, et al.
Veröffentlicht: (2025) -
Scaling Up Adaptive Filter Optimizers
von: Casebeer, Jonah, et al.
Veröffentlicht: (2024) -
Learning to Upsample and Upmix Audio in the Latent Domain
von: Bralios, Dimitrios, et al.
Veröffentlicht: (2025) -
Presto! Distilling Steps and Layers for Accelerating Music Generation
von: Novack, Zachary, et al.
Veröffentlicht: (2024) -
Audio Generation Through Score-Based Generative Modeling: Design Principles and Implementation
von: Zhu, Ge, et al.
Veröffentlicht: (2025)