Unsupervised Composable Representations for Audio
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bindi, Giovanni, Esling, Philippe |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Combining audio control and style transfer using latent diffusion
von: Demerlé, Nils, et al.
Veröffentlicht: (2024)
von: Demerlé, Nils, et al.
Veröffentlicht: (2024)
uaMix-MAE: Efficient Tuning of Pretrained Audio Transformers with Unsupervised Audio Mixtures
von: Tabassum, Afrina, et al.
Veröffentlicht: (2024)
von: Tabassum, Afrina, et al.
Veröffentlicht: (2024)
HEAR: Holistic Evaluation of Audio Representations
von: Turian, Joseph, et al.
Veröffentlicht: (2022)
von: Turian, Joseph, et al.
Veröffentlicht: (2022)
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer
von: Takeuchi, Daiki, et al.
Veröffentlicht: (2025)
von: Takeuchi, Daiki, et al.
Veröffentlicht: (2025)
An Unsupervised Domain Adaptation Method for Locating Manipulated Region in partially fake Audio
von: Zeng, Siding, et al.
Veröffentlicht: (2024)
von: Zeng, Siding, et al.
Veröffentlicht: (2024)
Zero-Shot Unsupervised and Text-Based Audio Editing Using DDPM Inversion
von: Manor, Hila, et al.
Veröffentlicht: (2024)
von: Manor, Hila, et al.
Veröffentlicht: (2024)
Learning Music Audio Representations With Limited Data
von: Plachouras, Christos, et al.
Veröffentlicht: (2025)
von: Plachouras, Christos, et al.
Veröffentlicht: (2025)
Learning Disentangled Audio Representations through Controlled Synthesis
von: Brima, Yusuf, et al.
Veröffentlicht: (2024)
von: Brima, Yusuf, et al.
Veröffentlicht: (2024)
Motif Mining and Unsupervised Representation Learning for BirdCLEF 2022
von: Miyaguchi, Anthony, et al.
Veröffentlicht: (2022)
von: Miyaguchi, Anthony, et al.
Veröffentlicht: (2022)
COCOLA: Coherence-Oriented Contrastive Learning of Musical Audio Representations
von: Ciranni, Ruben, et al.
Veröffentlicht: (2024)
von: Ciranni, Ruben, et al.
Veröffentlicht: (2024)
PBSCR: The Piano Bootleg Score Composer Recognition Dataset
von: Jain, Arhan, et al.
Veröffentlicht: (2024)
von: Jain, Arhan, et al.
Veröffentlicht: (2024)
Towards Robust Few-shot Class Incremental Learning in Audio Classification using Contrastive Representation
von: Singh, Riyansha, et al.
Veröffentlicht: (2024)
von: Singh, Riyansha, et al.
Veröffentlicht: (2024)
A2SB: Audio-to-Audio Schrodinger Bridges
von: Kong, Zhifeng, et al.
Veröffentlicht: (2025)
von: Kong, Zhifeng, et al.
Veröffentlicht: (2025)
Zero Shot Audio to Audio Emotion Transfer With Speaker Disentanglement
von: Dutta, Soumya, et al.
Veröffentlicht: (2024)
von: Dutta, Soumya, et al.
Veröffentlicht: (2024)
TACOS: Temporally-aligned Audio CaptiOnS for Language-Audio Pretraining
von: Primus, Paul, et al.
Veröffentlicht: (2025)
von: Primus, Paul, et al.
Veröffentlicht: (2025)
Integrating Text-to-Music Models with Language Models: Composing Long Structured Music Pieces
von: Atassi, Lilac
Veröffentlicht: (2024)
von: Atassi, Lilac
Veröffentlicht: (2024)
The Rarity of Musical Audio Signals Within the Space of Possible Audio Generation
von: Collins, Nick
Veröffentlicht: (2024)
von: Collins, Nick
Veröffentlicht: (2024)
Estimated Audio-Caption Correspondences Improve Language-Based Audio Retrieval
von: Primus, Paul, et al.
Veröffentlicht: (2024)
von: Primus, Paul, et al.
Veröffentlicht: (2024)
Fusing Audio and Metadata Embeddings Improves Language-based Audio Retrieval
von: Primus, Paul, et al.
Veröffentlicht: (2024)
von: Primus, Paul, et al.
Veröffentlicht: (2024)
Composer's Assistant 2: Interactive Multi-Track MIDI Infilling with Fine-Grained User Control
von: Malandro, Martin E.
Veröffentlicht: (2024)
von: Malandro, Martin E.
Veröffentlicht: (2024)
Audio Match Cutting: Finding and Creating Matching Audio Transitions in Movies and Videos
von: Fedorishin, Dennis, et al.
Veröffentlicht: (2024)
von: Fedorishin, Dennis, et al.
Veröffentlicht: (2024)
Diffusion Models for Audio Restoration
von: Lemercier, Jean-Marie, et al.
Veröffentlicht: (2024)
von: Lemercier, Jean-Marie, et al.
Veröffentlicht: (2024)
Multi-bit Audio Watermarking
von: Lanzendörfer, Luca A., et al.
Veröffentlicht: (2025)
von: Lanzendörfer, Luca A., et al.
Veröffentlicht: (2025)
Instabilities in Convnets for Raw Audio
von: Haider, Daniel, et al.
Veröffentlicht: (2023)
von: Haider, Daniel, et al.
Veröffentlicht: (2023)
Automatic Contextual Audio Denoising
von: Luong, Diep, et al.
Veröffentlicht: (2026)
von: Luong, Diep, et al.
Veröffentlicht: (2026)
Enhancing Audio-Language Models through Self-Supervised Post-Training with Text-Audio Pairs
von: Sinha, Anshuman, et al.
Veröffentlicht: (2024)
von: Sinha, Anshuman, et al.
Veröffentlicht: (2024)
Can Synthetic Audio From Generative Foundation Models Assist Audio Recognition and Speech Modeling?
von: Feng, Tiantian, et al.
Veröffentlicht: (2024)
von: Feng, Tiantian, et al.
Veröffentlicht: (2024)
Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities
von: Kong, Zhifeng, et al.
Veröffentlicht: (2024)
von: Kong, Zhifeng, et al.
Veröffentlicht: (2024)
"I am bad": Interpreting Stealthy, Universal and Robust Audio Jailbreaks in Audio-Language Models
von: Gupta, Isha, et al.
Veröffentlicht: (2025)
von: Gupta, Isha, et al.
Veröffentlicht: (2025)
Audio Decoding by Inverse Problem Solving
von: T., Pedro J. Villasana, et al.
Veröffentlicht: (2024)
von: T., Pedro J. Villasana, et al.
Veröffentlicht: (2024)
Does Audio Deepfake Detection Generalize?
von: Müller, Nicolas M., et al.
Veröffentlicht: (2022)
von: Müller, Nicolas M., et al.
Veröffentlicht: (2022)
Aligning Audio Captions with Human Preferences
von: Hegde, Kartik, et al.
Veröffentlicht: (2025)
von: Hegde, Kartik, et al.
Veröffentlicht: (2025)
Multi-Task Corrupted Prediction for Learning Robust Audio-Visual Speech Representation
von: Kim, Sungnyun, et al.
Veröffentlicht: (2025)
von: Kim, Sungnyun, et al.
Veröffentlicht: (2025)
AudioComposer: Towards Fine-grained Audio Generation with Natural Language Descriptions
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2024)
Targeted Augmented Data for Audio Deepfake Detection
von: Astrid, Marcella, et al.
Veröffentlicht: (2024)
von: Astrid, Marcella, et al.
Veröffentlicht: (2024)
Learning Spatially-Aware Language and Audio Embeddings
von: Devnani, Bhavika, et al.
Veröffentlicht: (2024)
von: Devnani, Bhavika, et al.
Veröffentlicht: (2024)
Fast Timing-Conditioned Latent Audio Diffusion
von: Evans, Zach, et al.
Veröffentlicht: (2024)
von: Evans, Zach, et al.
Veröffentlicht: (2024)
Towards Audio Codec-based Speech Separation
von: Yip, Jia Qi, et al.
Veröffentlicht: (2024)
von: Yip, Jia Qi, et al.
Veröffentlicht: (2024)
SNAC: Multi-Scale Neural Audio Codec
von: Siuzdak, Hubert, et al.
Veröffentlicht: (2024)
von: Siuzdak, Hubert, et al.
Veröffentlicht: (2024)
Contrastive Learning from Synthetic Audio Doppelgängers
von: Cherep, Manuel, et al.
Veröffentlicht: (2024)
von: Cherep, Manuel, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Combining audio control and style transfer using latent diffusion
von: Demerlé, Nils, et al.
Veröffentlicht: (2024) -
uaMix-MAE: Efficient Tuning of Pretrained Audio Transformers with Unsupervised Audio Mixtures
von: Tabassum, Afrina, et al.
Veröffentlicht: (2024) -
HEAR: Holistic Evaluation of Audio Representations
von: Turian, Joseph, et al.
Veröffentlicht: (2022) -
CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer
von: Takeuchi, Daiki, et al.
Veröffentlicht: (2025) -
An Unsupervised Domain Adaptation Method for Locating Manipulated Region in partially fake Audio
von: Zeng, Siding, et al.
Veröffentlicht: (2024)