PoDAR: Power-Disentangled Audio Representation for Generative Modeling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Luebs, Alejandro, Vaidya, Mithilesh, Kumar, Ishaan, Badam, Sumukh, Bailey, Stephen W., Bendel, Matthew, Sotelo, Jose, He, Xingzhe |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Deep Learning for Assessment of Oral Reading Fluency
von: Vaidya, Mithilesh, et al.
Veröffentlicht: (2024)
von: Vaidya, Mithilesh, et al.
Veröffentlicht: (2024)
Self-Supervised Multi-View Learning for Disentangled Music Audio Representations
von: Wilkins, Julia, et al.
Veröffentlicht: (2024)
von: Wilkins, Julia, et al.
Veröffentlicht: (2024)
Audio-text Retrieval with Transformer-based Hierarchical Alignment and Disentangled Cross-modal Representation
von: Xin, Yifei, et al.
Veröffentlicht: (2024)
von: Xin, Yifei, et al.
Veröffentlicht: (2024)
Goodbye Drift: Anchored Tree Sampling for Long-Horizon Video-to-Video Generation
von: Bendel, Matthew, et al.
Veröffentlicht: (2026)
von: Bendel, Matthew, et al.
Veröffentlicht: (2026)
Learning Disentangled Audio Representations through Controlled Synthesis
von: Brima, Yusuf, et al.
Veröffentlicht: (2024)
von: Brima, Yusuf, et al.
Veröffentlicht: (2024)
Discrete Audio Representations for Automated Audio Captioning
von: Tian, Jingguang, et al.
Veröffentlicht: (2025)
von: Tian, Jingguang, et al.
Veröffentlicht: (2025)
HASRD: Hierarchical Acoustic and Semantic Representation Disentanglement
von: Hussein, Amir, et al.
Veröffentlicht: (2025)
von: Hussein, Amir, et al.
Veröffentlicht: (2025)
Compositional Audio Representation Learning
von: Sridhar, Sripathi, et al.
Veröffentlicht: (2024)
von: Sridhar, Sripathi, et al.
Veröffentlicht: (2024)
DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
Disentangled Representation Learning for Environment-agnostic Speaker Recognition
von: Nam, KiHyun, et al.
Veröffentlicht: (2024)
von: Nam, KiHyun, et al.
Veröffentlicht: (2024)
Audio Fingerprinting with Holographic Reduced Representations
von: Fujita, Yusuke, et al.
Veröffentlicht: (2024)
von: Fujita, Yusuke, et al.
Veröffentlicht: (2024)
DSpAST: Disentangled Representations for Spatial Audio Reasoning with Large Language Models
von: Wilkinghoff, Kevin, et al.
Veröffentlicht: (2025)
von: Wilkinghoff, Kevin, et al.
Veröffentlicht: (2025)
Learning Multidimensional Disentangled Representations of Instrumental Sounds for Musical Similarity Assessment
von: Hashizume, Yuka, et al.
Veröffentlicht: (2024)
von: Hashizume, Yuka, et al.
Veröffentlicht: (2024)
Balancing Information Preservation and Disentanglement in Self-Supervised Music Representation Learning
von: Wilkins, Julia, et al.
Veröffentlicht: (2025)
von: Wilkins, Julia, et al.
Veröffentlicht: (2025)
Learning Disentangled Speech Representations with Contrastive Learning and Time-Invariant Retrieval
von: Deng, Yimin, et al.
Veröffentlicht: (2024)
von: Deng, Yimin, et al.
Veröffentlicht: (2024)
Geometric Analysis of Speech Representation Spaces: Topological Disentanglement and Confound Detection
von: Kashyap, Bipasha, et al.
Veröffentlicht: (2026)
von: Kashyap, Bipasha, et al.
Veröffentlicht: (2026)
Natural Language Supervision for General-Purpose Audio Representations
von: Elizalde, Benjamin, et al.
Veröffentlicht: (2023)
von: Elizalde, Benjamin, et al.
Veröffentlicht: (2023)
Assessing the Alignment of Audio Representations with Timbre Similarity Ratings
von: Tian, Haokun, et al.
Veröffentlicht: (2025)
von: Tian, Haokun, et al.
Veröffentlicht: (2025)
Zero Shot Audio to Audio Emotion Transfer With Speaker Disentanglement
von: Dutta, Soumya, et al.
Veröffentlicht: (2024)
von: Dutta, Soumya, et al.
Veröffentlicht: (2024)
Quantifying Dimensional Independence in Speech: An Information-Theoretic Framework for Disentangled Representation Learning
von: Kashyap, Bipasha, et al.
Veröffentlicht: (2026)
von: Kashyap, Bipasha, et al.
Veröffentlicht: (2026)
DIVINE: Coordinating Multimodal Disentangled Representations for Oro-Facial Neurological Disorder Assessment
von: Akhtar, Mohd Mujtaba, et al.
Veröffentlicht: (2026)
von: Akhtar, Mohd Mujtaba, et al.
Veröffentlicht: (2026)
Learning Expressive Disentangled Speech Representations with Soft Speech Units and Adversarial Style Augmentation
von: Deng, Yimin, et al.
Veröffentlicht: (2024)
von: Deng, Yimin, et al.
Veröffentlicht: (2024)
SemanticAudio: Audio Generation and Editing in Semantic Space
von: Dai, Zheqi, et al.
Veröffentlicht: (2026)
von: Dai, Zheqi, et al.
Veröffentlicht: (2026)
PAT: Parameter-Free Audio-Text Aligner to Boost Zero-Shot Audio Classification
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions
von: Zhu, Xinfa, et al.
Veröffentlicht: (2025)
von: Zhu, Xinfa, et al.
Veröffentlicht: (2025)
STFTCodec: High-Fidelity Audio Compression through Time-Frequency Domain Representation
von: Feng, Tao, et al.
Veröffentlicht: (2025)
von: Feng, Tao, et al.
Veröffentlicht: (2025)
SONAR: Self-Distilled Continual Pre-training for Domain Adaptive Audio Representation
von: Zhang, Yizhou, et al.
Veröffentlicht: (2025)
von: Zhang, Yizhou, et al.
Veröffentlicht: (2025)
Fx-Encoder++: Extracting Instrument-Wise Audio Effects Representations from Mixtures
von: Yeh, Yen-Tung, et al.
Veröffentlicht: (2025)
von: Yeh, Yen-Tung, et al.
Veröffentlicht: (2025)
Leveraging Self-supervised Audio Representations for Data-Efficient Acoustic Scene Classification
von: Cai, Yiqiang, et al.
Veröffentlicht: (2024)
von: Cai, Yiqiang, et al.
Veröffentlicht: (2024)
Learning Disentangled Speech Representations
von: Brima, Yusuf, et al.
Veröffentlicht: (2023)
von: Brima, Yusuf, et al.
Veröffentlicht: (2023)
Audio Array-Based 3D UAV Trajectory Estimation with LiDAR Pseudo-Labeling
von: Lei, Allen, et al.
Veröffentlicht: (2024)
von: Lei, Allen, et al.
Veröffentlicht: (2024)
Evaluating CNN with Stacked Feature Representations and Audio Spectrogram Transformer Models for Sound Classification
von: Dehaghania, Parinaz Binandeh, et al.
Veröffentlicht: (2026)
von: Dehaghania, Parinaz Binandeh, et al.
Veröffentlicht: (2026)
Audio-Visual Representation Learning via Knowledge Distillation from Speech Foundation Models
von: Zhang, Jing-Xuan, et al.
Veröffentlicht: (2025)
von: Zhang, Jing-Xuan, et al.
Veröffentlicht: (2025)
Rethinking Continual Learning for Speech and Audio: A Representation-Centric Taxonomy and Open Problems
von: Xiao, Yang, et al.
Veröffentlicht: (2026)
von: Xiao, Yang, et al.
Veröffentlicht: (2026)
MoE Adapter for Large Audio Language Models: Sparsity, Disentanglement, and Gradient-Conflict-Free
von: Lei, Yishu, et al.
Veröffentlicht: (2026)
von: Lei, Yishu, et al.
Veröffentlicht: (2026)
Robust Audio Tagging under Class-wise Supervision Unreliability
von: Hou, Yuanbo, et al.
Veröffentlicht: (2026)
von: Hou, Yuanbo, et al.
Veröffentlicht: (2026)
Identifying Hearing Difficulty Moments in Conversational Audio
von: Collins, Jack, et al.
Veröffentlicht: (2025)
von: Collins, Jack, et al.
Veröffentlicht: (2025)
UniAudio: An Audio Foundation Model Toward Universal Audio Generation
von: Yang, Dongchao, et al.
Veröffentlicht: (2023)
von: Yang, Dongchao, et al.
Veröffentlicht: (2023)
Synthetic Audio Forensics Evaluation (SAFE) Challenge
von: Trapeznikov, Kirill, et al.
Veröffentlicht: (2025)
von: Trapeznikov, Kirill, et al.
Veröffentlicht: (2025)
OMAR-RQ: Open Music Audio Representation Model Trained with Multi-Feature Masked Token Prediction
von: Alonso-Jiménez, Pablo, et al.
Veröffentlicht: (2025)
von: Alonso-Jiménez, Pablo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Deep Learning for Assessment of Oral Reading Fluency
von: Vaidya, Mithilesh, et al.
Veröffentlicht: (2024) -
Self-Supervised Multi-View Learning for Disentangled Music Audio Representations
von: Wilkins, Julia, et al.
Veröffentlicht: (2024) -
Audio-text Retrieval with Transformer-based Hierarchical Alignment and Disentangled Cross-modal Representation
von: Xin, Yifei, et al.
Veröffentlicht: (2024) -
Goodbye Drift: Anchored Tree Sampling for Long-Horizon Video-to-Video Generation
von: Bendel, Matthew, et al.
Veröffentlicht: (2026) -
Learning Disentangled Audio Representations through Controlled Synthesis
von: Brima, Yusuf, et al.
Veröffentlicht: (2024)