Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound
Fuente:
arXiv
Saved in:
| Main Authors: | Tjandra, Andros, Wu, Yi-Chiao, Guo, Baishan, Hoffman, John, Ellis, Brian, Vyas, Apoorv, Shi, Bowen, Chen, Sanyuan, Le, Matt, Zacharov, Nick, Wood, Carleigh, Lee, Ann, Hsu, Wei-Ning |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Generative Pre-training for Speech with Flow Matching
by: Liu, Alexander H., et al.
Published: (2023)
by: Liu, Alexander H., et al.
Published: (2023)
Learning Fine-Grained Controllability on Speech Generation via Efficient Fine-Tuning
by: Chien, Chung-Ming, et al.
Published: (2024)
by: Chien, Chung-Ming, et al.
Published: (2024)
SAM Audio Judge: A Unified Multimodal Framework for Perceptual Evaluation of Audio Separation
by: Wang, Helin, et al.
Published: (2026)
by: Wang, Helin, et al.
Published: (2026)
Audiobox TTA-RAG: Improving Zero-Shot and Few-Shot Text-To-Audio with Retrieval-Augmented Generation
by: Yang, Mu, et al.
Published: (2024)
by: Yang, Mu, et al.
Published: (2024)
MusicFlow: Cascaded Flow Matching for Text Guided Music Generation
by: Prajwal, K R, et al.
Published: (2024)
by: Prajwal, K R, et al.
Published: (2024)
MR-FlowDPO: Multi-Reward Direct Preference Optimization for Flow-Matching Text-to-Music Generation
by: Ziv, Alon, et al.
Published: (2025)
by: Ziv, Alon, et al.
Published: (2025)
SAM Audio: Segment Anything in Audio
by: Shi, Bowen, et al.
Published: (2025)
by: Shi, Bowen, et al.
Published: (2025)
The AudioMOS Challenge 2025
by: Huang, Wen-Chin, et al.
Published: (2025)
by: Huang, Wen-Chin, et al.
Published: (2025)
Pushing the Frontier of Audiovisual Perception with Large-Scale Multimodal Correspondence Learning
by: Vyas, Apoorv, et al.
Published: (2025)
by: Vyas, Apoorv, et al.
Published: (2025)
SMILE: Speech Meta In-Context Learning for Low-Resource Language Automatic Speech Recognition
by: Hsu, Ming-Hao, et al.
Published: (2024)
by: Hsu, Ming-Hao, et al.
Published: (2024)
UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions
by: Qiang, Chunyu, et al.
Published: (2026)
by: Qiang, Chunyu, et al.
Published: (2026)
Book Review: Offenders and the Sexual Abuse of Children: Interventions and Limitations
by: Carleigh Bristol Slater
Published: (2025)
by: Carleigh Bristol Slater
Published: (2025)
‘When They See Us’: La televisión como contra-narrativa sobre las experiencias racistas
by: Andros Pineda Machuca
Published: (2020)
by: Andros Pineda Machuca
Published: (2020)
Music for the Bubblegum Set
by: Hoffman, Frank W.
Published: (1975)
by: Hoffman, Frank W.
Published: (1975)
An Order-Complexity Aesthetic Assessment Model for Aesthetic-aware Music Recommendation
by: Jin, Xin, et al.
Published: (2024)
by: Jin, Xin, et al.
Published: (2024)
Multiple regimes in the preferences for redistribution
by: Andros Kourtellos, et al.
Published: (2024)
by: Andros Kourtellos, et al.
Published: (2024)
On the Problem of Text-To-Speech Model Selection for Synthetic Data Generation in Automatic Speech Recognition
by: Rossenbach, Nick, et al.
Published: (2024)
by: Rossenbach, Nick, et al.
Published: (2024)
High order schemes for solving partial differential equations on a quantum computer
by: Arseniev, Boris, et al.
Published: (2024)
by: Arseniev, Boris, et al.
Published: (2024)
Dimetindene—Is the minimum toxic dose for children too strict?
by: Michal Čečrle, et al.
Published: (2024)
by: Michal Čečrle, et al.
Published: (2024)
Affine Gauge Theory: A Diffeomorphism Invariant Gauge Theory of Gravity
by: Tjandra, Kurniawan, et al.
Published: (2025)
by: Tjandra, Kurniawan, et al.
Published: (2025)
Diffeomorphism Invariance and Background Independence
by: Tjandra, Kurniawan, et al.
Published: (2025)
by: Tjandra, Kurniawan, et al.
Published: (2025)
Automatic Speech Recognition (ASR) for the Diagnosis of pronunciation of Speech Sound Disorders in Korean children
by: Ahn, Taekyung, et al.
Published: (2024)
by: Ahn, Taekyung, et al.
Published: (2024)
Unified Semi-Supervised Pipeline for Automatic Speech Recognition
by: Tadevosyan, Nune, et al.
Published: (2025)
by: Tadevosyan, Nune, et al.
Published: (2025)
Dynamic ASR Pathways: An Adaptive Masking Approach Towards Efficient Pruning of A Multilingual ASR Model
by: Xie, Jiamin, et al.
Published: (2023)
by: Xie, Jiamin, et al.
Published: (2023)
Explaining the Musical Advantage in Speech Perception Through Beat Perception and Working Memory
by: Maxime Perron, et al.
Published: (2026)
by: Maxime Perron, et al.
Published: (2026)
MALI-Dev/1DSeaLevelModel_FWTW: v1.0.0-alpha.1-emulation-paper
by: Holly Han, et al.
Published: (2025)
by: Holly Han, et al.
Published: (2025)
Overcoming geographic barriers with telemedicine: Digital transformation strategies for preventing cardiovascular disease in older adults
by: Sanyuan Wei, et al.
Published: (2025)
by: Sanyuan Wei, et al.
Published: (2025)
Sámi Ethnoentomology and the More-than-Musical Aesthetic of Mosquitoes
by: Renzi, Nicola, et al.
Published: (2026)
by: Renzi, Nicola, et al.
Published: (2026)
On the Effect of Purely Synthetic Training Data for Different Automatic Speech Recognition Architectures
by: Hilmes, Benedikt, et al.
Published: (2024)
by: Hilmes, Benedikt, et al.
Published: (2024)
Thinking fast and slow -- a cognitive inspired framework for decision intelligence for power systems
by: Mathur, Apoorv
Published: (2026)
by: Mathur, Apoorv
Published: (2026)
Musical Nuances and the Aesthetic Experience of Popular Music Hooks: Theoretical Considerations and Analytical Approaches
by: Bernhard Steinbrecher
Published: (2021)
by: Bernhard Steinbrecher
Published: (2021)
Musical Nuances and the Aesthetic Experience of Popular Music Hooks: Theoretical Considerations and Analytical Approaches
by: Bernhard Steinbrecher
Published: (2021)
by: Bernhard Steinbrecher
Published: (2021)
Metaheuristic optimization scheme for quantum kernel classifiers using entanglement‐directed graphs
by: Yozef Tjandra, et al.
Published: (2024)
by: Yozef Tjandra, et al.
Published: (2024)
Detecting Musical Deepfakes
by: Sunday, Nick
Published: (2025)
by: Sunday, Nick
Published: (2025)
Exponential bounds for monochromatic sums equal to products
by: Bowen, Matt
Published: (2024)
by: Bowen, Matt
Published: (2024)
Monochromatic non-commuting products
by: Bowen, Matt
Published: (2024)
by: Bowen, Matt
Published: (2024)
‘I Wish I Fought for Myself More Instead of Just Letting Doctors Dismiss Me’: A Combined Qualitative Analysis of Four Cohorts of Aotearoa New Zealand Endometriosis Patients
by: Katherine Ellis, et al.
Published: (2025)
by: Katherine Ellis, et al.
Published: (2025)
RepVF: A Unified Vector Fields Representation for Multi-task 3D Perception
by: Li, Chunliang, et al.
Published: (2024)
by: Li, Chunliang, et al.
Published: (2024)
Controlling Whisper: Universal Acoustic Adversarial Attacks to Control Speech Foundation Models
by: Raina, Vyas, et al.
Published: (2024)
by: Raina, Vyas, et al.
Published: (2024)
The ICASSP 2026 Automatic Song Aesthetics Evaluation Challenge
by: Ma, Guobin, et al.
Published: (2026)
by: Ma, Guobin, et al.
Published: (2026)
Similar Items
-
Generative Pre-training for Speech with Flow Matching
by: Liu, Alexander H., et al.
Published: (2023) -
Learning Fine-Grained Controllability on Speech Generation via Efficient Fine-Tuning
by: Chien, Chung-Ming, et al.
Published: (2024) -
SAM Audio Judge: A Unified Multimodal Framework for Perceptual Evaluation of Audio Separation
by: Wang, Helin, et al.
Published: (2026) -
Audiobox TTA-RAG: Improving Zero-Shot and Few-Shot Text-To-Audio with Retrieval-Augmented Generation
by: Yang, Mu, et al.
Published: (2024) -
MusicFlow: Cascaded Flow Matching for Text Guided Music Generation
by: Prajwal, K R, et al.
Published: (2024)