Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Tjandra, Andros, Wu, Yi-Chiao, Guo, Baishan, Hoffman, John, Ellis, Brian, Vyas, Apoorv, Shi, Bowen, Chen, Sanyuan, Le, Matt, Zacharov, Nick, Wood, Carleigh, Lee, Ann, Hsu, Wei-Ning |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Generative Pre-training for Speech with Flow Matching
par: Liu, Alexander H., et autres
Publié: (2023)
par: Liu, Alexander H., et autres
Publié: (2023)
Learning Fine-Grained Controllability on Speech Generation via Efficient Fine-Tuning
par: Chien, Chung-Ming, et autres
Publié: (2024)
par: Chien, Chung-Ming, et autres
Publié: (2024)
SAM Audio Judge: A Unified Multimodal Framework for Perceptual Evaluation of Audio Separation
par: Wang, Helin, et autres
Publié: (2026)
par: Wang, Helin, et autres
Publié: (2026)
Audiobox TTA-RAG: Improving Zero-Shot and Few-Shot Text-To-Audio with Retrieval-Augmented Generation
par: Yang, Mu, et autres
Publié: (2024)
par: Yang, Mu, et autres
Publié: (2024)
MusicFlow: Cascaded Flow Matching for Text Guided Music Generation
par: Prajwal, K R, et autres
Publié: (2024)
par: Prajwal, K R, et autres
Publié: (2024)
MR-FlowDPO: Multi-Reward Direct Preference Optimization for Flow-Matching Text-to-Music Generation
par: Ziv, Alon, et autres
Publié: (2025)
par: Ziv, Alon, et autres
Publié: (2025)
SAM Audio: Segment Anything in Audio
par: Shi, Bowen, et autres
Publié: (2025)
par: Shi, Bowen, et autres
Publié: (2025)
The AudioMOS Challenge 2025
par: Huang, Wen-Chin, et autres
Publié: (2025)
par: Huang, Wen-Chin, et autres
Publié: (2025)
Pushing the Frontier of Audiovisual Perception with Large-Scale Multimodal Correspondence Learning
par: Vyas, Apoorv, et autres
Publié: (2025)
par: Vyas, Apoorv, et autres
Publié: (2025)
SMILE: Speech Meta In-Context Learning for Low-Resource Language Automatic Speech Recognition
par: Hsu, Ming-Hao, et autres
Publié: (2024)
par: Hsu, Ming-Hao, et autres
Publié: (2024)
UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions
par: Qiang, Chunyu, et autres
Publié: (2026)
par: Qiang, Chunyu, et autres
Publié: (2026)
Book Review: Offenders and the Sexual Abuse of Children: Interventions and Limitations
par: Carleigh Bristol Slater
Publié: (2025)
par: Carleigh Bristol Slater
Publié: (2025)
‘When They See Us’: La televisión como contra-narrativa sobre las experiencias racistas
par: Andros Pineda Machuca
Publié: (2020)
par: Andros Pineda Machuca
Publié: (2020)
Music for the Bubblegum Set
par: Hoffman, Frank W.
Publié: (1975)
par: Hoffman, Frank W.
Publié: (1975)
An Order-Complexity Aesthetic Assessment Model for Aesthetic-aware Music Recommendation
par: Jin, Xin, et autres
Publié: (2024)
par: Jin, Xin, et autres
Publié: (2024)
Multiple regimes in the preferences for redistribution
par: Andros Kourtellos, et autres
Publié: (2024)
par: Andros Kourtellos, et autres
Publié: (2024)
On the Problem of Text-To-Speech Model Selection for Synthetic Data Generation in Automatic Speech Recognition
par: Rossenbach, Nick, et autres
Publié: (2024)
par: Rossenbach, Nick, et autres
Publié: (2024)
High order schemes for solving partial differential equations on a quantum computer
par: Arseniev, Boris, et autres
Publié: (2024)
par: Arseniev, Boris, et autres
Publié: (2024)
Dimetindene—Is the minimum toxic dose for children too strict?
par: Michal Čečrle, et autres
Publié: (2024)
par: Michal Čečrle, et autres
Publié: (2024)
Affine Gauge Theory: A Diffeomorphism Invariant Gauge Theory of Gravity
par: Tjandra, Kurniawan, et autres
Publié: (2025)
par: Tjandra, Kurniawan, et autres
Publié: (2025)
Diffeomorphism Invariance and Background Independence
par: Tjandra, Kurniawan, et autres
Publié: (2025)
par: Tjandra, Kurniawan, et autres
Publié: (2025)
Automatic Speech Recognition (ASR) for the Diagnosis of pronunciation of Speech Sound Disorders in Korean children
par: Ahn, Taekyung, et autres
Publié: (2024)
par: Ahn, Taekyung, et autres
Publié: (2024)
Unified Semi-Supervised Pipeline for Automatic Speech Recognition
par: Tadevosyan, Nune, et autres
Publié: (2025)
par: Tadevosyan, Nune, et autres
Publié: (2025)
Dynamic ASR Pathways: An Adaptive Masking Approach Towards Efficient Pruning of A Multilingual ASR Model
par: Xie, Jiamin, et autres
Publié: (2023)
par: Xie, Jiamin, et autres
Publié: (2023)
Explaining the Musical Advantage in Speech Perception Through Beat Perception and Working Memory
par: Maxime Perron, et autres
Publié: (2026)
par: Maxime Perron, et autres
Publié: (2026)
MALI-Dev/1DSeaLevelModel_FWTW: v1.0.0-alpha.1-emulation-paper
par: Holly Han, et autres
Publié: (2025)
par: Holly Han, et autres
Publié: (2025)
Overcoming geographic barriers with telemedicine: Digital transformation strategies for preventing cardiovascular disease in older adults
par: Sanyuan Wei, et autres
Publié: (2025)
par: Sanyuan Wei, et autres
Publié: (2025)
Sámi Ethnoentomology and the More-than-Musical Aesthetic of Mosquitoes
par: Renzi, Nicola, et autres
Publié: (2026)
par: Renzi, Nicola, et autres
Publié: (2026)
On the Effect of Purely Synthetic Training Data for Different Automatic Speech Recognition Architectures
par: Hilmes, Benedikt, et autres
Publié: (2024)
par: Hilmes, Benedikt, et autres
Publié: (2024)
Thinking fast and slow -- a cognitive inspired framework for decision intelligence for power systems
par: Mathur, Apoorv
Publié: (2026)
par: Mathur, Apoorv
Publié: (2026)
Musical Nuances and the Aesthetic Experience of Popular Music Hooks: Theoretical Considerations and Analytical Approaches
par: Bernhard Steinbrecher
Publié: (2021)
par: Bernhard Steinbrecher
Publié: (2021)
Musical Nuances and the Aesthetic Experience of Popular Music Hooks: Theoretical Considerations and Analytical Approaches
par: Bernhard Steinbrecher
Publié: (2021)
par: Bernhard Steinbrecher
Publié: (2021)
Metaheuristic optimization scheme for quantum kernel classifiers using entanglement‐directed graphs
par: Yozef Tjandra, et autres
Publié: (2024)
par: Yozef Tjandra, et autres
Publié: (2024)
Detecting Musical Deepfakes
par: Sunday, Nick
Publié: (2025)
par: Sunday, Nick
Publié: (2025)
Exponential bounds for monochromatic sums equal to products
par: Bowen, Matt
Publié: (2024)
par: Bowen, Matt
Publié: (2024)
Monochromatic non-commuting products
par: Bowen, Matt
Publié: (2024)
par: Bowen, Matt
Publié: (2024)
‘I Wish I Fought for Myself More Instead of Just Letting Doctors Dismiss Me’: A Combined Qualitative Analysis of Four Cohorts of Aotearoa New Zealand Endometriosis Patients
par: Katherine Ellis, et autres
Publié: (2025)
par: Katherine Ellis, et autres
Publié: (2025)
RepVF: A Unified Vector Fields Representation for Multi-task 3D Perception
par: Li, Chunliang, et autres
Publié: (2024)
par: Li, Chunliang, et autres
Publié: (2024)
Controlling Whisper: Universal Acoustic Adversarial Attacks to Control Speech Foundation Models
par: Raina, Vyas, et autres
Publié: (2024)
par: Raina, Vyas, et autres
Publié: (2024)
The ICASSP 2026 Automatic Song Aesthetics Evaluation Challenge
par: Ma, Guobin, et autres
Publié: (2026)
par: Ma, Guobin, et autres
Publié: (2026)
Documents similaires
-
Generative Pre-training for Speech with Flow Matching
par: Liu, Alexander H., et autres
Publié: (2023) -
Learning Fine-Grained Controllability on Speech Generation via Efficient Fine-Tuning
par: Chien, Chung-Ming, et autres
Publié: (2024) -
SAM Audio Judge: A Unified Multimodal Framework for Perceptual Evaluation of Audio Separation
par: Wang, Helin, et autres
Publié: (2026) -
Audiobox TTA-RAG: Improving Zero-Shot and Few-Shot Text-To-Audio with Retrieval-Augmented Generation
par: Yang, Mu, et autres
Publié: (2024) -
MusicFlow: Cascaded Flow Matching for Text Guided Music Generation
par: Prajwal, K R, et autres
Publié: (2024)