Decoding the Ear: A Framework for Objectifying Expressiveness from Human Preference Through Efficient Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Zhiyu, Yang, Jingwen, Zhao, Jiale, Liu, Meng, Li, Sunzhu, Wang, Benyou |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization
by: Banerjee, Adhiraj, et al.
Published: (2026)
by: Banerjee, Adhiraj, et al.
Published: (2026)
Style Mixture of Experts for Expressive Text-To-Speech Synthesis
by: Jawaid, Ahad, et al.
Published: (2024)
by: Jawaid, Ahad, et al.
Published: (2024)
Audio MultiChallenge: A Multi-Turn Evaluation of Spoken Dialogue Systems on Natural Human Interaction
by: Gosai, Advait, et al.
Published: (2025)
by: Gosai, Advait, et al.
Published: (2025)
Putting HUMANS first: Efficient LAM Evaluation with Human Preference Alignment
by: Gan, Woody Haosheng, et al.
Published: (2026)
by: Gan, Woody Haosheng, et al.
Published: (2026)
U-Codec: Ultra Low Frame-rate Neural Speech Codec for Fast High-fidelity Speech Generation
by: Yang, Xusheng, et al.
Published: (2025)
by: Yang, Xusheng, et al.
Published: (2025)
GTR-Voice: Articulatory Phonetics Informed Controllable Expressive Speech Synthesis
by: Li, Zehua Kcriss, et al.
Published: (2024)
by: Li, Zehua Kcriss, et al.
Published: (2024)
Rasa: Building Expressive Speech Synthesis Systems for Indian Languages in Low-resource Settings
by: Varadhan, Praveen Srinivasa, et al.
Published: (2024)
by: Varadhan, Praveen Srinivasa, et al.
Published: (2024)
Label-Looping: Highly Efficient Decoding for Transducers
by: Bataev, Vladimir, et al.
Published: (2024)
by: Bataev, Vladimir, et al.
Published: (2024)
MambaVoiceCloning: Efficient and Expressive Text-to-Speech via State-Space Modeling and Diffusion Control
by: Kumar, Sahil, et al.
Published: (2026)
by: Kumar, Sahil, et al.
Published: (2026)
Assessing Factual Music Comprehension in Large Audio Language Models
by: Lin, Daniel Chenyu, et al.
Published: (2025)
by: Lin, Daniel Chenyu, et al.
Published: (2025)
Benchmarking Music Generation Models and Metrics via Human Preference Studies
by: Grötschla, Florian, et al.
Published: (2025)
by: Grötschla, Florian, et al.
Published: (2025)
Task Vector in TTS: Toward Emotionally Expressive Dialectal Speech Synthesis
by: Feng, Pengchao, et al.
Published: (2025)
by: Feng, Pengchao, et al.
Published: (2025)
Efficient Training of Self-Supervised Speech Foundation Models on a Compute Budget
by: Liu, Andy T., et al.
Published: (2024)
by: Liu, Andy T., et al.
Published: (2024)
Soundwave: Less is More for Speech-Text Alignment in LLMs
by: Zhang, Yuhao, et al.
Published: (2025)
by: Zhang, Yuhao, et al.
Published: (2025)
Semantic Codebooks as Effective Priors for Neural Speech Compression
by: Bai, Liuyang, et al.
Published: (2025)
by: Bai, Liuyang, et al.
Published: (2025)
UALM: Unified Audio Language Model for Understanding, Generation and Reasoning
by: Tian, Jinchuan, et al.
Published: (2025)
by: Tian, Jinchuan, et al.
Published: (2025)
WEE-Therapy: A Mixture of Weak Encoders Framework for Psychological Counseling Dialogue Analysis
by: Kang, Yongqi, et al.
Published: (2025)
by: Kang, Yongqi, et al.
Published: (2025)
CAARMA: Class Augmentation with Adversarial Mixup Regularization
by: Baali, Massa, et al.
Published: (2025)
by: Baali, Massa, et al.
Published: (2025)
Guiding Frame-Level CTC Alignments Using Self-knowledge Distillation
by: Kim, Eungbeom, et al.
Published: (2024)
by: Kim, Eungbeom, et al.
Published: (2024)
Early Attentive Sparsification Accelerates Neural Speech Transcription
by: Xu, Zifei, et al.
Published: (2025)
by: Xu, Zifei, et al.
Published: (2025)
Speech-Worthy Alignment for Japanese SpeechLLMs via Direct Preference Optimization
by: Zhao, Mengjie, et al.
Published: (2026)
by: Zhao, Mengjie, et al.
Published: (2026)
Efficient Adapter Finetuning for Tail Languages in Streaming Multilingual ASR
by: Bai, Junwen, et al.
Published: (2024)
by: Bai, Junwen, et al.
Published: (2024)
Towards Early Prediction of Self-Supervised Speech Model Performance
by: Whetten, Ryan, et al.
Published: (2025)
by: Whetten, Ryan, et al.
Published: (2025)
Learning When to Think While Listening in Large Audio-Language Models
by: Song, Zhiyuan, et al.
Published: (2026)
by: Song, Zhiyuan, et al.
Published: (2026)
RUMAA: Repeat-Aware Unified Music Audio Analysis for Score-Performance Alignment, Transcription, and Mistake Detection
by: Chang, Sungkyun, et al.
Published: (2025)
by: Chang, Sungkyun, et al.
Published: (2025)
AlignAtt: Using Attention-based Audio-Translation Alignments as a Guide for Simultaneous Speech Translation
by: Papi, Sara, et al.
Published: (2023)
by: Papi, Sara, et al.
Published: (2023)
Moonshine v2: Ergodic Streaming Encoder ASR for Latency-Critical Speech Applications
by: Kudlur, Manjunath, et al.
Published: (2026)
by: Kudlur, Manjunath, et al.
Published: (2026)
Flavors of Moonshine: Tiny Specialized ASR Models for Edge Devices
by: King, Evan, et al.
Published: (2025)
by: King, Evan, et al.
Published: (2025)
WavLink: Compact Audio-Text Embeddings with a Global Whisper Token
by: Kumar, Gokul Karthik, et al.
Published: (2026)
by: Kumar, Gokul Karthik, et al.
Published: (2026)
Large Language Model Data Generation for Enhanced Intent Recognition in German Speech
by: Rosin, Theresa Pekarek, et al.
Published: (2025)
by: Rosin, Theresa Pekarek, et al.
Published: (2025)
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025
by: Ferreira, Alef Iury Siqueira, et al.
Published: (2025)
by: Ferreira, Alef Iury Siqueira, et al.
Published: (2025)
Huntington Disease Automatic Speech Recognition with Biomarker Supervision
by: Wang, Charles L., et al.
Published: (2026)
by: Wang, Charles L., et al.
Published: (2026)
Investigation for Relative Voice Impression Estimation
by: Fujita, Kenichi, et al.
Published: (2026)
by: Fujita, Kenichi, et al.
Published: (2026)
RO-N3WS: Enhancing Generalization in Low-Resource ASR with Diverse Romanian Speech Benchmarks
by: Diaconu, Alexandra, et al.
Published: (2026)
by: Diaconu, Alexandra, et al.
Published: (2026)
EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents
by: Bogavelli, Tara, et al.
Published: (2026)
by: Bogavelli, Tara, et al.
Published: (2026)
From Alignment to Advancement: Bootstrapping Audio-Language Alignment with Synthetic Data
by: Kuan, Chun-Yi, et al.
Published: (2025)
by: Kuan, Chun-Yi, et al.
Published: (2025)
Spontaneous Style Text-to-Speech Synthesis with Controllable Spontaneous Behaviors Based on Language Models
by: Li, Weiqin, et al.
Published: (2024)
by: Li, Weiqin, et al.
Published: (2024)
Aligning Audio Captions with Human Preferences
by: Hegde, Kartik, et al.
Published: (2025)
by: Hegde, Kartik, et al.
Published: (2025)
HyperTTS: Parameter Efficient Adaptation in Text to Speech using Hypernetworks
by: Li, Yingting, et al.
Published: (2024)
by: Li, Yingting, et al.
Published: (2024)
Competitive Audio-Language Models with Data-Efficient Single-Stage Training on Public Data
by: Kumar, Gokul Karthik, et al.
Published: (2025)
by: Kumar, Gokul Karthik, et al.
Published: (2025)
Similar Items
-
PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization
by: Banerjee, Adhiraj, et al.
Published: (2026) -
Style Mixture of Experts for Expressive Text-To-Speech Synthesis
by: Jawaid, Ahad, et al.
Published: (2024) -
Audio MultiChallenge: A Multi-Turn Evaluation of Spoken Dialogue Systems on Natural Human Interaction
by: Gosai, Advait, et al.
Published: (2025) -
Putting HUMANS first: Efficient LAM Evaluation with Human Preference Alignment
by: Gan, Woody Haosheng, et al.
Published: (2026) -
U-Codec: Ultra Low Frame-rate Neural Speech Codec for Fast High-fidelity Speech Generation
by: Yang, Xusheng, et al.
Published: (2025)