Putting HUMANS first: Efficient LAM Evaluation with Human Preference Alignment
Fuente:
arXiv
Guardado en:
| Autores principales: | Gan, Woody Haosheng, Held, William, Yang, Diyi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Scaling Open Discrete Audio Foundation Models with Interleaved Semantic, Acoustic, and Text Tokens
por: Manakul, Potsawee, et al.
Publicado: (2026)
por: Manakul, Potsawee, et al.
Publicado: (2026)
AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation
por: Manakul, Potsawee, et al.
Publicado: (2025)
por: Manakul, Potsawee, et al.
Publicado: (2025)
GrowLoop: Self-Evolving Conversation Evaluation Seeded by Human
por: Lin, Yihang, et al.
Publicado: (2026)
por: Lin, Yihang, et al.
Publicado: (2026)
Towards Fine-Grained Code-Switch Speech Translation with Semantic Space Alignment
por: Gao, Yan, et al.
Publicado: (2025)
por: Gao, Yan, et al.
Publicado: (2025)
Interactive ASR: Towards Human-Like Interaction and Semantic Coherence Evaluation for Agentic Speech Recognition
por: Wang, Peng, et al.
Publicado: (2026)
por: Wang, Peng, et al.
Publicado: (2026)
Soundwave: Less is More for Speech-Text Alignment in LLMs
por: Zhang, Yuhao, et al.
Publicado: (2025)
por: Zhang, Yuhao, et al.
Publicado: (2025)
Data-efficient Targeted Token-level Preference Optimization for LLM-based Text-to-Speech
por: Kotoge, Rikuto, et al.
Publicado: (2025)
por: Kotoge, Rikuto, et al.
Publicado: (2025)
Emotion-Aligned Generation in Diffusion Text to Speech Models via Preference-Guided Optimization
por: Shi, Jiacheng, et al.
Publicado: (2025)
por: Shi, Jiacheng, et al.
Publicado: (2025)
Efficient Training for Cross-lingual Speech Language Models
por: Zhou, Yan, et al.
Publicado: (2026)
por: Zhou, Yan, et al.
Publicado: (2026)
EchoDistill:Alignment Noisy-to-Clean Self-Distillation for Robust Audio LLMs
por: Lin, Liang, et al.
Publicado: (2026)
por: Lin, Liang, et al.
Publicado: (2026)
WESR: Scaling and Evaluating Word-level Event-Speech Recognition
por: Yang, Chenchen, et al.
Publicado: (2026)
por: Yang, Chenchen, et al.
Publicado: (2026)
SpeechJudge: Towards Human-Level Judgment for Speech Naturalness
por: Zhang, Xueyao, et al.
Publicado: (2025)
por: Zhang, Xueyao, et al.
Publicado: (2025)
A Critical Review of the Need for Knowledge-Centric Evaluation of Quranic Recitation
por: Al-Kharusi, Mohammed Hilal, et al.
Publicado: (2025)
por: Al-Kharusi, Mohammed Hilal, et al.
Publicado: (2025)
Language Family Matters: Evaluating LLM-Based ASR Across Linguistic Boundaries
por: Zhang, Yuchen, et al.
Publicado: (2026)
por: Zhang, Yuchen, et al.
Publicado: (2026)
CAF-Score: Calibrating CLAP with LALMs for Reference-free Audio Captioning Evaluation
por: Lee, Insung, et al.
Publicado: (2026)
por: Lee, Insung, et al.
Publicado: (2026)
Finding My Voice: Generative Reconstruction of Disordered Speech for Automated Clinical Evaluation
por: Rosero, Karen, et al.
Publicado: (2025)
por: Rosero, Karen, et al.
Publicado: (2025)
ASKD-Whisper: Adaptive Self-knowledge Distillation for Efficient and Low-Latency Automatic Speech Recognition
por: Lee, Junseok, et al.
Publicado: (2026)
por: Lee, Junseok, et al.
Publicado: (2026)
Still Between Us? Evaluating and Improving Voice Assistant Robustness to Third-Party Interruptions
por: Lee, Dongwook, et al.
Publicado: (2026)
por: Lee, Dongwook, et al.
Publicado: (2026)
VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents
por: Wu, Weihao, et al.
Publicado: (2025)
por: Wu, Weihao, et al.
Publicado: (2025)
Beyond Prompting: Efficient and Robust Contextual Biasing for Speech LLMs via Logit-Space Integration (LOGIC)
por: Wang, Peidong
Publicado: (2026)
por: Wang, Peidong
Publicado: (2026)
Audio Jailbreaks in Large Audio-Language Models: Taxonomy, Attack-Defense Analysis, and Cost-Aware Evaluation
por: Feng, Bo-Han, et al.
Publicado: (2026)
por: Feng, Bo-Han, et al.
Publicado: (2026)
Length-Aware Rotary Position Embedding for Text-Speech Alignment
por: Kim, Hyeongju, et al.
Publicado: (2025)
por: Kim, Hyeongju, et al.
Publicado: (2025)
Mind the Gap! Static and Interactive Evaluations of Large Audio Models
por: Li, Minzhi, et al.
Publicado: (2025)
por: Li, Minzhi, et al.
Publicado: (2025)
Assessing the Ability of Neural TTS Systems to Model Consonant-Induced F0 Perturbation
por: Yang, Tianle, et al.
Publicado: (2026)
por: Yang, Tianle, et al.
Publicado: (2026)
MOSS-VoiceGenerator: Create Realistic Voices with Natural Language Descriptions
por: Huang, Kexin, et al.
Publicado: (2026)
por: Huang, Kexin, et al.
Publicado: (2026)
Decoding the Ear: A Framework for Objectifying Expressiveness from Human Preference Through Efficient Alignment
por: Lin, Zhiyu, et al.
Publicado: (2025)
por: Lin, Zhiyu, et al.
Publicado: (2025)
MOSS-TTS Technical Report
por: Gong, Yitian, et al.
Publicado: (2026)
por: Gong, Yitian, et al.
Publicado: (2026)
No Verifiable Reward for Prosody: Toward Preference-Guided Prosody Learning in TTS
por: Shin, Seungyoun, et al.
Publicado: (2025)
por: Shin, Seungyoun, et al.
Publicado: (2025)
MOSS-TTSD: Text to Spoken Dialogue Generation
por: Zhang, Yuqian, et al.
Publicado: (2026)
por: Zhang, Yuqian, et al.
Publicado: (2026)
From Scores to Preferences: Redefining MOS Benchmarking for Speech Quality Reward Modeling
por: Cao, Yifei, et al.
Publicado: (2025)
por: Cao, Yifei, et al.
Publicado: (2025)
Seamless Dysfluent Speech Text Alignment for Disordered Speech Analysis
por: Ye, Zongli, et al.
Publicado: (2025)
por: Ye, Zongli, et al.
Publicado: (2025)
Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference Optimization
por: Majumder, Navonil, et al.
Publicado: (2024)
por: Majumder, Navonil, et al.
Publicado: (2024)
Abusive music and song transformation using GenAI and LLMs
por: Choi, Jiyang, et al.
Publicado: (2026)
por: Choi, Jiyang, et al.
Publicado: (2026)
A novel LSTM music generator based on the fractional time-frequency feature extraction
por: Ya, Li, et al.
Publicado: (2026)
por: Ya, Li, et al.
Publicado: (2026)
Do LLM Decoders Listen Fairly? Benchmarking How Language Model Priors Shape Bias in Speech Recognition
por: Ginjala, Srishti, et al.
Publicado: (2026)
por: Ginjala, Srishti, et al.
Publicado: (2026)
MedMosaic: A Challenging Large Scale Benchmark of Diverse Medical Audio
por: Rajgarhia, Harshit, et al.
Publicado: (2026)
por: Rajgarhia, Harshit, et al.
Publicado: (2026)
Raon-Speech Technical Report
por: Kim, Beomsoo, et al.
Publicado: (2026)
por: Kim, Beomsoo, et al.
Publicado: (2026)
Prosody-Guided Harmonic Attention for Phase-Coherent Neural Vocoding in the Complex Spectrum
por: Al-Radhi, Mohammed Salah, et al.
Publicado: (2026)
por: Al-Radhi, Mohammed Salah, et al.
Publicado: (2026)
VorTEX: Various overlap ratio for Target speech EXtraction
por: Oh, Ro-hoon, et al.
Publicado: (2026)
por: Oh, Ro-hoon, et al.
Publicado: (2026)
Synchronization and Turn-Taking in Full-Duplex Speech Dialogue Models
por: Riera, Pablo, et al.
Publicado: (2026)
por: Riera, Pablo, et al.
Publicado: (2026)
Ejemplares similares
-
Scaling Open Discrete Audio Foundation Models with Interleaved Semantic, Acoustic, and Text Tokens
por: Manakul, Potsawee, et al.
Publicado: (2026) -
AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation
por: Manakul, Potsawee, et al.
Publicado: (2025) -
GrowLoop: Self-Evolving Conversation Evaluation Seeded by Human
por: Lin, Yihang, et al.
Publicado: (2026) -
Towards Fine-Grained Code-Switch Speech Translation with Semantic Space Alignment
por: Gao, Yan, et al.
Publicado: (2025) -
Interactive ASR: Towards Human-Like Interaction and Semantic Coherence Evaluation for Agentic Speech Recognition
por: Wang, Peng, et al.
Publicado: (2026)