Do Audio LLMs Really LISTEN, or Just Transcribe? Measuring Lexical vs. Acoustic Emotion Cues Reliance
Fuente:
arXiv
Salvato in:
| Autori principali: | Chen, Jingyi, Guo, Zhimeng, Chun, Jiyun, Wang, Pichao, Perrault, Andrew, Elsner, Micha |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Fine-Tuning Text-to-Speech Diffusion Models Using Reinforcement Learning with Human Feedback
di: Chen, Jingyi, et al.
Pubblicazione: (2025)
di: Chen, Jingyi, et al.
Pubblicazione: (2025)
Shortcomings of LLMs for Low-Resource Translation: Retrieval and Understanding are Both the Problem
di: Court, Sara, et al.
Pubblicazione: (2024)
di: Court, Sara, et al.
Pubblicazione: (2024)
DLPO: Diffusion Model Loss-Guided Reinforcement Learning for Fine-Tuning Text-to-Speech Diffusion Models
di: Chen, Jingyi, et al.
Pubblicazione: (2024)
di: Chen, Jingyi, et al.
Pubblicazione: (2024)
Prompt and circumstance: A word-by-word LLM prompting approach to interlinear glossing for low-resource languages
di: Elsner, Micha, et al.
Pubblicazione: (2025)
di: Elsner, Micha, et al.
Pubblicazione: (2025)
Beyond Length: Context-Aware Expansion and Independence as Developmentally Sensitive Evaluation in Child Utterances
di: Chun, Jiyun, et al.
Pubblicazione: (2026)
di: Chun, Jiyun, et al.
Pubblicazione: (2026)
ARES: Alternating Reinforcement Learning and Supervised Fine-Tuning for Enhanced Multi-Modal Chain-of-Thought Reasoning Through Diverse AI Feedback
di: Byun, Ju-Seung, et al.
Pubblicazione: (2024)
di: Byun, Ju-Seung, et al.
Pubblicazione: (2024)
Why is "Chicago" Predictive of Deceptive Reviews? Using LLMs to Discover Language Phenomena from Lexical Cues
di: Qu, Jiaming, et al.
Pubblicazione: (2025)
di: Qu, Jiaming, et al.
Pubblicazione: (2025)
Speech LLMs are Contextual Reasoning Transcribers
di: Deng, Keqi, et al.
Pubblicazione: (2026)
di: Deng, Keqi, et al.
Pubblicazione: (2026)
Do LLMs Really Memorize Personally Identifiable Information? Revisiting PII Leakage with a Cue-Controlled Memorization Framework
di: Luo, Xiaoyu, et al.
Pubblicazione: (2026)
di: Luo, Xiaoyu, et al.
Pubblicazione: (2026)
LISTEN to Your Preferences: An LLM Framework for Multi-Objective Selection
di: Jovine, Adam S., et al.
Pubblicazione: (2025)
di: Jovine, Adam S., et al.
Pubblicazione: (2025)
Do LLMs Really Adapt to Domains? An Ontology Learning Perspective
di: Mai, Huu Tan, et al.
Pubblicazione: (2024)
di: Mai, Huu Tan, et al.
Pubblicazione: (2024)
Do LLMs Really Think Step-by-step In Implicit Reasoning?
di: Yu, Yijiong
Pubblicazione: (2024)
di: Yu, Yijiong
Pubblicazione: (2024)
The Role of Prosodic and Lexical Cues in Turn-Taking with Self-Supervised Speech Representations
di: Russell, Sam OConnor, et al.
Pubblicazione: (2026)
di: Russell, Sam OConnor, et al.
Pubblicazione: (2026)
All That Glitters Is Not Audio: Rethinking Text Priors and Audio Reliance in Audio-Language Evaluation
di: Foo, Leonardo Haw-Yang, et al.
Pubblicazione: (2026)
di: Foo, Leonardo Haw-Yang, et al.
Pubblicazione: (2026)
Do Emotions Really Affect Argument Convincingness? A Dynamic Approach with LLM-based Manipulation Checks
di: Chen, Yanran, et al.
Pubblicazione: (2025)
di: Chen, Yanran, et al.
Pubblicazione: (2025)
Acquiring Pronunciation Knowledge from Transcribed Speech Audio via Multi-task Learning
di: Sun, Siqi, et al.
Pubblicazione: (2024)
di: Sun, Siqi, et al.
Pubblicazione: (2024)
Do LLMs Really Forget? Evaluating Unlearning with Knowledge Correlation and Confidence Awareness
di: Wei, Rongzhe, et al.
Pubblicazione: (2025)
di: Wei, Rongzhe, et al.
Pubblicazione: (2025)
Dispersion Measures as Predictors of Lexical Decision Time, Word Familiarity, and Lexical Complexity
di: Nohejl, Adam, et al.
Pubblicazione: (2025)
di: Nohejl, Adam, et al.
Pubblicazione: (2025)
Do MLLMs Really Understand the Charts?
di: Zhang, Xiao, et al.
Pubblicazione: (2025)
di: Zhang, Xiao, et al.
Pubblicazione: (2025)
Counterfactual Cultural Cues Reduce Medical QA Accuracy in LLMs: Identifier vs Context Effects
di: Rezaei, Amirhossein Haji Mohammad, et al.
Pubblicazione: (2026)
di: Rezaei, Amirhossein Haji Mohammad, et al.
Pubblicazione: (2026)
Do LLMs Know What Luxembourgish Borrows? Probing Lexical Neology in Low-Resource Multilingual Models
di: Hosseini-Kivanani, Nina
Pubblicazione: (2026)
di: Hosseini-Kivanani, Nina
Pubblicazione: (2026)
GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
di: Chen, Guoguo, et al.
Pubblicazione: (2021)
di: Chen, Guoguo, et al.
Pubblicazione: (2021)
On the Contribution of Lexical Features to Speech Emotion Recognition
di: Combei, David
Pubblicazione: (2025)
di: Combei, David
Pubblicazione: (2025)
Parameters vs. Context: Fine-Grained Control of Knowledge Reliance in Language Models
di: Bi, Baolong, et al.
Pubblicazione: (2025)
di: Bi, Baolong, et al.
Pubblicazione: (2025)
Do the Right Thing, Just Debias! Multi-Category Bias Mitigation Using LLMs
di: Roy, Amartya, et al.
Pubblicazione: (2024)
di: Roy, Amartya, et al.
Pubblicazione: (2024)
Rubato: Transcribing Piano Music with Timestamps
di: Tamer, Nazif Can, et al.
Pubblicazione: (2026)
di: Tamer, Nazif Can, et al.
Pubblicazione: (2026)
Meta-Tuning LLMs to Leverage Lexical Knowledge for Generalizable Language Style Understanding
di: Guo, Ruohao, et al.
Pubblicazione: (2023)
di: Guo, Ruohao, et al.
Pubblicazione: (2023)
Are LLMs Really Not Knowledgeable? Mining the Submerged Knowledge in LLMs' Memory
di: Tao, Xingjian, et al.
Pubblicazione: (2024)
di: Tao, Xingjian, et al.
Pubblicazione: (2024)
MSR-86K: An Evolving, Multilingual Corpus with 86,300 Hours of Transcribed Audio for Speech Recognition Research
di: Li, Song, et al.
Pubblicazione: (2024)
di: Li, Song, et al.
Pubblicazione: (2024)
Do LLMs "Feel"? Emotion Circuits Discovery and Control
di: Wang, Chenxi, et al.
Pubblicazione: (2025)
di: Wang, Chenxi, et al.
Pubblicazione: (2025)
HearSay Benchmark: Do Audio LLMs Leak What They Hear?
di: Wang, Jin, et al.
Pubblicazione: (2026)
di: Wang, Jin, et al.
Pubblicazione: (2026)
Fake Alignment: Are LLMs Really Aligned Well?
di: Wang, Yixu, et al.
Pubblicazione: (2023)
di: Wang, Yixu, et al.
Pubblicazione: (2023)
Resolving Transcription Ambiguity in Spanish: A Hybrid Acoustic-Lexical System for Punctuation Restoration
di: Zhu, Xiliang, et al.
Pubblicazione: (2024)
di: Zhu, Xiliang, et al.
Pubblicazione: (2024)
Do Efficient Transformers Really Save Computation?
di: Yang, Kai, et al.
Pubblicazione: (2024)
di: Yang, Kai, et al.
Pubblicazione: (2024)
Self-Train Before You Transcribe
di: Flynn, Robert, et al.
Pubblicazione: (2024)
di: Flynn, Robert, et al.
Pubblicazione: (2024)
Do LLMs Really Need 10+ Thoughts for "Find the Time 1000 Days Later"? Towards Structural Understanding of LLM Overthinking
di: Zhang, Xinliang Frederick, et al.
Pubblicazione: (2025)
di: Zhang, Xinliang Frederick, et al.
Pubblicazione: (2025)
Do LLMs Really Know What They Don't Know? Internal States Mainly Reflect Knowledge Recall Rather Than Truthfulness
di: Cheang, Chi Seng, et al.
Pubblicazione: (2025)
di: Cheang, Chi Seng, et al.
Pubblicazione: (2025)
Offloading Score: Measuring AI Reliance Through Counterfactual Workflows
di: Padmakumar, Vishakh, et al.
Pubblicazione: (2026)
di: Padmakumar, Vishakh, et al.
Pubblicazione: (2026)
TOGGL: Transcribing Overlapping Speech with Staggered Labeling
di: Li, Chak-Fai, et al.
Pubblicazione: (2024)
di: Li, Chak-Fai, et al.
Pubblicazione: (2024)
Beyond Classification: Towards Speech Emotion Reasoning with Multitask AudioLLMs
di: Zhang, Wenyu, et al.
Pubblicazione: (2025)
di: Zhang, Wenyu, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Fine-Tuning Text-to-Speech Diffusion Models Using Reinforcement Learning with Human Feedback
di: Chen, Jingyi, et al.
Pubblicazione: (2025) -
Shortcomings of LLMs for Low-Resource Translation: Retrieval and Understanding are Both the Problem
di: Court, Sara, et al.
Pubblicazione: (2024) -
DLPO: Diffusion Model Loss-Guided Reinforcement Learning for Fine-Tuning Text-to-Speech Diffusion Models
di: Chen, Jingyi, et al.
Pubblicazione: (2024) -
Prompt and circumstance: A word-by-word LLM prompting approach to interlinear glossing for low-resource languages
di: Elsner, Micha, et al.
Pubblicazione: (2025) -
Beyond Length: Context-Aware Expansion and Independence as Developmentally Sensitive Evaluation in Child Utterances
di: Chun, Jiyun, et al.
Pubblicazione: (2026)