"Sorry, I Didn't Catch That": How Speech Models Miss What Matters Most
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Kaitlyn, Bartelds, Martijn, Bianchi, Federico, Zou, James |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Model Agreed, But Didn't Learn: Diagnosing Surface Compliance in Large Language Models
by: Gu, Xiaojie, et al.
Published: (2026)
by: Gu, Xiaojie, et al.
Published: (2026)
Too Long, Didn't Model: Decomposing LLM Long-Context Understanding With Novels
by: Hamilton, Sil, et al.
Published: (2025)
by: Hamilton, Sil, et al.
Published: (2025)
Voice "Cloning" is Style Transfer
by: Zhou, Kaitlyn, et al.
Published: (2026)
by: Zhou, Kaitlyn, et al.
Published: (2026)
Belief in the Machine: Investigating Epistemological Blind Spots of Language Models
by: Suzgun, Mirac, et al.
Published: (2024)
by: Suzgun, Mirac, et al.
Published: (2024)
Large Language Models are Vulnerable to Bait-and-Switch Attacks for Generating Harmful Content
by: Bianchi, Federico, et al.
Published: (2024)
by: Bianchi, Federico, et al.
Published: (2024)
What News Recommendation Research Did (But Mostly Didn't) Teach Us About Building A News Recommender
by: Higley, Karl, et al.
Published: (2025)
by: Higley, Karl, et al.
Published: (2025)
Measuring What Matters -- or What's Convenient?: Robustness of LLM-Based Scoring Systems to Construct-Irrelevant Factors
by: Walsh, Cole, et al.
Published: (2026)
by: Walsh, Cole, et al.
Published: (2026)
Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal
by: Prakash, Nirmalendu, et al.
Published: (2025)
by: Prakash, Nirmalendu, et al.
Published: (2025)
Legal Fact Prediction: The Missing Piece in Legal Judgment Prediction
by: Liu, Junkai, et al.
Published: (2024)
by: Liu, Junkai, et al.
Published: (2024)
"What's Up, Doc?": Analyzing How Users Seek Health Information in Large-Scale Conversational AI Datasets
by: Paruchuri, Akshay, et al.
Published: (2025)
by: Paruchuri, Akshay, et al.
Published: (2025)
Green AI: Which Programming Language Consumes the Most?
by: Marini, Niccolò, et al.
Published: (2024)
by: Marini, Niccolò, et al.
Published: (2024)
Confirming Correct, Missing the Rest: LLM Tutoring Agents Struggle Where Feedback Matters Most
by: Yasir, Tahreem, et al.
Published: (2026)
by: Yasir, Tahreem, et al.
Published: (2026)
Beyond Brainstorming: What Drives High-Quality Scientific Ideas? Lessons from Multi-Agent Collaboration
by: Chen, Nuo, et al.
Published: (2025)
by: Chen, Nuo, et al.
Published: (2025)
Why Slop Matters
by: Kommers, Cody, et al.
Published: (2025)
by: Kommers, Cody, et al.
Published: (2025)
How Large Language Models are Designed to Hallucinate
by: Ackermann, Richard, et al.
Published: (2025)
by: Ackermann, Richard, et al.
Published: (2025)
Walking in Others' Shoes: How Perspective-Taking Guides Large Language Models in Reducing Toxicity and Bias
by: Xu, Rongwu, et al.
Published: (2024)
by: Xu, Rongwu, et al.
Published: (2024)
LLM Generated Persona is a Promise with a Catch
by: Li, Ang, et al.
Published: (2025)
by: Li, Ang, et al.
Published: (2025)
Prompt Selection Matters: Enhancing Text Annotations for Social Sciences with Large Language Models
by: Abraham, Louis, et al.
Published: (2024)
by: Abraham, Louis, et al.
Published: (2024)
Stars, Stripes, and Silicon: Unravelling the ChatGPT's All-American, Monochrome, Cis-centric Bias
by: Torrielli, Federico
Published: (2024)
by: Torrielli, Federico
Published: (2024)
How to Connect Speech Foundation Models and Large Language Models? What Matters and What Does Not
by: Verdini, Francesco, et al.
Published: (2024)
by: Verdini, Francesco, et al.
Published: (2024)
Whose Journey Matters? Investigating Identity Biases in Large Language Models (LLMs) for Travel Planning Assistance
by: Ren, Ruiping, et al.
Published: (2024)
by: Ren, Ruiping, et al.
Published: (2024)
Position: The Most Expensive Part of an LLM should be its Training Data
by: Kandpal, Nikhil, et al.
Published: (2025)
by: Kandpal, Nikhil, et al.
Published: (2025)
Towards Fairness Assessment of Dutch Hate Speech Detection
by: Bauer, Julie, et al.
Published: (2025)
by: Bauer, Julie, et al.
Published: (2025)
Ask LLMs Directly, "What shapes your bias?": Measuring Social Bias in Large Language Models
by: Shin, Jisu, et al.
Published: (2024)
by: Shin, Jisu, et al.
Published: (2024)
How Do Language Models Process Ethical Instructions? Deliberation, Consistency, and Other-Recognition Across Four Models
by: Fukui, Hiroki
Published: (2026)
by: Fukui, Hiroki
Published: (2026)
Large language models in medicine: the potentials and pitfalls
by: Omiye, Jesutofunmi A., et al.
Published: (2023)
by: Omiye, Jesutofunmi A., et al.
Published: (2023)
Towards Weakly-Supervised Hate Speech Classification Across Datasets
by: Jin, Yiping, et al.
Published: (2023)
by: Jin, Yiping, et al.
Published: (2023)
Ads in AI Chatbots? An Analysis of How Large Language Models Navigate Conflicts of Interest
by: Wu, Addison J., et al.
Published: (2026)
by: Wu, Addison J., et al.
Published: (2026)
Teaching Language Models How to Code Like Learners: Conversational Serialization for Student Simulation
by: Koutcheme, Charles, et al.
Published: (2026)
by: Koutcheme, Charles, et al.
Published: (2026)
Navigating Dialectal Bias and Ethical Complexities in Levantine Arabic Hate Speech Detection
by: Ahmed, Ahmed Haj, et al.
Published: (2024)
by: Ahmed, Ahmed Haj, et al.
Published: (2024)
Place Matters: Comparing LLM Hallucination Rates for Place-Based Legal Queries
by: Curran, Damian, et al.
Published: (2025)
by: Curran, Damian, et al.
Published: (2025)
What can large language models do for sustainable food?
by: Thomas, Anna T., et al.
Published: (2025)
by: Thomas, Anna T., et al.
Published: (2025)
What Is The Political Content in LLMs' Pre- and Post-Training Data?
by: Ceron, Tanise, et al.
Published: (2025)
by: Ceron, Tanise, et al.
Published: (2025)
Empathy and the Right to Be an Exception: What LLMs Can and Cannot Do
by: Kidder, William, et al.
Published: (2024)
by: Kidder, William, et al.
Published: (2024)
AI Didn't Start the Fire: Examining the Stack Exchange Moderator and Contributor Strike
by: Wu, Yiwei, et al.
Published: (2025)
by: Wu, Yiwei, et al.
Published: (2025)
Where It Really Matters: Few-Shot Environmental Conservation Media Monitoring for Low-Resource Languages
by: Jain, Sameer, et al.
Published: (2024)
by: Jain, Sameer, et al.
Published: (2024)
"I understand why I got this grade": Automatic Short Answer Grading with Feedback
by: Aggarwal, Dishank, et al.
Published: (2024)
by: Aggarwal, Dishank, et al.
Published: (2024)
In Silico Sociology: Forecasting COVID-19 Polarization with Large Language Models
by: Kozlowski, Austin C., et al.
Published: (2024)
by: Kozlowski, Austin C., et al.
Published: (2024)
Classroom AI: Large Language Models as Grade-Specific Teachers
by: Oh, Jio, et al.
Published: (2026)
by: Oh, Jio, et al.
Published: (2026)
Towards medical AI misalignment: a preliminary study
by: Puccio, Barbara, et al.
Published: (2025)
by: Puccio, Barbara, et al.
Published: (2025)
Similar Items
-
The Model Agreed, But Didn't Learn: Diagnosing Surface Compliance in Large Language Models
by: Gu, Xiaojie, et al.
Published: (2026) -
Too Long, Didn't Model: Decomposing LLM Long-Context Understanding With Novels
by: Hamilton, Sil, et al.
Published: (2025) -
Voice "Cloning" is Style Transfer
by: Zhou, Kaitlyn, et al.
Published: (2026) -
Belief in the Machine: Investigating Epistemological Blind Spots of Language Models
by: Suzgun, Mirac, et al.
Published: (2024) -
Large Language Models are Vulnerable to Bait-and-Switch Attacks for Generating Harmful Content
by: Bianchi, Federico, et al.
Published: (2024)