Measuring What VLMs Don't Say: Validation Metrics Hide Clinical Terminology Erasure in Radiology Report Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Parikh, Aditya, Feragen, Aasa, Das, Sneha, Frank, Stella |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Who Does Your Algorithm Fail? Investigating Age and Ethnic Bias in the MAMA-MIA Dataset
by: Parikh, Aditya, et al.
Published: (2025)
by: Parikh, Aditya, et al.
Published: (2025)
Investigating Label Bias and Representational Sources of Age-Related Disparities in Medical Segmentation
by: Parikh, Aditya, et al.
Published: (2025)
by: Parikh, Aditya, et al.
Published: (2025)
Towards Fairness under Label Bias in Image Segmentation: Impact, Measurement and Mitigation
by: Parikh, Aditya, et al.
Published: (2026)
by: Parikh, Aditya, et al.
Published: (2026)
Fair Lung Disease Diagnosis from Chest CT via Gender-Adversarial Attention Multiple Instance Learning
by: Parikh, Aditya, et al.
Published: (2026)
by: Parikh, Aditya, et al.
Published: (2026)
Reasoning Models Don't Always Say What They Think
by: Chen, Yanda, et al.
Published: (2025)
by: Chen, Yanda, et al.
Published: (2025)
Generalizing Fairness to Generative Language Models via Reformulation of Non-discrimination Criteria
by: Sterlie, Sara, et al.
Published: (2024)
by: Sterlie, Sara, et al.
Published: (2024)
What Models Know, How Well They Know It: Knowledge-Weighted Fine-Tuning for Learning When to Say "I Don't Know"
by: Lee, Joosung, et al.
Published: (2026)
by: Lee, Joosung, et al.
Published: (2026)
Clinically Grounded Agent-based Report Evaluation: An Interpretable Metric for Radiology Report Generation
by: Dua, Radhika, et al.
Published: (2025)
by: Dua, Radhika, et al.
Published: (2025)
Can AI Assistants Know What They Don't Know?
by: Cheng, Qinyuan, et al.
Published: (2024)
by: Cheng, Qinyuan, et al.
Published: (2024)
What Language Models Know But Don't Say: Non-Generative Prior Extraction for Generalization
by: Rezaeimanesh, Sara, et al.
Published: (2026)
by: Rezaeimanesh, Sara, et al.
Published: (2026)
What Prompts Don't Say: Understanding and Managing Underspecification in LLM Prompts
by: Yang, Chenyang, et al.
Published: (2025)
by: Yang, Chenyang, et al.
Published: (2025)
Implicit Intelligence -- Evaluating Agents on What Users Don't Say
by: Sirdeshmukh, Ved, et al.
Published: (2026)
by: Sirdeshmukh, Ved, et al.
Published: (2026)
When Words Don't Mean What They Say: Figurative Understanding in Bengali Idioms
by: Sakhawat, Adib, et al.
Published: (2026)
by: Sakhawat, Adib, et al.
Published: (2026)
Honest AI: Fine-Tuning "Small" Language Models to Say "I Don't Know", and Reducing Hallucination in RAG
by: Chen, Xinxi, et al.
Published: (2024)
by: Chen, Xinxi, et al.
Published: (2024)
Don't Say No: Jailbreaking LLM by Suppressing Refusal
by: Zhou, Yukai, et al.
Published: (2024)
by: Zhou, Yukai, et al.
Published: (2024)
Why Don't Prompt-Based Fairness Metrics Correlate?
by: Zayed, Abdelrahman, et al.
Published: (2024)
by: Zayed, Abdelrahman, et al.
Published: (2024)
Don't Pay Attention
by: Hammoud, Mohammad, et al.
Published: (2025)
by: Hammoud, Mohammad, et al.
Published: (2025)
CRIMSON: A Clinically-Grounded LLM-Based Metric for Generative Radiology Report Evaluation
by: Baharoon, Mohammed, et al.
Published: (2026)
by: Baharoon, Mohammed, et al.
Published: (2026)
What Generative Artificial Intelligence Means for Terminological Definitions
by: Martín, Antonio San
Published: (2024)
by: Martín, Antonio San
Published: (2024)
Large Language Models Must Be Taught to Know What They Don't Know
by: Kapoor, Sanyam, et al.
Published: (2024)
by: Kapoor, Sanyam, et al.
Published: (2024)
Development and Validation of a Large Language Model for Generating Fully-Structured Radiology Reports
by: Niu, Chuang, et al.
Published: (2024)
by: Niu, Chuang, et al.
Published: (2024)
Patronus: Interpretable Diffusion Models with Prototypes
by: Weng, Nina, et al.
Published: (2025)
by: Weng, Nina, et al.
Published: (2025)
RadReason: Radiology Report Evaluation Metric with Reasons and Sub-Scores
by: Li, Yingshu, et al.
Published: (2025)
by: Li, Yingshu, et al.
Published: (2025)
Reasoning Models Reason Well, Until They Don't
by: Rameshkumar, Revanth, et al.
Published: (2025)
by: Rameshkumar, Revanth, et al.
Published: (2025)
Show, Don't Tell: Uncovering Implicit Character Portrayal using LLMs
by: Jaipersaud, Brandon, et al.
Published: (2024)
by: Jaipersaud, Brandon, et al.
Published: (2024)
MedCT: A Clinical Terminology Graph for Generative AI Applications in Healthcare
by: Chen, Ye, et al.
Published: (2025)
by: Chen, Ye, et al.
Published: (2025)
When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't
by: Nemitz, Jonathan, et al.
Published: (2026)
by: Nemitz, Jonathan, et al.
Published: (2026)
Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth
by: Gur-Arieh, Yoav, et al.
Published: (2026)
by: Gur-Arieh, Yoav, et al.
Published: (2026)
Language Models Don't Learn the Physical Manifestation of Language
by: Lee, Bruce W., et al.
Published: (2024)
by: Lee, Bruce W., et al.
Published: (2024)
R-Tuning: Instructing Large Language Models to Say `I Don't Know'
by: Zhang, Hanning, et al.
Published: (2023)
by: Zhang, Hanning, et al.
Published: (2023)
Don't Think Twice! Over-Reasoning Impairs Confidence Calibration
by: Lacombe, Romain, et al.
Published: (2025)
by: Lacombe, Romain, et al.
Published: (2025)
sDPO: Don't Use Your Data All at Once
by: Kim, Dahyun, et al.
Published: (2024)
by: Kim, Dahyun, et al.
Published: (2024)
GREEN: Generative Radiology Report Evaluation and Error Notation
by: Ostmeier, Sophie, et al.
Published: (2024)
by: Ostmeier, Sophie, et al.
Published: (2024)
Generating High Quality Synthetic Data for Dutch Medical Conversations
by: Kuan, Cecilia, et al.
Published: (2026)
by: Kuan, Cecilia, et al.
Published: (2026)
Why Models Know But Don't Say: Chain-of-Thought Faithfulness Divergence Between Thinking Tokens and Answers in Open-Weight Reasoning Models
by: Young, Richard J.
Published: (2026)
by: Young, Richard J.
Published: (2026)
Do Retrieval Augmented Language Models Know When They Don't Know?
by: Zhou, Youchao, et al.
Published: (2025)
by: Zhou, Youchao, et al.
Published: (2025)
Don't Overthink it. Preferring Shorter Thinking Chains for Improved LLM Reasoning
by: Hassid, Michael, et al.
Published: (2025)
by: Hassid, Michael, et al.
Published: (2025)
Don't Buy it! Reassessing the Ad Understanding Abilities of Contrastive Multimodal Models
by: Bavaresco, A., et al.
Published: (2024)
by: Bavaresco, A., et al.
Published: (2024)
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
by: Mayne, Harry, et al.
Published: (2025)
by: Mayne, Harry, et al.
Published: (2025)
CLEAR: A Clinically-Grounded Tabular Framework for Radiology Report Evaluation
by: Jiang, Yuyang, et al.
Published: (2025)
by: Jiang, Yuyang, et al.
Published: (2025)
Similar Items
-
Who Does Your Algorithm Fail? Investigating Age and Ethnic Bias in the MAMA-MIA Dataset
by: Parikh, Aditya, et al.
Published: (2025) -
Investigating Label Bias and Representational Sources of Age-Related Disparities in Medical Segmentation
by: Parikh, Aditya, et al.
Published: (2025) -
Towards Fairness under Label Bias in Image Segmentation: Impact, Measurement and Mitigation
by: Parikh, Aditya, et al.
Published: (2026) -
Fair Lung Disease Diagnosis from Chest CT via Gender-Adversarial Attention Multiple Instance Learning
by: Parikh, Aditya, et al.
Published: (2026) -
Reasoning Models Don't Always Say What They Think
by: Chen, Yanda, et al.
Published: (2025)