Measuring What VLMs Don't Say: Validation Metrics Hide Clinical Terminology Erasure in Radiology Report Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Parikh, Aditya, Feragen, Aasa, Das, Sneha, Frank, Stella |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Who Does Your Algorithm Fail? Investigating Age and Ethnic Bias in the MAMA-MIA Dataset
di: Parikh, Aditya, et al.
Pubblicazione: (2025)
di: Parikh, Aditya, et al.
Pubblicazione: (2025)
Investigating Label Bias and Representational Sources of Age-Related Disparities in Medical Segmentation
di: Parikh, Aditya, et al.
Pubblicazione: (2025)
di: Parikh, Aditya, et al.
Pubblicazione: (2025)
Towards Fairness under Label Bias in Image Segmentation: Impact, Measurement and Mitigation
di: Parikh, Aditya, et al.
Pubblicazione: (2026)
di: Parikh, Aditya, et al.
Pubblicazione: (2026)
Fair Lung Disease Diagnosis from Chest CT via Gender-Adversarial Attention Multiple Instance Learning
di: Parikh, Aditya, et al.
Pubblicazione: (2026)
di: Parikh, Aditya, et al.
Pubblicazione: (2026)
Reasoning Models Don't Always Say What They Think
di: Chen, Yanda, et al.
Pubblicazione: (2025)
di: Chen, Yanda, et al.
Pubblicazione: (2025)
Generalizing Fairness to Generative Language Models via Reformulation of Non-discrimination Criteria
di: Sterlie, Sara, et al.
Pubblicazione: (2024)
di: Sterlie, Sara, et al.
Pubblicazione: (2024)
What Models Know, How Well They Know It: Knowledge-Weighted Fine-Tuning for Learning When to Say "I Don't Know"
di: Lee, Joosung, et al.
Pubblicazione: (2026)
di: Lee, Joosung, et al.
Pubblicazione: (2026)
Clinically Grounded Agent-based Report Evaluation: An Interpretable Metric for Radiology Report Generation
di: Dua, Radhika, et al.
Pubblicazione: (2025)
di: Dua, Radhika, et al.
Pubblicazione: (2025)
Can AI Assistants Know What They Don't Know?
di: Cheng, Qinyuan, et al.
Pubblicazione: (2024)
di: Cheng, Qinyuan, et al.
Pubblicazione: (2024)
What Language Models Know But Don't Say: Non-Generative Prior Extraction for Generalization
di: Rezaeimanesh, Sara, et al.
Pubblicazione: (2026)
di: Rezaeimanesh, Sara, et al.
Pubblicazione: (2026)
What Prompts Don't Say: Understanding and Managing Underspecification in LLM Prompts
di: Yang, Chenyang, et al.
Pubblicazione: (2025)
di: Yang, Chenyang, et al.
Pubblicazione: (2025)
Implicit Intelligence -- Evaluating Agents on What Users Don't Say
di: Sirdeshmukh, Ved, et al.
Pubblicazione: (2026)
di: Sirdeshmukh, Ved, et al.
Pubblicazione: (2026)
When Words Don't Mean What They Say: Figurative Understanding in Bengali Idioms
di: Sakhawat, Adib, et al.
Pubblicazione: (2026)
di: Sakhawat, Adib, et al.
Pubblicazione: (2026)
Honest AI: Fine-Tuning "Small" Language Models to Say "I Don't Know", and Reducing Hallucination in RAG
di: Chen, Xinxi, et al.
Pubblicazione: (2024)
di: Chen, Xinxi, et al.
Pubblicazione: (2024)
Don't Say No: Jailbreaking LLM by Suppressing Refusal
di: Zhou, Yukai, et al.
Pubblicazione: (2024)
di: Zhou, Yukai, et al.
Pubblicazione: (2024)
Why Don't Prompt-Based Fairness Metrics Correlate?
di: Zayed, Abdelrahman, et al.
Pubblicazione: (2024)
di: Zayed, Abdelrahman, et al.
Pubblicazione: (2024)
Don't Pay Attention
di: Hammoud, Mohammad, et al.
Pubblicazione: (2025)
di: Hammoud, Mohammad, et al.
Pubblicazione: (2025)
CRIMSON: A Clinically-Grounded LLM-Based Metric for Generative Radiology Report Evaluation
di: Baharoon, Mohammed, et al.
Pubblicazione: (2026)
di: Baharoon, Mohammed, et al.
Pubblicazione: (2026)
What Generative Artificial Intelligence Means for Terminological Definitions
di: Martín, Antonio San
Pubblicazione: (2024)
di: Martín, Antonio San
Pubblicazione: (2024)
Large Language Models Must Be Taught to Know What They Don't Know
di: Kapoor, Sanyam, et al.
Pubblicazione: (2024)
di: Kapoor, Sanyam, et al.
Pubblicazione: (2024)
Development and Validation of a Large Language Model for Generating Fully-Structured Radiology Reports
di: Niu, Chuang, et al.
Pubblicazione: (2024)
di: Niu, Chuang, et al.
Pubblicazione: (2024)
Patronus: Interpretable Diffusion Models with Prototypes
di: Weng, Nina, et al.
Pubblicazione: (2025)
di: Weng, Nina, et al.
Pubblicazione: (2025)
RadReason: Radiology Report Evaluation Metric with Reasons and Sub-Scores
di: Li, Yingshu, et al.
Pubblicazione: (2025)
di: Li, Yingshu, et al.
Pubblicazione: (2025)
Reasoning Models Reason Well, Until They Don't
di: Rameshkumar, Revanth, et al.
Pubblicazione: (2025)
di: Rameshkumar, Revanth, et al.
Pubblicazione: (2025)
Show, Don't Tell: Uncovering Implicit Character Portrayal using LLMs
di: Jaipersaud, Brandon, et al.
Pubblicazione: (2024)
di: Jaipersaud, Brandon, et al.
Pubblicazione: (2024)
MedCT: A Clinical Terminology Graph for Generative AI Applications in Healthcare
di: Chen, Ye, et al.
Pubblicazione: (2025)
di: Chen, Ye, et al.
Pubblicazione: (2025)
When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't
di: Nemitz, Jonathan, et al.
Pubblicazione: (2026)
di: Nemitz, Jonathan, et al.
Pubblicazione: (2026)
Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth
di: Gur-Arieh, Yoav, et al.
Pubblicazione: (2026)
di: Gur-Arieh, Yoav, et al.
Pubblicazione: (2026)
Language Models Don't Learn the Physical Manifestation of Language
di: Lee, Bruce W., et al.
Pubblicazione: (2024)
di: Lee, Bruce W., et al.
Pubblicazione: (2024)
R-Tuning: Instructing Large Language Models to Say `I Don't Know'
di: Zhang, Hanning, et al.
Pubblicazione: (2023)
di: Zhang, Hanning, et al.
Pubblicazione: (2023)
Don't Think Twice! Over-Reasoning Impairs Confidence Calibration
di: Lacombe, Romain, et al.
Pubblicazione: (2025)
di: Lacombe, Romain, et al.
Pubblicazione: (2025)
sDPO: Don't Use Your Data All at Once
di: Kim, Dahyun, et al.
Pubblicazione: (2024)
di: Kim, Dahyun, et al.
Pubblicazione: (2024)
GREEN: Generative Radiology Report Evaluation and Error Notation
di: Ostmeier, Sophie, et al.
Pubblicazione: (2024)
di: Ostmeier, Sophie, et al.
Pubblicazione: (2024)
Generating High Quality Synthetic Data for Dutch Medical Conversations
di: Kuan, Cecilia, et al.
Pubblicazione: (2026)
di: Kuan, Cecilia, et al.
Pubblicazione: (2026)
Why Models Know But Don't Say: Chain-of-Thought Faithfulness Divergence Between Thinking Tokens and Answers in Open-Weight Reasoning Models
di: Young, Richard J.
Pubblicazione: (2026)
di: Young, Richard J.
Pubblicazione: (2026)
Do Retrieval Augmented Language Models Know When They Don't Know?
di: Zhou, Youchao, et al.
Pubblicazione: (2025)
di: Zhou, Youchao, et al.
Pubblicazione: (2025)
Don't Overthink it. Preferring Shorter Thinking Chains for Improved LLM Reasoning
di: Hassid, Michael, et al.
Pubblicazione: (2025)
di: Hassid, Michael, et al.
Pubblicazione: (2025)
Don't Buy it! Reassessing the Ad Understanding Abilities of Contrastive Multimodal Models
di: Bavaresco, A., et al.
Pubblicazione: (2024)
di: Bavaresco, A., et al.
Pubblicazione: (2024)
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
di: Mayne, Harry, et al.
Pubblicazione: (2025)
di: Mayne, Harry, et al.
Pubblicazione: (2025)
CLEAR: A Clinically-Grounded Tabular Framework for Radiology Report Evaluation
di: Jiang, Yuyang, et al.
Pubblicazione: (2025)
di: Jiang, Yuyang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Who Does Your Algorithm Fail? Investigating Age and Ethnic Bias in the MAMA-MIA Dataset
di: Parikh, Aditya, et al.
Pubblicazione: (2025) -
Investigating Label Bias and Representational Sources of Age-Related Disparities in Medical Segmentation
di: Parikh, Aditya, et al.
Pubblicazione: (2025) -
Towards Fairness under Label Bias in Image Segmentation: Impact, Measurement and Mitigation
di: Parikh, Aditya, et al.
Pubblicazione: (2026) -
Fair Lung Disease Diagnosis from Chest CT via Gender-Adversarial Attention Multiple Instance Learning
di: Parikh, Aditya, et al.
Pubblicazione: (2026) -
Reasoning Models Don't Always Say What They Think
di: Chen, Yanda, et al.
Pubblicazione: (2025)