Benchmark Success, Clinical Failure: When Reinforcement Learning Optimizes for Benchmarks, Not Patients
Fuente:
arXiv
Saved in:
| Main Authors: | Berger, Armin, Bergau, Manuela, Schneider, Helen, Ahmad, Saad, Lagones, Tom Anglim, Brugnara, Gianluca, Foltyn-Dumitru, Martha, Schlamp, Kai, Vollmuth, Philipp, Sifa, Rafet |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing
by: Berghaus, David, et al.
Published: (2025)
by: Berghaus, David, et al.
Published: (2025)
Reasoning LLMs in the Medical Domain: A Literature Survey
by: Berger, Armin, et al.
Published: (2025)
by: Berger, Armin, et al.
Published: (2025)
History Rhymes: Macro-Contextual Retrieval for Robust Financial Forecasting
by: Khanna, Sarthak, et al.
Published: (2025)
by: Khanna, Sarthak, et al.
Published: (2025)
Quantum Computing from Hopfield Nets
by: Bauckhage, Christian, et al.
Published: (2025)
by: Bauckhage, Christian, et al.
Published: (2025)
Generalizing Abstention for Noise-Robust Learning in Medical Image Segmentation
by: Moustafa, Wesam, et al.
Published: (2026)
by: Moustafa, Wesam, et al.
Published: (2026)
A Survey on Current Trends and Recent Advances in Text Anonymization
by: Deußer, Tobias, et al.
Published: (2025)
by: Deußer, Tobias, et al.
Published: (2025)
Towards Unified Multimodal Financial Forecasting: Integrating Sentiment Embeddings and Market Indicators via Cross-Modal Attention
by: Khanna, Sarthak, et al.
Published: (2025)
by: Khanna, Sarthak, et al.
Published: (2025)
[Vision Paper] PRObot: Enhancing Patient-Reported Outcome Measures for Diabetic Retinopathy using Chatbots and Generative AI
by: Pielka, Maren, et al.
Published: (2024)
by: Pielka, Maren, et al.
Published: (2024)
From Retinal Pixels to Patients: Evolution of Deep Learning Research in Diabetic Retinopathy Screening
by: Chopra, Muskaan, et al.
Published: (2025)
by: Chopra, Muskaan, et al.
Published: (2025)
Knowing When Not to Predict: Self Supervised Learning and Abstention for Safer DR Screening
by: Chopra, Muskaan, et al.
Published: (2026)
by: Chopra, Muskaan, et al.
Published: (2026)
SynCED-EnDe 2025: A Synthetic and Curated English - German Dataset for Critical Error Detection in Machine Translation
by: Chopra, Muskaan, et al.
Published: (2025)
by: Chopra, Muskaan, et al.
Published: (2025)
Towards Reliable Machine Translation: Scaling LLMs for Critical Error Detection and Safety
by: Chopra, Muskaan, et al.
Published: (2026)
by: Chopra, Muskaan, et al.
Published: (2026)
Model-agnostic Body Part Relevance Assessment for Pedestrian Detection
by: Günder, Maurice, et al.
Published: (2023)
by: Günder, Maurice, et al.
Published: (2023)
Pointer-Guided Pre-Training: Infusing Large Language Models with Paragraph-Level Contextual Awareness
by: Hillebrand, Lars, et al.
Published: (2024)
by: Hillebrand, Lars, et al.
Published: (2024)
Interpretable Topic Extraction and Word Embedding Learning using row-stochastic DEDICOM
by: Hillebrand, Lars, et al.
Published: (2025)
by: Hillebrand, Lars, et al.
Published: (2025)
How Small Can You Go? Compact Language Models for On-Device Critical Error Detection in Machine Translation
by: Chopra, Muskaan, et al.
Published: (2025)
by: Chopra, Muskaan, et al.
Published: (2025)
Advancing Risk and Quality Assurance: A RAG Chatbot for Improved Regulatory Compliance
by: Hillebrand, Lars, et al.
Published: (2025)
by: Hillebrand, Lars, et al.
Published: (2025)
Domain-Adaptation through Synthetic Data: Fine-Tuning Large Language Models for German Law
by: Bashir, Ali Hamza, et al.
Published: (2026)
by: Bashir, Ali Hamza, et al.
Published: (2026)
Die Nachhaltigkeit und der Mittelwald
by: Vollmuth, David Willi
Published: (2021)
by: Vollmuth, David Willi
Published: (2021)
Informed Deep Abstaining Classifier: Investigating noise-robust training for diagnostic decision support systems
by: Schneider, Helen, et al.
Published: (2024)
by: Schneider, Helen, et al.
Published: (2024)
Reformy w polskim szkolnictwie wyższym po 1990 r. w świetle nauki o polityce publicznej
by: Agnieszka Dziedziczak-Foltyn
Published: (2018)
by: Agnieszka Dziedziczak-Foltyn
Published: (2018)
Is continuous CoT better suited for multi-lingual reasoning?
by: Bashir, Ali Hamza, et al.
Published: (2026)
by: Bashir, Ali Hamza, et al.
Published: (2026)
Exact Generalisation Error Exposes Benchmarks Skew Graph Neural Networks Success (or Failure)
by: Ayday, Nil, et al.
Published: (2025)
by: Ayday, Nil, et al.
Published: (2025)
Towards Automated Regulatory Compliance Verification in Financial Auditing with Large Language Models
by: Berger, Armin, et al.
Published: (2025)
by: Berger, Armin, et al.
Published: (2025)
Queue-based Eco-Driving at Roundabouts with Reinforcement Learning
by: Schlamp, Anna-Lena, et al.
Published: (2024)
by: Schlamp, Anna-Lena, et al.
Published: (2024)
Dataset for: Does Organized Misconduct Shape Network Topology? Part III: Topology of Retracted Article Citation Networks
by: IRMAK, Rafet
Published: (2026)
by: IRMAK, Rafet
Published: (2026)
Integración de pruebas y manejo de la restricción del crecimiento fetal
by: Sifa Turan
Published: (2008)
by: Sifa Turan
Published: (2008)
Benchmarking Sensor-Fault Robustness in Forecasting
by: Windmann, Alexander, et al.
Published: (2026)
by: Windmann, Alexander, et al.
Published: (2026)
Which organisational context factors help women to obtain and retain leadership positions in the 21st century? A systematic review and research agenda for human resource management
by: Lioba A. Gierke, et al.
Published: (2024)
by: Lioba A. Gierke, et al.
Published: (2024)
Chapter Deep Learning Training and Benchmarks for Earth Observation Images: Data Sets, Features, and Procedures
by: Schwarz, Gottfried, et al.
Published: (2021)
by: Schwarz, Gottfried, et al.
Published: (2021)
SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios
by: Clark, Jackson, et al.
Published: (2026)
by: Clark, Jackson, et al.
Published: (2026)
When Judgment Becomes Noise: How Design Failures in LLM Judge Benchmarks Silently Undermine Validity
by: Feuer, Benjamin, et al.
Published: (2025)
by: Feuer, Benjamin, et al.
Published: (2025)
LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks
by: Ullah, Saad, et al.
Published: (2023)
by: Ullah, Saad, et al.
Published: (2023)
Forgetful but Faithful: A Cognitive Memory Architecture and Benchmark for Privacy-Aware Generative Agents
by: Alqithami, Saad
Published: (2025)
by: Alqithami, Saad
Published: (2025)
Valoración de la relación entre funciones ejecutivas y conductas agresivas de niños sordos: impacto de la educación especial temprana
by: Rafet Firat Sipal
Published: (2010)
by: Rafet Firat Sipal
Published: (2010)
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
by: Akhtar, Mubashara, et al.
Published: (2026)
by: Akhtar, Mubashara, et al.
Published: (2026)
Benchmarking Failures in Tool-Augmented Language Models
by: Treviño, Eduardo, et al.
Published: (2025)
by: Treviño, Eduardo, et al.
Published: (2025)
SugarViT -- Multi-objective Regression of UAV Images with Vision Transformers and Deep Label Distribution Learning Demonstrated on Disease Severity Prediction in Sugar Beet
by: Günder, Maurice, et al.
Published: (2023)
by: Günder, Maurice, et al.
Published: (2023)
The KMAT: Benchmarking Knowledge Management.
by: de Jager, Martha
Published: (1998)
by: de Jager, Martha
Published: (1998)
IGAff: Benchmarking Adversarial Iterative and Genetic Affine Algorithms on Deep Neural Networks
by: Echim, Sebastian-Vasile, et al.
Published: (2025)
by: Echim, Sebastian-Vasile, et al.
Published: (2025)
Similar Items
-
Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing
by: Berghaus, David, et al.
Published: (2025) -
Reasoning LLMs in the Medical Domain: A Literature Survey
by: Berger, Armin, et al.
Published: (2025) -
History Rhymes: Macro-Contextual Retrieval for Robust Financial Forecasting
by: Khanna, Sarthak, et al.
Published: (2025) -
Quantum Computing from Hopfield Nets
by: Bauckhage, Christian, et al.
Published: (2025) -
Generalizing Abstention for Noise-Robust Learning in Medical Image Segmentation
by: Moustafa, Wesam, et al.
Published: (2026)