Beyond Benchmarks: Dynamic, Automatic And Systematic Red-Teaming Agents For Trustworthy Medical Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pan, Jiazhen, Jian, Bailiang, Hager, Paul, Zhang, Yundi, Liu, Che, Jungmann, Friedrike, Li, Hongwei Bran, You, Chenyu, Wu, Junde, Zhu, Jiayuan, Liu, Fenglin, Liu, Yuyuan, Bubeck, Niklas, Wachinger, Christian, Chen, Gong, Zhenyu, Ouyang, Cheng, Kaissis, Georgios, Wiestler, Benedikt, Rueckert, Daniel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Disentangling Progress in Medical Image Registration: Beyond Trend-Driven Architectures towards Domain-Specific Strategies
von: Jian, Bailiang, et al.
Veröffentlicht: (2025)
von: Jian, Bailiang, et al.
Veröffentlicht: (2025)
Mamba? Catch The Hype Or Rethink What Really Helps for Image Registration
von: Jian, Bailiang, et al.
Veröffentlicht: (2024)
von: Jian, Bailiang, et al.
Veröffentlicht: (2024)
Ask Patients with Patience: Enabling LLMs for Human-Centric Medical Dialogue with Grounded Reasoning
von: Zhu, Jiayuan, et al.
Veröffentlicht: (2025)
von: Zhu, Jiayuan, et al.
Veröffentlicht: (2025)
MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning
von: Pan, Jiazhen, et al.
Veröffentlicht: (2025)
von: Pan, Jiazhen, et al.
Veröffentlicht: (2025)
Reconstruct or Generate: Exploring the Spectrum of Generative Modeling for Cardiac MRI
von: Bubeck, Niklas, et al.
Veröffentlicht: (2025)
von: Bubeck, Niklas, et al.
Veröffentlicht: (2025)
TimeFlow: Temporal Conditioning for Longitudinal Brain MRI Registration and Aging Analysis
von: Jian, Bailiang, et al.
Veröffentlicht: (2025)
von: Jian, Bailiang, et al.
Veröffentlicht: (2025)
No Image, No Problem: End-to-End Multi-Task Cardiac Analysis from Undersampled k-Space
von: Zhang, Yundi, et al.
Veröffentlicht: (2026)
von: Zhang, Yundi, et al.
Veröffentlicht: (2026)
Towards Cardiac MRI Foundation Models: Comprehensive Visual-Tabular Representations for Whole-Heart Assessment and Beyond
von: Zhang, Yundi, et al.
Veröffentlicht: (2025)
von: Zhang, Yundi, et al.
Veröffentlicht: (2025)
Unsupervised whole-heart function assessment
von: Zhang, Yundi, et al.
Veröffentlicht: (2025)
von: Zhang, Yundi, et al.
Veröffentlicht: (2025)
Generating Reports or Repeating Templates? Measuring and Mitigating Template Collapse in 3D CT Report Generation
von: Maye-Lasserre, Tom, et al.
Veröffentlicht: (2026)
von: Maye-Lasserre, Tom, et al.
Veröffentlicht: (2026)
Evo-MedAgent: Beyond One-Shot Diagnosis with Agents That Remember, Reflect, and Improve
von: Shen, Weixiang, et al.
Veröffentlicht: (2026)
von: Shen, Weixiang, et al.
Veröffentlicht: (2026)
Git Context Controller: Manage the Context of LLM-based Agents like Git
von: Wu, Junde, et al.
Veröffentlicht: (2025)
von: Wu, Junde, et al.
Veröffentlicht: (2025)
Incentivising the federation: gradient-based metrics for data selection and valuation in private decentralised training
von: Usynin, Dmitrii, et al.
Veröffentlicht: (2023)
von: Usynin, Dmitrii, et al.
Veröffentlicht: (2023)
How Much Is One Recurrence Worth? Iso-Depth Scaling Laws for Looped Language Models
von: Schwethelm, Kristian, et al.
Veröffentlicht: (2026)
von: Schwethelm, Kristian, et al.
Veröffentlicht: (2026)
ChEX: Interactive Localization and Region Description in Chest X-rays
von: Müller, Philip, et al.
Veröffentlicht: (2024)
von: Müller, Philip, et al.
Veröffentlicht: (2024)
Survival In-Context: Amortized Bayesian Survival Analysis via Prior-Fitted Networks
von: Seletkov, Dmitrii, et al.
Veröffentlicht: (2026)
von: Seletkov, Dmitrii, et al.
Veröffentlicht: (2026)
Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools
von: Wu, Junde, et al.
Veröffentlicht: (2025)
von: Wu, Junde, et al.
Veröffentlicht: (2025)
VariViT: A Vision Transformer for Variable Image Sizes
von: Varma, Aswathi, et al.
Veröffentlicht: (2026)
von: Varma, Aswathi, et al.
Veröffentlicht: (2026)
Whole Heart 3D+T Representation Learning Through Sparse 2D Cardiac MR Images
von: Zhang, Yundi, et al.
Veröffentlicht: (2024)
von: Zhang, Yundi, et al.
Veröffentlicht: (2024)
Gradient-Weight Alignment as a Train-Time Proxy for Generalization in Classification Tasks
von: Hölzl, Florian A., et al.
Veröffentlicht: (2025)
von: Hölzl, Florian A., et al.
Veröffentlicht: (2025)
Efficient numeracy in language models through single-token number embeddings
von: Kreitner, Linus, et al.
Veröffentlicht: (2025)
von: Kreitner, Linus, et al.
Veröffentlicht: (2025)
From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding
von: Liu, Yuyuan, et al.
Veröffentlicht: (2026)
von: Liu, Yuyuan, et al.
Veröffentlicht: (2026)
Evaluating the Impact of Medical Image Reconstruction on Downstream AI Fairness and Performance
von: Wohlrapp, Matteo, et al.
Veröffentlicht: (2026)
von: Wohlrapp, Matteo, et al.
Veröffentlicht: (2026)
Weakly Supervised Object Detection in Chest X-Rays with Differentiable ROI Proposal Networks and Soft ROI Pooling
von: Müller, Philip, et al.
Veröffentlicht: (2024)
von: Müller, Philip, et al.
Veröffentlicht: (2024)
Kernel Normalized Convolutional Networks
von: Nasirigerdeh, Reza, et al.
Veröffentlicht: (2022)
von: Nasirigerdeh, Reza, et al.
Veröffentlicht: (2022)
Laplace Sample Information: Data Informativeness Through a Bayesian Lens
von: Kaiser, Johannes, et al.
Veröffentlicht: (2025)
von: Kaiser, Johannes, et al.
Veröffentlicht: (2025)
Does DINOv3 Set a New Medical Vision Standard? Benchmarking 2D and 3D Classification, Segmentation, and Registration
von: Liu, Che, et al.
Veröffentlicht: (2025)
von: Liu, Che, et al.
Veröffentlicht: (2025)
Agentic Large Language Models for Training-Free Neuro-Radiological Image Analysis
von: Erdur, Ayhan Can, et al.
Veröffentlicht: (2026)
von: Erdur, Ayhan Can, et al.
Veröffentlicht: (2026)
Evo: Autoregressive-Diffusion Large Language Models with Evolving Balance
von: Wu, Junde, et al.
Veröffentlicht: (2026)
von: Wu, Junde, et al.
Veröffentlicht: (2026)
MedOpenClaw and MedFlowBench: Auditing Medical Agents in Full-Study Workflows
von: Shen, Weixiang, et al.
Veröffentlicht: (2026)
von: Shen, Weixiang, et al.
Veröffentlicht: (2026)
Single-subject Multi-contrast MRI Super-resolution via Implicit Neural Representations
von: McGinnis, Julian, et al.
Veröffentlicht: (2023)
von: McGinnis, Julian, et al.
Veröffentlicht: (2023)
MultiMAE for Brain MRIs: Robustness to Missing Inputs Using Multi-Modal Masked Autoencoder
von: Erdur, Ayhan Can, et al.
Veröffentlicht: (2025)
von: Erdur, Ayhan Can, et al.
Veröffentlicht: (2025)
Learning Brain Tumor Representation in 3D High-Resolution MR Images via Interpretable State Space Models
von: Hu, Qingqiao, et al.
Veröffentlicht: (2024)
von: Hu, Qingqiao, et al.
Veröffentlicht: (2024)
Beyond the Calibration Point: Mechanism Comparison in Differential Privacy
von: Kaissis, Georgios, et al.
Veröffentlicht: (2024)
von: Kaissis, Georgios, et al.
Veröffentlicht: (2024)
BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents
von: Wu, Jinge, et al.
Veröffentlicht: (2026)
von: Wu, Jinge, et al.
Veröffentlicht: (2026)
Direct Cardiac Segmentation from Undersampled K-space Using Transformers
von: Zhang, Yundi, et al.
Veröffentlicht: (2024)
von: Zhang, Yundi, et al.
Veröffentlicht: (2024)
Reconstruction-free segmentation from undersampled k-space using transformers
von: Zhang, Yundi, et al.
Veröffentlicht: (2025)
von: Zhang, Yundi, et al.
Veröffentlicht: (2025)
Towards Universal Unsupervised Anomaly Detection in Medical Imaging
von: Bercea, Cosmin I., et al.
Veröffentlicht: (2024)
von: Bercea, Cosmin I., et al.
Veröffentlicht: (2024)
Diffusion Models with Implicit Guidance for Medical Anomaly Detection
von: Bercea, Cosmin I., et al.
Veröffentlicht: (2024)
von: Bercea, Cosmin I., et al.
Veröffentlicht: (2024)
Latent Interpolation Learning Using Diffusion Models for Cardiac Volume Reconstruction
von: Bubeck, Niklas, et al.
Veröffentlicht: (2025)
von: Bubeck, Niklas, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Disentangling Progress in Medical Image Registration: Beyond Trend-Driven Architectures towards Domain-Specific Strategies
von: Jian, Bailiang, et al.
Veröffentlicht: (2025) -
Mamba? Catch The Hype Or Rethink What Really Helps for Image Registration
von: Jian, Bailiang, et al.
Veröffentlicht: (2024) -
Ask Patients with Patience: Enabling LLMs for Human-Centric Medical Dialogue with Grounded Reasoning
von: Zhu, Jiayuan, et al.
Veröffentlicht: (2025) -
MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning
von: Pan, Jiazhen, et al.
Veröffentlicht: (2025) -
Reconstruct or Generate: Exploring the Spectrum of Generative Modeling for Cardiac MRI
von: Bubeck, Niklas, et al.
Veröffentlicht: (2025)