An autonomous agent for auditing and improving the reliability of clinical AI models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Kuhn, Lukas, Buettner, Florian |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Non-Contrastive Vision-Language Learning with Predictive Embedding Alignment
par: Kuhn, Lukas, et autres
Publié: (2026)
par: Kuhn, Lukas, et autres
Publié: (2026)
From Entropy to Calibrated Uncertainty: Training Language Models to Reason About Uncertainty
par: Jenane, Azza, et autres
Publié: (2026)
par: Jenane, Azza, et autres
Publié: (2026)
LVLM-Aided Alignment of Task-Specific Vision Models
par: Koebler, Alexander, et autres
Publié: (2025)
par: Koebler, Alexander, et autres
Publié: (2025)
Context is all you need: Towards autonomous model-based process design using agentic AI in flowsheet simulations
par: Schäfer, Pascal, et autres
Publié: (2026)
par: Schäfer, Pascal, et autres
Publié: (2026)
Improving Perturbation-based Explanations by Understanding the Role of Uncertainty Calibration
par: Decker, Thomas, et autres
Publié: (2025)
par: Decker, Thomas, et autres
Publié: (2025)
Why Uncertainty Calibration Matters for Reliable Perturbation-based Explanations
par: Decker, Thomas, et autres
Publié: (2025)
par: Decker, Thomas, et autres
Publié: (2025)
How to Leverage Predictive Uncertainty Estimates for Reducing Catastrophic Forgetting in Online Continual Learning
par: Serra, Giuseppe, et autres
Publié: (2024)
par: Serra, Giuseppe, et autres
Publié: (2024)
Mind the Gap: A Framework for Assessing Pitfalls in Multimodal Active Learning
par: Eisenhardt, Dustin, et autres
Publié: (2026)
par: Eisenhardt, Dustin, et autres
Publié: (2026)
A collaborative agent with two lightweight synergistic models for autonomous crystal materials research
par: Shi, Tongyu, et autres
Publié: (2026)
par: Shi, Tongyu, et autres
Publié: (2026)
Building reliable sim driving agents by scaling self-play
par: Cornelisse, Daphne, et autres
Publié: (2025)
par: Cornelisse, Daphne, et autres
Publié: (2025)
Incremental Uncertainty-aware Performance Monitoring with Active Labeling Intervention
par: Koebler, Alexander, et autres
Publié: (2025)
par: Koebler, Alexander, et autres
Publié: (2025)
Grasping Partially Occluded Objects Using Autoencoder-Based Point Cloud Inpainting
par: Koebler, Alexander, et autres
Publié: (2025)
par: Koebler, Alexander, et autres
Publié: (2025)
How frontier AI companies could implement an internal audit function
par: Gomez, Francesca, et autres
Publié: (2025)
par: Gomez, Francesca, et autres
Publié: (2025)
Quantifying the Gaps Between Translation and Native Perception in Training for Multimodal, Multilingual Retrieval
par: Buettner, Kyle, et autres
Publié: (2024)
par: Buettner, Kyle, et autres
Publié: (2024)
Supervising the search process produces reliable and generalizable information-seeking agents
par: Xiong, Guangzhi, et autres
Publié: (2025)
par: Xiong, Guangzhi, et autres
Publié: (2025)
Adaptive auditing of AI systems with anytime-valid guarantees
par: Zhou, Siyu, et autres
Publié: (2026)
par: Zhou, Siyu, et autres
Publié: (2026)
Participatory provenance as representational auditing for AI-mediated public consultation
par: Mahajan, Sachit
Publié: (2026)
par: Mahajan, Sachit
Publié: (2026)
Zero-shot adaptable task planning for autonomous construction robots: a comparative study of lightweight single and multi-AI agent systems
par: Naderi, Hossein, et autres
Publié: (2026)
par: Naderi, Hossein, et autres
Publié: (2026)
Inadequacy of common stochastic neural networks for reliable clinical decision support
par: Lindenmeyer, Adrian, et autres
Publié: (2024)
par: Lindenmeyer, Adrian, et autres
Publié: (2024)
Hallucination, reliability, and the role of generative AI in science
par: Rathkopf, Charles
Publié: (2025)
par: Rathkopf, Charles
Publié: (2025)
Blueprint-Bench: Comparing spatial intelligence of LLMs, agents and image models
par: Petersson, Lukas, et autres
Publié: (2025)
par: Petersson, Lukas, et autres
Publié: (2025)
An AI-native experimental laboratory for autonomous biomolecular engineering
par: Wu, Mingyu, et autres
Publié: (2025)
par: Wu, Mingyu, et autres
Publié: (2025)
A Behavior Tree-inspired programming language for autonomous agents
par: Biggar, Oliver, et autres
Publié: (2024)
par: Biggar, Oliver, et autres
Publié: (2024)
Agentic AI for autonomous anomaly management in complex systems
par: Barenji, Reza Vatankhah, et autres
Publié: (2025)
par: Barenji, Reza Vatankhah, et autres
Publié: (2025)
AI co-mathematician: Accelerating mathematicians with agentic AI
par: Zheng, Daniel, et autres
Publié: (2026)
par: Zheng, Daniel, et autres
Publié: (2026)
A Technological Perspective on Misuse of Available AI
par: Pöhler, Lukas, et autres
Publié: (2024)
par: Pöhler, Lukas, et autres
Publié: (2024)
SiriuS: Self-improving Multi-agent Systems via Bootstrapped Reasoning
par: Zhao, Wanjia, et autres
Publié: (2025)
par: Zhao, Wanjia, et autres
Publié: (2025)
Provably Better Explanations with Optimized Aggregation of Feature Attributions
par: Decker, Thomas, et autres
Publié: (2024)
par: Decker, Thomas, et autres
Publié: (2024)
Continually self-improving AI
par: Yang, Zitong
Publié: (2026)
par: Yang, Zitong
Publié: (2026)
Can LLM-Augmented autonomous agents cooperate?, An evaluation of their cooperative capabilities through Melting Pot
par: Mosquera, Manuel, et autres
Publié: (2024)
par: Mosquera, Manuel, et autres
Publié: (2024)
Beyond Overconfidence: Foundation Models Redefine Calibration in Deep Neural Networks
par: Hekler, Achim, et autres
Publié: (2025)
par: Hekler, Achim, et autres
Publié: (2025)
Adaptive AI decision interface for autonomous electronic material discovery
par: Dai, Yahao, et autres
Publié: (2025)
par: Dai, Yahao, et autres
Publié: (2025)
From large language models to multimodal AI: A scoping review on the potential of generative AI in medicine
par: Buess, Lukas, et autres
Publié: (2025)
par: Buess, Lukas, et autres
Publié: (2025)
A Multimodal Recaptioning Framework to Account for Perceptual Diversity Across Languages in Vision-Language Modeling
par: Buettner, Kyle, et autres
Publié: (2025)
par: Buettner, Kyle, et autres
Publié: (2025)
An empirical evaluation of the risks of AI model updates using clinical data: stability, arbitrariness, and fairness
par: Bilionis, Ioannis, et autres
Publié: (2026)
par: Bilionis, Ioannis, et autres
Publié: (2026)
Towards autonomous quantum physics research using LLM agents with access to intelligent tools
par: Arlt, Sören, et autres
Publié: (2025)
par: Arlt, Sören, et autres
Publié: (2025)
A3D: Agentic AI flow for autonomous Accelerator Design
par: Nallathambi, Abinand, et autres
Publié: (2026)
par: Nallathambi, Abinand, et autres
Publié: (2026)
Design Conductor: An agent autonomously builds a 1.5 GHz Linux-capable RISC-V CPU
par: The Verkor Team, et autres
Publié: (2026)
par: The Verkor Team, et autres
Publié: (2026)
Automated QoR improvement in OpenROAD with coding agents
par: Ghose, Amur, et autres
Publié: (2026)
par: Ghose, Amur, et autres
Publié: (2026)
Addressing the regulatory gap: moving towards an EU AI audit ecosystem beyond the AI Act by including civil society
par: Hartmann, David, et autres
Publié: (2024)
par: Hartmann, David, et autres
Publié: (2024)
Documents similaires
-
Non-Contrastive Vision-Language Learning with Predictive Embedding Alignment
par: Kuhn, Lukas, et autres
Publié: (2026) -
From Entropy to Calibrated Uncertainty: Training Language Models to Reason About Uncertainty
par: Jenane, Azza, et autres
Publié: (2026) -
LVLM-Aided Alignment of Task-Specific Vision Models
par: Koebler, Alexander, et autres
Publié: (2025) -
Context is all you need: Towards autonomous model-based process design using agentic AI in flowsheet simulations
par: Schäfer, Pascal, et autres
Publié: (2026) -
Improving Perturbation-based Explanations by Understanding the Role of Uncertainty Calibration
par: Decker, Thomas, et autres
Publié: (2025)