Robustness tests for biomedical foundation models should tailor to specifications
Fuente:
arXiv
Salvato in:
| Autori principali: | Xian, R. Patrick, Baker, Noah R., David, Tom, Cui, Qiming, Holmgren, A. Jay, Bauer, Stefan, Sushil, Madhumita, Abbasi-Asl, Reza |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Measuring temporal effects of agent knowledge by date-controlled tool use
di: Xian, R. Patrick, et al.
Pubblicazione: (2025)
di: Xian, R. Patrick, et al.
Pubblicazione: (2025)
Assessing biomedical knowledge robustness in large language models by query-efficient sampling attacks
di: Xian, R. Patrick, et al.
Pubblicazione: (2024)
di: Xian, R. Patrick, et al.
Pubblicazione: (2024)
Reliable agent engineering should integrate machine-compatible organizational principles
di: Xian, R. Patrick, et al.
Pubblicazione: (2025)
di: Xian, R. Patrick, et al.
Pubblicazione: (2025)
7T MRI Synthesization from 3T Acquisitions
di: Cui, Qiming, et al.
Pubblicazione: (2024)
di: Cui, Qiming, et al.
Pubblicazione: (2024)
Language model developers should report train-test overlap
di: Zhang, Andy K, et al.
Pubblicazione: (2024)
di: Zhang, Andy K, et al.
Pubblicazione: (2024)
MARD: A Multi-Agent Framework for Robust Android Malware Detection
di: Zeng, Xueying, et al.
Pubblicazione: (2026)
di: Zeng, Xueying, et al.
Pubblicazione: (2026)
Informed and Assessable Observability Design Decisions in Cloud-native Microservice Applications
di: Borges, Maria C., et al.
Pubblicazione: (2024)
di: Borges, Maria C., et al.
Pubblicazione: (2024)
Robust Mutation Analysis of Quantum Programs Under Noise
di: Fortz, Sophie, et al.
Pubblicazione: (2026)
di: Fortz, Sophie, et al.
Pubblicazione: (2026)
Semantic Grounding of Digital Twin Metamodels Using RDF Graphs
di: Abbasi, Faima, et al.
Pubblicazione: (2025)
di: Abbasi, Faima, et al.
Pubblicazione: (2025)
NESSiE: The Necessary Safety Benchmark -- Identifying Errors that should not Exist
di: Bertram, Johannes, et al.
Pubblicazione: (2026)
di: Bertram, Johannes, et al.
Pubblicazione: (2026)
Choosing and Using Text-to-Speech Software
di: Peters, Tom, et al.
Pubblicazione: (2007)
di: Peters, Tom, et al.
Pubblicazione: (2007)
RITFIS: Robust input testing framework for LLMs-based intelligent software
di: Xiao, Mingxuan, et al.
Pubblicazione: (2024)
di: Xiao, Mingxuan, et al.
Pubblicazione: (2024)
Stability prediction of the software requirements specification
di: del Sagrado, J., et al.
Pubblicazione: (2024)
di: del Sagrado, J., et al.
Pubblicazione: (2024)
Augmenting unit test suites from integration tests
di: Paltoglou, Katerina, et al.
Pubblicazione: (2026)
di: Paltoglou, Katerina, et al.
Pubblicazione: (2026)
METRION: A Framework for Accurate Software Energy Measurement
di: Weigell, Benjamin, et al.
Pubblicazione: (2025)
di: Weigell, Benjamin, et al.
Pubblicazione: (2025)
No Free Lunch: Research Software Testing in Teaching
di: Dorner, Michael, et al.
Pubblicazione: (2024)
di: Dorner, Michael, et al.
Pubblicazione: (2024)
Teralizer: Semantics-Based Test Generalization from Conventional Unit Tests to Property-Based Tests
di: Glock, Johann, et al.
Pubblicazione: (2025)
di: Glock, Johann, et al.
Pubblicazione: (2025)
Program Decomposition and Translation with Static Analysis
di: Ibrahimzada, Ali Reza
Pubblicazione: (2024)
di: Ibrahimzada, Ali Reza
Pubblicazione: (2024)
Code Collaborate: Dissecting Team Dynamics in First-Semester Programming Students
di: Berrezueta-Guzman, Santiago, et al.
Pubblicazione: (2024)
di: Berrezueta-Guzman, Santiago, et al.
Pubblicazione: (2024)
OXN -- Automated Observability Assessments for Cloud-Native Applications
di: Borges, Maria C., et al.
Pubblicazione: (2024)
di: Borges, Maria C., et al.
Pubblicazione: (2024)
Enhancing Deployment-Time Predictive Model Robustness for Code Analysis and Optimization
di: Wang, Huanting, et al.
Pubblicazione: (2024)
di: Wang, Huanting, et al.
Pubblicazione: (2024)
Agile Retrospectives: What went well? What didn't go well? What should we do?
di: Spichkova, Maria, et al.
Pubblicazione: (2025)
di: Spichkova, Maria, et al.
Pubblicazione: (2025)
Demystifying Device-specific Compatibility Issues in Android Apps
di: Chen, Junfeng, et al.
Pubblicazione: (2024)
di: Chen, Junfeng, et al.
Pubblicazione: (2024)
An Effective Approach to Embedding Source Code by Combining Large Language and Sentence Embedding Models
di: Xian, Zixiang, et al.
Pubblicazione: (2024)
di: Xian, Zixiang, et al.
Pubblicazione: (2024)
Quality attributes of test cases and test suites -- importance & challenges from practitioners' perspectives
di: Tran, Huynh Khanh Vi, et al.
Pubblicazione: (2025)
di: Tran, Huynh Khanh Vi, et al.
Pubblicazione: (2025)
A Soundness and Precision Benchmark for Java Debloating Tools
di: Klauke, Jonas, et al.
Pubblicazione: (2025)
di: Klauke, Jonas, et al.
Pubblicazione: (2025)
A Multi-Layer Testing Framework for Automated Data Quality Assurance in Cloud-Native ELT Pipelines
di: Gargouri, Ismail, et al.
Pubblicazione: (2026)
di: Gargouri, Ismail, et al.
Pubblicazione: (2026)
Evaluating Design Conformance Through Trace Comparison
di: Anderson, Reid, et al.
Pubblicazione: (2026)
di: Anderson, Reid, et al.
Pubblicazione: (2026)
Dr. Boot: Bootstrapping Program Synthesis Language Models to Perform Repairing
di: van der Vleuten, Noah
Pubblicazione: (2025)
di: van der Vleuten, Noah
Pubblicazione: (2025)
Generating Accurate OpenAPI Descriptions from Java Source Code
di: Lercher, Alexander, et al.
Pubblicazione: (2024)
di: Lercher, Alexander, et al.
Pubblicazione: (2024)
Model-driven realization of IDTA submodel specifications: The good, the bad, the incompatible?
di: Eichelberger, Holger, et al.
Pubblicazione: (2024)
di: Eichelberger, Holger, et al.
Pubblicazione: (2024)
Bridging the Gap Between Domain-specific Frameworks and Multiple Hardware Devices
di: Wen, Xu, et al.
Pubblicazione: (2024)
di: Wen, Xu, et al.
Pubblicazione: (2024)
The role of slicing in test-driven development
di: Dieste, Oscar, et al.
Pubblicazione: (2024)
di: Dieste, Oscar, et al.
Pubblicazione: (2024)
An Overview and Catalogue of Dependency Challenges in Open Source Software Package Registries
di: Mens, Tom, et al.
Pubblicazione: (2024)
di: Mens, Tom, et al.
Pubblicazione: (2024)
Same App, Different Behaviors: Uncovering Device-specific Behaviors in Android Apps
di: Dong, Zikan, et al.
Pubblicazione: (2024)
di: Dong, Zikan, et al.
Pubblicazione: (2024)
Opportunities in deep learning methods development for computational biology
di: Lee, Alex Jihun, et al.
Pubblicazione: (2024)
di: Lee, Alex Jihun, et al.
Pubblicazione: (2024)
Observation-based unit test generation at Meta
di: Alshahwan, Nadia, et al.
Pubblicazione: (2024)
di: Alshahwan, Nadia, et al.
Pubblicazione: (2024)
On (Mis)perceptions of testing effectiveness: an empirical study
di: Vegas, Sira, et al.
Pubblicazione: (2024)
di: Vegas, Sira, et al.
Pubblicazione: (2024)
Example-driven development: bridging tests and documentation
di: Nierstrasz, Oscar, et al.
Pubblicazione: (2024)
di: Nierstrasz, Oscar, et al.
Pubblicazione: (2024)
CODEPROMPTZIP: Code-specific Prompt Compression for Retrieval-Augmented Generation in Coding Tasks with LMs
di: He, Pengfei, et al.
Pubblicazione: (2025)
di: He, Pengfei, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Measuring temporal effects of agent knowledge by date-controlled tool use
di: Xian, R. Patrick, et al.
Pubblicazione: (2025) -
Assessing biomedical knowledge robustness in large language models by query-efficient sampling attacks
di: Xian, R. Patrick, et al.
Pubblicazione: (2024) -
Reliable agent engineering should integrate machine-compatible organizational principles
di: Xian, R. Patrick, et al.
Pubblicazione: (2025) -
7T MRI Synthesization from 3T Acquisitions
di: Cui, Qiming, et al.
Pubblicazione: (2024) -
Language model developers should report train-test overlap
di: Zhang, Andy K, et al.
Pubblicazione: (2024)