Scalable Stewardship of an LLM-Assisted Clinical Benchmark with Physician Oversight
Fuente:
arXiv
Guardado en:
| Autores principales: | Ye, Junze, Tawfik, Daniel, Goodell, Alex J., Kotha, Nikhil V., Buyyounouski, Mark K., Bayati, Mohsen |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Calibrating Conservatism for Scalable Oversight
por: Overman, William, et al.
Publicado: (2026)
por: Overman, William, et al.
Publicado: (2026)
Scaling Clinician-Grade Feature Generation from Clinical Notes with Multi-Agent Language Models
por: Wang, Jiayi, et al.
Publicado: (2025)
por: Wang, Jiayi, et al.
Publicado: (2025)
The Oversight Game: Learning to Cooperatively Balance an AI Agent's Safety and Autonomy
por: Overman, William, et al.
Publicado: (2025)
por: Overman, William, et al.
Publicado: (2025)
Efficient Detection of Bad Benchmark Items with Novel Scalability Coefficients
por: Hardy, Michael, et al.
Publicado: (2026)
por: Hardy, Michael, et al.
Publicado: (2026)
Rectification Difficulty and Optimal Sample Allocation in LLM-Augmented Surveys
por: Ye, Zikun, et al.
Publicado: (2026)
por: Ye, Zikun, et al.
Publicado: (2026)
Towards Reliable LLM Evaluation: Correcting the Winner's Curse in Adaptive Benchmarking
por: Xu, Yang, et al.
Publicado: (2026)
por: Xu, Yang, et al.
Publicado: (2026)
SEED-SET: Scalable Evolving Experimental Design for System-level Ethical Testing
por: Parashar, Anjali, et al.
Publicado: (2026)
por: Parashar, Anjali, et al.
Publicado: (2026)
RJUA-MedDQA: A Multimodal Benchmark for Medical Document Question Answering and Clinical Reasoning
por: Jin, Congyun, et al.
Publicado: (2024)
por: Jin, Congyun, et al.
Publicado: (2024)
Integrating Dynamic Correlation Shifts and Weighted Benchmarking in Extreme Value Analysis
por: Panagoulias, Dimitrios P., et al.
Publicado: (2024)
por: Panagoulias, Dimitrios P., et al.
Publicado: (2024)
Ranking Policy Learning via Marketplace Expected Value Estimation From Observational Data
por: Ebrahimzadeh, Ehsan, et al.
Publicado: (2024)
por: Ebrahimzadeh, Ehsan, et al.
Publicado: (2024)
AI-Assisted Decision-Making for Clinical Assessment of Auto-Segmented Contour Quality
por: Wang, Biling, et al.
Publicado: (2025)
por: Wang, Biling, et al.
Publicado: (2025)
StatLLM: A Dataset for Evaluating the Performance of Large Language Models in Statistical Analysis
por: Song, Xinyi, et al.
Publicado: (2025)
por: Song, Xinyi, et al.
Publicado: (2025)
A Benchmark for Scalable Oversight Protocols
por: Sudhir, Abhimanyu Pallavi, et al.
Publicado: (2025)
por: Sudhir, Abhimanyu Pallavi, et al.
Publicado: (2025)
Simulating and Experimenting with Social Media Mobilization Using LLM Agents
por: Shirani, Sadegh, et al.
Publicado: (2025)
por: Shirani, Sadegh, et al.
Publicado: (2025)
Decomposing Physician Disagreement in HealthBench
por: Borgohain, Satya, et al.
Publicado: (2026)
por: Borgohain, Satya, et al.
Publicado: (2026)
Incorporating LLM Embeddings for Variation Across the Human Genome
por: Niu, Hongqian, et al.
Publicado: (2025)
por: Niu, Hongqian, et al.
Publicado: (2025)
Eligibility-Aware Evidence Synthesis: An Agentic Framework for Clinical Trial Meta-Analysis
por: Zhao, Yao, et al.
Publicado: (2026)
por: Zhao, Yao, et al.
Publicado: (2026)
A network analysis of decision strategies of human experts in steel manufacturing
por: Merten, Daniel Christopher, et al.
Publicado: (2021)
por: Merten, Daniel Christopher, et al.
Publicado: (2021)
Post Launch Evaluation of Policies in a High-Dimensional Setting
por: Nassiri, Shima, et al.
Publicado: (2024)
por: Nassiri, Shima, et al.
Publicado: (2024)
Scalable Spatiotemporal Prediction with Bayesian Neural Fields
por: Saad, Feras, et al.
Publicado: (2024)
por: Saad, Feras, et al.
Publicado: (2024)
Improving LLM Leaderboards with Psychometrical Methodology
por: Federiakin, Denis
Publicado: (2025)
por: Federiakin, Denis
Publicado: (2025)
AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems
por: Hardy, Michael, et al.
Publicado: (2026)
por: Hardy, Michael, et al.
Publicado: (2026)
Quantitative Technology Forecasting: a Review of Trend Extrapolation Methods
por: Tsai, Peng-Hung, et al.
Publicado: (2024)
por: Tsai, Peng-Hung, et al.
Publicado: (2024)
How Should We Represent History in Interpretable Models of Clinical Policies?
por: Matsson, Anton, et al.
Publicado: (2024)
por: Matsson, Anton, et al.
Publicado: (2024)
Conformal Arbitrage: Risk-Controlled Balancing of Competing Objectives in Language Models
por: Overman, William, et al.
Publicado: (2025)
por: Overman, William, et al.
Publicado: (2025)
AI-Assisted Conversational Interviewing: Effects on Data Quality and Respondent Experience
por: Barari, Soubhik, et al.
Publicado: (2025)
por: Barari, Soubhik, et al.
Publicado: (2025)
DeepScore: A Comprehensive Approach to Measuring Quality in AI-Generated Clinical Documentation
por: Oleson, Jon
Publicado: (2024)
por: Oleson, Jon
Publicado: (2024)
Prune 'n Predict: Optimizing LLM Decision-making with Conformal Prediction
por: Vishwakarma, Harit, et al.
Publicado: (2024)
por: Vishwakarma, Harit, et al.
Publicado: (2024)
ImplicitRM: Unbiased Reward Modeling from Implicit Preference Data for LLM alignment
por: Wang, Hao, et al.
Publicado: (2026)
por: Wang, Hao, et al.
Publicado: (2026)
Wafer-Level Etch Spatial Profiling for Process Monitoring from Time-Series with Time-LLM
por: Kim, Hyunwoo, et al.
Publicado: (2026)
por: Kim, Hyunwoo, et al.
Publicado: (2026)
On the Mechanistic Interpretability of Neural Networks for Causality in Bio-statistics
por: Conan, Jean-Baptiste A.
Publicado: (2025)
por: Conan, Jean-Baptiste A.
Publicado: (2025)
Decision Quality Evaluation Framework at Pinterest
por: Tian, Yuqi, et al.
Publicado: (2026)
por: Tian, Yuqi, et al.
Publicado: (2026)
HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics
por: Luettgau, Lennart, et al.
Publicado: (2025)
por: Luettgau, Lennart, et al.
Publicado: (2025)
Data-Driven Bayesian Network Models of Hurricane Evacuation Decision Making
por: Wang, Hui Sophie, et al.
Publicado: (2023)
por: Wang, Hui Sophie, et al.
Publicado: (2023)
Decade-long Emission Forecasting with an Ensemble Model in Taiwan
por: Hung, Gordon, et al.
Publicado: (2025)
por: Hung, Gordon, et al.
Publicado: (2025)
ChatGPT and post-test probability
por: Weisenthal, Samuel J.
Publicado: (2023)
por: Weisenthal, Samuel J.
Publicado: (2023)
Process-Aware Analysis of Treatment Paths in Heart Failure Patients: A Case Study
por: Beyel, Harry H., et al.
Publicado: (2024)
por: Beyel, Harry H., et al.
Publicado: (2024)
Surrogate-Based Prevalence Measurement for Large-Scale A/B Testing
por: Xu, Zehao, et al.
Publicado: (2026)
por: Xu, Zehao, et al.
Publicado: (2026)
Calculating Customer Lifetime Value and Churn using Beta Geometric Negative Binomial and Gamma-Gamma Distribution in a NFT based setting
por: Das, Sagarnil
Publicado: (2025)
por: Das, Sagarnil
Publicado: (2025)
Unlocking the Potential of Past Research: Using Generative AI to Reconstruct Healthcare Simulation Models
por: Monks, Thomas, et al.
Publicado: (2025)
por: Monks, Thomas, et al.
Publicado: (2025)
Ejemplares similares
-
Calibrating Conservatism for Scalable Oversight
por: Overman, William, et al.
Publicado: (2026) -
Scaling Clinician-Grade Feature Generation from Clinical Notes with Multi-Agent Language Models
por: Wang, Jiayi, et al.
Publicado: (2025) -
The Oversight Game: Learning to Cooperatively Balance an AI Agent's Safety and Autonomy
por: Overman, William, et al.
Publicado: (2025) -
Efficient Detection of Bad Benchmark Items with Novel Scalability Coefficients
por: Hardy, Michael, et al.
Publicado: (2026) -
Rectification Difficulty and Optimal Sample Allocation in LLM-Augmented Surveys
por: Ye, Zikun, et al.
Publicado: (2026)