Large language models require a new form of oversight: capability-based monitoring
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kellogg, Katherine C., Ye, Bingyang, Hu, Yifan, Savova, Guergana K., Wallace, Byron, Bitterman, Danielle S. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Assessing the Macro and Micro Effects of Random Seeds on Fine-Tuning Large Language Models
von: Bui, Nghia, et al.
Veröffentlicht: (2025)
von: Bui, Nghia, et al.
Veröffentlicht: (2025)
Evaluation of ChatGPT Family of Models for Biomedical Reasoning and Classification
von: Chen, Shan, et al.
Veröffentlicht: (2023)
von: Chen, Shan, et al.
Veröffentlicht: (2023)
Identifying Task Groupings for Multi-Task Learning Using Pointwise V-Usable Information
von: Li, Yingya, et al.
Veröffentlicht: (2024)
von: Li, Yingya, et al.
Veröffentlicht: (2024)
Artificial Intelligence in the Legal Field: Law Students Perspective
von: Andreeva, Daniela, et al.
Veröffentlicht: (2024)
von: Andreeva, Daniela, et al.
Veröffentlicht: (2024)
Large Language Models to Identify Social Determinants of Health in Electronic Health Records
von: Guevara, Marco, et al.
Veröffentlicht: (2023)
von: Guevara, Marco, et al.
Veröffentlicht: (2023)
Proof of Time: A Benchmark for Evaluating Scientific Idea Judgments
von: Ye, Bingyang, et al.
Veröffentlicht: (2026)
von: Ye, Bingyang, et al.
Veröffentlicht: (2026)
Auxiliary task demands mask the capabilities of smaller language models
von: Hu, Jennifer, et al.
Veröffentlicht: (2024)
von: Hu, Jennifer, et al.
Veröffentlicht: (2024)
AI Act and Large Language Models (LLMs): When critical issues and privacy impact require human and ethical oversight
von: Fabiano, Nicola
Veröffentlicht: (2024)
von: Fabiano, Nicola
Veröffentlicht: (2024)
A Field Guide to Deploying AI Agents in Clinical Practice
von: Gallifant, Jack, et al.
Veröffentlicht: (2025)
von: Gallifant, Jack, et al.
Veröffentlicht: (2025)
Do Automatic Factuality Metrics Measure Factuality? A Critical Evaluation
von: Ramprasad, Sanjana, et al.
Veröffentlicht: (2024)
von: Ramprasad, Sanjana, et al.
Veröffentlicht: (2024)
Confirmation bias: A challenge for scalable oversight
von: Recchia, Gabriel, et al.
Veröffentlicht: (2025)
von: Recchia, Gabriel, et al.
Veröffentlicht: (2025)
ZNO-Eval: Benchmarking reasoning capabilities of large language models in Ukrainian
von: Syromiatnikov, Mykyta, et al.
Veröffentlicht: (2025)
von: Syromiatnikov, Mykyta, et al.
Veröffentlicht: (2025)
Safety challenges of AI in medicine in the era of large language models
von: Wang, Xiaoye, et al.
Veröffentlicht: (2024)
von: Wang, Xiaoye, et al.
Veröffentlicht: (2024)
Towards physician-centered oversight of conversational diagnostic AI
von: Vedadi, Elahe, et al.
Veröffentlicht: (2025)
von: Vedadi, Elahe, et al.
Veröffentlicht: (2025)
Can large language models assist choice modelling? Insights into prompting strategies and current models capabilities
von: Sfeir, Georges, et al.
Veröffentlicht: (2025)
von: Sfeir, Georges, et al.
Veröffentlicht: (2025)
Long-form factuality in large language models
von: Wei, Jerry, et al.
Veröffentlicht: (2024)
von: Wei, Jerry, et al.
Veröffentlicht: (2024)
Fine-tuning large language models for domain adaptation: Exploration of training strategies, scaling, model merging and synergistic capabilities
von: Lu, Wei, et al.
Veröffentlicht: (2024)
von: Lu, Wei, et al.
Veröffentlicht: (2024)
WorldMedQA-V: a multilingual, multimodal medical examination dataset for multimodal language models evaluation
von: Matos, João, et al.
Veröffentlicht: (2024)
von: Matos, João, et al.
Veröffentlicht: (2024)
Automatically Extracting Numerical Results from Randomized Controlled Trials with Large Language Models
von: Yun, Hye Sun, et al.
Veröffentlicht: (2024)
von: Yun, Hye Sun, et al.
Veröffentlicht: (2024)
AI Epidemiology: achieving explainable AI through expert oversight patterns
von: Tempest-Walters, Kit
Veröffentlicht: (2025)
von: Tempest-Walters, Kit
Veröffentlicht: (2025)
Adapting Abstract Meaning Representation Parsing to the Clinical Narrative -- the SPRING THYME parser
von: Cai, Jon Z., et al.
Veröffentlicht: (2024)
von: Cai, Jon Z., et al.
Veröffentlicht: (2024)
The use of large language models to enhance cancer clinical trial educational materials
von: Gao, Mingye, et al.
Veröffentlicht: (2024)
von: Gao, Mingye, et al.
Veröffentlicht: (2024)
Evidence of interrelated cognitive-like capabilities in large language models: Indications of artificial general intelligence or achievement?
von: Ilić, David, et al.
Veröffentlicht: (2023)
von: Ilić, David, et al.
Veröffentlicht: (2023)
ClinicalGPT-R1: Pushing reasoning capability of generalist disease diagnosis with large language model
von: Lan, Wuyang, et al.
Veröffentlicht: (2025)
von: Lan, Wuyang, et al.
Veröffentlicht: (2025)
What a diff makes: automating code migration with large language models
von: Rosenfeld, Katherine A., et al.
Veröffentlicht: (2025)
von: Rosenfeld, Katherine A., et al.
Veröffentlicht: (2025)
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
von: Wijk, Hjalmar, et al.
Veröffentlicht: (2024)
von: Wijk, Hjalmar, et al.
Veröffentlicht: (2024)
Sparse Autoencoder Features for Classifications and Transferability
von: Gallifant, Jack, et al.
Veröffentlicht: (2025)
von: Gallifant, Jack, et al.
Veröffentlicht: (2025)
KScope: A Framework for Characterizing the Knowledge Status of Language Models
von: Xiao, Yuxin, et al.
Veröffentlicht: (2025)
von: Xiao, Yuxin, et al.
Veröffentlicht: (2025)
Medical Education and artificial intelligence: Responsible and effective practice requires human oversight
von: Kevin W. Eva
Veröffentlicht: (2024)
von: Kevin W. Eva
Veröffentlicht: (2024)
Assay2Mol: large language model-based drug design using BioAssay context
von: Deng, Yifan, et al.
Veröffentlicht: (2025)
von: Deng, Yifan, et al.
Veröffentlicht: (2025)
Simulated patient systems powered by large language model-based AI agents offer potential for transforming medical education
von: Yu, Huizi, et al.
Veröffentlicht: (2024)
von: Yu, Huizi, et al.
Veröffentlicht: (2024)
Evaluating the Factuality of Zero-shot Summarizers Across Varied Domains
von: Ramprasad, Sanjana, et al.
Veröffentlicht: (2024)
von: Ramprasad, Sanjana, et al.
Veröffentlicht: (2024)
Free-form language-based robotic reasoning and grasping
von: Jiao, Runyu, et al.
Veröffentlicht: (2025)
von: Jiao, Runyu, et al.
Veröffentlicht: (2025)
When Raw Data Prevails: Are Large Language Model Embeddings Effective in Numerical Data Representation for Medical Machine Learning Applications?
von: Gao, Yanjun, et al.
Veröffentlicht: (2024)
von: Gao, Yanjun, et al.
Veröffentlicht: (2024)
Large language models and linguistic intentionality
von: Grindrod, Jumbly
Veröffentlicht: (2024)
von: Grindrod, Jumbly
Veröffentlicht: (2024)
Inferring Pluggable Types with Machine Learning
von: Siddiqui, Kazi Amanul Islam, et al.
Veröffentlicht: (2024)
von: Siddiqui, Kazi Amanul Islam, et al.
Veröffentlicht: (2024)
The Dual-Route Model of Induction
von: Feucht, Sheridan, et al.
Veröffentlicht: (2025)
von: Feucht, Sheridan, et al.
Veröffentlicht: (2025)
Multimodal large language model for wheat breeding: a new exploration of smart breeding
von: Yang, Guofeng, et al.
Veröffentlicht: (2024)
von: Yang, Guofeng, et al.
Veröffentlicht: (2024)
Position Paper On Diagnostic Uncertainty Estimation from Large Language Models: Next-Word Probability Is Not Pre-test Probability
von: Gao, Yanjun, et al.
Veröffentlicht: (2024)
von: Gao, Yanjun, et al.
Veröffentlicht: (2024)
Over One Hundred Years' Service to Science
von: Savova, Elena
Veröffentlicht: (1971)
von: Savova, Elena
Veröffentlicht: (1971)
Ähnliche Einträge
-
Assessing the Macro and Micro Effects of Random Seeds on Fine-Tuning Large Language Models
von: Bui, Nghia, et al.
Veröffentlicht: (2025) -
Evaluation of ChatGPT Family of Models for Biomedical Reasoning and Classification
von: Chen, Shan, et al.
Veröffentlicht: (2023) -
Identifying Task Groupings for Multi-Task Learning Using Pointwise V-Usable Information
von: Li, Yingya, et al.
Veröffentlicht: (2024) -
Artificial Intelligence in the Legal Field: Law Students Perspective
von: Andreeva, Daniela, et al.
Veröffentlicht: (2024) -
Large Language Models to Identify Social Determinants of Health in Electronic Health Records
von: Guevara, Marco, et al.
Veröffentlicht: (2023)