RAmBLA: A Framework for Evaluating the Reliability of LLMs as Assistants in the Biomedical Domain
Fuente:
arXiv
Saved in:
| Main Authors: | Bolton, William James, Poyiadzi, Rafael, Morrell, Edward R., Bueno, Gabriela van Bergen Gonzalez, Goetz, Lea |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Semantic-KG: Using Knowledge Graphs to Construct Benchmarks for Measuring Semantic Similarity
by: Wei, Qiyao, et al.
Published: (2025)
by: Wei, Qiyao, et al.
Published: (2025)
Evaluating Strategic Reasoning in Forecasting Agents
by: Liptay, Tom, et al.
Published: (2026)
by: Liptay, Tom, et al.
Published: (2026)
Pollen analysis of sediment profile BLA00 from Blankensee
by: Kleinmann, Angelika, et al.
Published: (2001)
by: Kleinmann, Angelika, et al.
Published: (2001)
Pollen analysis of sediment profile BLA0 from Blankensee
by: Kleinmann, Angelika, et al.
Published: (2002)
by: Kleinmann, Angelika, et al.
Published: (2002)
Inequities in naloxone administration among fatal overdose decedents by race and ethnicity in Pennsylvania, 2019–21
by: Erin Takemoto, et al.
Published: (2024)
by: Erin Takemoto, et al.
Published: (2024)
LA DERROTA DE BLA: FENOMENOLOGÍA DE ALGUNAS INTERACCIONES EDUCADORAS
by: Javier Orlando Lozano Escobar
Published: (2005)
by: Javier Orlando Lozano Escobar
Published: (2005)
LLMs to Support a Domain Specific Knowledge Assistant
by: Lovin, Maria-Flavia
Published: (2025)
by: Lovin, Maria-Flavia
Published: (2025)
The BLA Benchmark: Investigating Basic Language Abilities of Pre-Trained Multimodal Models
by: Chen, Xinyi, et al.
Published: (2023)
by: Chen, Xinyi, et al.
Published: (2023)
Chapter De raderbaar van militair-arts Cornelis de Mooy
by: van Bergen, Leo
Published: (2024)
by: van Bergen, Leo
Published: (2024)
Chapter De verpleegstersengel en de oorlogsduivel: De ‘gedachtenisprent’ ter ere van de Rode-Kruisambulances 1870-1871
by: van Bergen, Leo
Published: (2026)
by: van Bergen, Leo
Published: (2026)
IMM Paper 9: RPCS-1 — A Continuous Parametric Framework for Behavioral Profile Assessment
by: Bergen, Travis
Published: (2026)
by: Bergen, Travis
Published: (2026)
Comparative Physiology and Morphology of BLA‐Projecting NBM/SI Cholinergic Neurons in Mouse and Macaque
by: Feng Luo, et al.
Published: (2024)
by: Feng Luo, et al.
Published: (2024)
Fusion-Eval: Integrating Assistant Evaluators with LLMs
by: Shu, Lei, et al.
Published: (2023)
by: Shu, Lei, et al.
Published: (2023)
TREAT: A Code LLMs Trustworthiness / Reliability Evaluation and Testing Framework
by: Gao, Shuzheng, et al.
Published: (2025)
by: Gao, Shuzheng, et al.
Published: (2025)
On generalisability of segment anything model for nuclear instance segmentation in histology images
by: Xu, Kesi, et al.
Published: (2024)
by: Xu, Kesi, et al.
Published: (2024)
Natural Language Processing in the Patent Domain: A Survey
by: Jiang, Lekang, et al.
Published: (2024)
by: Jiang, Lekang, et al.
Published: (2024)
Distribution and abundance of nannofossils in DSDP Site 89-585
by: Bergen, James A
Published: (1986)
by: Bergen, James A
Published: (1986)
(Table 3) Distribution and abundance of early Cretaceous nannofossils in DSDP Hole 89-585A
by: Bergen, James A
Published: (1986)
by: Bergen, James A
Published: (1986)
(Table 4) Distribution and abundance of Cretaceous nannofossils in DSDP Hole 89-585
by: Bergen, James A
Published: (1986)
by: Bergen, James A
Published: (1986)
(Table 2) Distribution and abundance of late Cretaceous to early Paleocene nannofossils in DSDP Hole 89-585A
by: Bergen, James A
Published: (1986)
by: Bergen, James A
Published: (1986)
(Table 1) Distribution and abundance of late Cretaceous to middle Eocene nannofossils in DSDP Hole 89-585
by: Bergen, James A
Published: (1986)
by: Bergen, James A
Published: (1986)
(Table 5) Distribution and abundance of early Cretaceous nannofossils in DSDP Hole 89-585A
by: Bergen, James A
Published: (1986)
by: Bergen, James A
Published: (1986)
Benchmarking LLMs for Pairwise Causal Discovery in Biomedical and Multi-Domain Contexts
by: Anuyah, Sydney, et al.
Published: (2026)
by: Anuyah, Sydney, et al.
Published: (2026)
SOWING THE SACRED: MEXICAN PENTECOSTAL FARMWORKERS IN CALIFORNIA. By Lloyd DanielBarba. New York, NY: Oxford University Press, 2022. Pp. xix + 349. Paper, $29.95.
by: Emily Morrell
Published: (2024)
by: Emily Morrell
Published: (2024)
Beyond Ethical Alignment: Evaluating LLMs as Artificial Moral Assistants
by: Galatolo, Alessio, et al.
Published: (2025)
by: Galatolo, Alessio, et al.
Published: (2025)
Exploring the Potential of LLMs as Personalized Assistants: Dataset, Evaluation, and Analysis
by: Mok, Jisoo, et al.
Published: (2025)
by: Mok, Jisoo, et al.
Published: (2025)
AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Large Language Models
by: Jackson, Declan, et al.
Published: (2025)
by: Jackson, Declan, et al.
Published: (2025)
SimulatorArena: Are User Simulators Reliable Proxies for Multi-Turn Evaluation of AI Assistants?
by: Dou, Yao, et al.
Published: (2025)
by: Dou, Yao, et al.
Published: (2025)
MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants
by: Bansal, Hritik, et al.
Published: (2024)
by: Bansal, Hritik, et al.
Published: (2024)
SEOE: A Scalable and Reliable Semantic Evaluation Framework for Open Domain Event Detection
by: Lu, Yi-Fan, et al.
Published: (2025)
by: Lu, Yi-Fan, et al.
Published: (2025)
Individual Fairness Through Reweighting and Tuning
by: Mahamadou, Abdoul Jalil Djiberou, et al.
Published: (2024)
by: Mahamadou, Abdoul Jalil Djiberou, et al.
Published: (2024)
Evolving Roles of LLMs in Scientific Innovation: Assistant, Collaborator, Scientist, and Evaluator
by: Zhang, Haoxuan, et al.
Published: (2025)
by: Zhang, Haoxuan, et al.
Published: (2025)
Women's Studies Collections: A Checklist Evaluation
by: Bolton, Brooke A.
Published: (2009)
by: Bolton, Brooke A.
Published: (2009)
mGluR5 in EC CCK to BLA Circuit Modulates Depressive‐Like Phenotypes through CCK Signaling
by: Muhammad Asim, et al.
Published: (2026)
by: Muhammad Asim, et al.
Published: (2026)
User-Assistant Bias in LLMs
by: Pan, Xu, et al.
Published: (2025)
by: Pan, Xu, et al.
Published: (2025)
Evaluating a Pre-Session Homework Exercise in a Standalone Information Literacy Class
by: Goetz, Joseph E., et al.
Published: (2015)
by: Goetz, Joseph E., et al.
Published: (2015)
Pub-Guard-LLM: Detecting Retracted Biomedical Articles with Reliable Explanations
by: Chen, Lihu, et al.
Published: (2025)
by: Chen, Lihu, et al.
Published: (2025)
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
by: Lunardi, Riccardo, et al.
Published: (2025)
by: Lunardi, Riccardo, et al.
Published: (2025)
Newman, Stephen M. Feb. 17, 1911
by: Newman, Stephen Morrell
Published: (1911)
by: Newman, Stephen Morrell
Published: (1911)
Apsidal motion in binary systems as a tool for determination of stellar masses: system HD 93205 (O3V+O8V)
by: N. I. Morrell
Published: (2002)
by: N. I. Morrell
Published: (2002)
Similar Items
-
Semantic-KG: Using Knowledge Graphs to Construct Benchmarks for Measuring Semantic Similarity
by: Wei, Qiyao, et al.
Published: (2025) -
Evaluating Strategic Reasoning in Forecasting Agents
by: Liptay, Tom, et al.
Published: (2026) -
Pollen analysis of sediment profile BLA00 from Blankensee
by: Kleinmann, Angelika, et al.
Published: (2001) -
Pollen analysis of sediment profile BLA0 from Blankensee
by: Kleinmann, Angelika, et al.
Published: (2002) -
Inequities in naloxone administration among fatal overdose decedents by race and ethnicity in Pennsylvania, 2019–21
by: Erin Takemoto, et al.
Published: (2024)