Advancing AI Research Assistants with Expert-Involved Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Tianyu, Han, Simeng, Wang, Hanchen, Luo, Xiao, Lu, Pan, Zhu, Biqing, Wang, Yuge, Li, Keyi, Chen, Jiapeng, Qu, Rihao, Liu, Yufeng, Cui, Xinyue, Yaish, Aviv, Chen, Yuhang, Hao, Minsheng, Li, Chuhan, Li, Kexing, Lu, Yinsheng, Wei, Xinyu, Xing, Qinzhe, Panescu, Antonia, Wang, Mengbo, Annaswamy, Vibha, Sanchez, Alicia, Cloherty, Jack, Cohan, Arman, Xu, Hua, Gerstein, Mark, Zou, James, Zhao, Hongyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911570585452544
author Liu, Tianyu
Han, Simeng
Wang, Hanchen
Luo, Xiao
Lu, Pan
Zhu, Biqing
Wang, Yuge
Li, Keyi
Chen, Jiapeng
Qu, Rihao
Liu, Yufeng
Cui, Xinyue
Yaish, Aviv
Chen, Yuhang
Hao, Minsheng
Li, Chuhan
Li, Kexing
Lu, Yinsheng
Wei, Xinyu
Xing, Qinzhe
Panescu, Antonia
Wang, Mengbo
Annaswamy, Vibha
Sanchez, Alicia
Cloherty, Jack
Cohan, Arman
Xu, Hua
Gerstein, Mark
Zou, James
Zhao, Hongyu
author_facet Liu, Tianyu
Han, Simeng
Wang, Hanchen
Luo, Xiao
Lu, Pan
Zhu, Biqing
Wang, Yuge
Li, Keyi
Chen, Jiapeng
Qu, Rihao
Liu, Yufeng
Cui, Xinyue
Yaish, Aviv
Chen, Yuhang
Hao, Minsheng
Li, Chuhan
Li, Kexing
Lu, Yinsheng
Wei, Xinyu
Xing, Qinzhe
Panescu, Antonia
Wang, Mengbo
Annaswamy, Vibha
Sanchez, Alicia
Cloherty, Jack
Cohan, Arman
Xu, Hua
Gerstein, Mark
Zou, James
Zhao, Hongyu
contents Large language models (LLMs) and large multimodal models (LMMs) promise to accelerate biomedical discovery, yet their reliability remains unclear. We introduce ARIEL (AI Research Assistant for Expert-in-the-Loop Learning), an open-source evaluation and optimization framework that pairs a curated multimodal biomedical corpus with expert-vetted tasks to probe two capabilities: full-length article summarization and fine-grained figure interpretation. Using uniform protocols and blinded PhD-level evaluation, we find that state-of-the-art models generate fluent but incomplete summaries, whereas LMMs struggle with detailed visual reasoning. We later observe that prompt engineering and lightweight fine-tuning substantially improve textual coverage, and a compute-scaled inference strategy enhances visual question answering. We build an ARIEL agent that integrates textual and visual cues, and we show it can propose testable mechanistic hypotheses. ARIEL delineates current strengths and limitations of foundation models, and provides a reproducible platform for advancing trustworthy AI in biomedicine.
format Preprint
id arxiv_https___arxiv_org_abs_2505_04638
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Advancing AI Research Assistants with Expert-Involved Learning
Liu, Tianyu
Han, Simeng
Wang, Hanchen
Luo, Xiao
Lu, Pan
Zhu, Biqing
Wang, Yuge
Li, Keyi
Chen, Jiapeng
Qu, Rihao
Liu, Yufeng
Cui, Xinyue
Yaish, Aviv
Chen, Yuhang
Hao, Minsheng
Li, Chuhan
Li, Kexing
Lu, Yinsheng
Wei, Xinyu
Xing, Qinzhe
Panescu, Antonia
Wang, Mengbo
Annaswamy, Vibha
Sanchez, Alicia
Cloherty, Jack
Cohan, Arman
Xu, Hua
Gerstein, Mark
Zou, James
Zhao, Hongyu
Artificial Intelligence
Computation and Language
Information Retrieval
Large language models (LLMs) and large multimodal models (LMMs) promise to accelerate biomedical discovery, yet their reliability remains unclear. We introduce ARIEL (AI Research Assistant for Expert-in-the-Loop Learning), an open-source evaluation and optimization framework that pairs a curated multimodal biomedical corpus with expert-vetted tasks to probe two capabilities: full-length article summarization and fine-grained figure interpretation. Using uniform protocols and blinded PhD-level evaluation, we find that state-of-the-art models generate fluent but incomplete summaries, whereas LMMs struggle with detailed visual reasoning. We later observe that prompt engineering and lightweight fine-tuning substantially improve textual coverage, and a compute-scaled inference strategy enhances visual question answering. We build an ARIEL agent that integrates textual and visual cues, and we show it can propose testable mechanistic hypotheses. ARIEL delineates current strengths and limitations of foundation models, and provides a reproducible platform for advancing trustworthy AI in biomedicine.
title Advancing AI Research Assistants with Expert-Involved Learning
topic Artificial Intelligence
Computation and Language
Information Retrieval
url https://arxiv.org/abs/2505.04638