Advancing AI Research Assistants with Expert-Involved Learning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911570585452544 |
|---|---|
| author | Liu, Tianyu Han, Simeng Wang, Hanchen Luo, Xiao Lu, Pan Zhu, Biqing Wang, Yuge Li, Keyi Chen, Jiapeng Qu, Rihao Liu, Yufeng Cui, Xinyue Yaish, Aviv Chen, Yuhang Hao, Minsheng Li, Chuhan Li, Kexing Lu, Yinsheng Wei, Xinyu Xing, Qinzhe Panescu, Antonia Wang, Mengbo Annaswamy, Vibha Sanchez, Alicia Cloherty, Jack Cohan, Arman Xu, Hua Gerstein, Mark Zou, James Zhao, Hongyu |
| author_facet | Liu, Tianyu Han, Simeng Wang, Hanchen Luo, Xiao Lu, Pan Zhu, Biqing Wang, Yuge Li, Keyi Chen, Jiapeng Qu, Rihao Liu, Yufeng Cui, Xinyue Yaish, Aviv Chen, Yuhang Hao, Minsheng Li, Chuhan Li, Kexing Lu, Yinsheng Wei, Xinyu Xing, Qinzhe Panescu, Antonia Wang, Mengbo Annaswamy, Vibha Sanchez, Alicia Cloherty, Jack Cohan, Arman Xu, Hua Gerstein, Mark Zou, James Zhao, Hongyu |
| contents | Large language models (LLMs) and large multimodal models (LMMs) promise to accelerate biomedical discovery, yet their reliability remains unclear. We introduce ARIEL (AI Research Assistant for Expert-in-the-Loop Learning), an open-source evaluation and optimization framework that pairs a curated multimodal biomedical corpus with expert-vetted tasks to probe two capabilities: full-length article summarization and fine-grained figure interpretation. Using uniform protocols and blinded PhD-level evaluation, we find that state-of-the-art models generate fluent but incomplete summaries, whereas LMMs struggle with detailed visual reasoning. We later observe that prompt engineering and lightweight fine-tuning substantially improve textual coverage, and a compute-scaled inference strategy enhances visual question answering. We build an ARIEL agent that integrates textual and visual cues, and we show it can propose testable mechanistic hypotheses. ARIEL delineates current strengths and limitations of foundation models, and provides a reproducible platform for advancing trustworthy AI in biomedicine. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_04638 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Advancing AI Research Assistants with Expert-Involved Learning Liu, Tianyu Han, Simeng Wang, Hanchen Luo, Xiao Lu, Pan Zhu, Biqing Wang, Yuge Li, Keyi Chen, Jiapeng Qu, Rihao Liu, Yufeng Cui, Xinyue Yaish, Aviv Chen, Yuhang Hao, Minsheng Li, Chuhan Li, Kexing Lu, Yinsheng Wei, Xinyu Xing, Qinzhe Panescu, Antonia Wang, Mengbo Annaswamy, Vibha Sanchez, Alicia Cloherty, Jack Cohan, Arman Xu, Hua Gerstein, Mark Zou, James Zhao, Hongyu Artificial Intelligence Computation and Language Information Retrieval Large language models (LLMs) and large multimodal models (LMMs) promise to accelerate biomedical discovery, yet their reliability remains unclear. We introduce ARIEL (AI Research Assistant for Expert-in-the-Loop Learning), an open-source evaluation and optimization framework that pairs a curated multimodal biomedical corpus with expert-vetted tasks to probe two capabilities: full-length article summarization and fine-grained figure interpretation. Using uniform protocols and blinded PhD-level evaluation, we find that state-of-the-art models generate fluent but incomplete summaries, whereas LMMs struggle with detailed visual reasoning. We later observe that prompt engineering and lightweight fine-tuning substantially improve textual coverage, and a compute-scaled inference strategy enhances visual question answering. We build an ARIEL agent that integrates textual and visual cues, and we show it can propose testable mechanistic hypotheses. ARIEL delineates current strengths and limitations of foundation models, and provides a reproducible platform for advancing trustworthy AI in biomedicine. |
| title | Advancing AI Research Assistants with Expert-Involved Learning |
| topic | Artificial Intelligence Computation and Language Information Retrieval |
| url | https://arxiv.org/abs/2505.04638 |