Five Years of SciCap: What We Learned and Future Directions for Scientific Figure Captioning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Ting-Hao 'Kenneth', Rossi, Ryan A., Kim, Sungchul, Yu, Tong, Hsu, Ting-Yao E., Yin, Ho, Ng, Giles, C. Lee
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908768459030528
author Huang, Ting-Hao 'Kenneth'
Rossi, Ryan A.
Kim, Sungchul
Yu, Tong
Hsu, Ting-Yao E.
Yin, Ho
Ng
Giles, C. Lee
author_facet Huang, Ting-Hao 'Kenneth'
Rossi, Ryan A.
Kim, Sungchul
Yu, Tong
Hsu, Ting-Yao E.
Yin, Ho
Ng
Giles, C. Lee
contents Between 2021 and 2025, the SciCap project grew from a small seed-funded idea at The Pennsylvania State University (Penn State) into one of the central efforts shaping the scientific figure-captioning landscape. Supported by a Penn State seed grant, Adobe, and the Alfred P. Sloan Foundation, what began as our attempt to test whether domain-specific training, which was successful in text models like SciBERT, could also work for figure captions expanded into a multi-institution collaboration. Over these five years, we curated, released, and continually updated a large collection of figure-caption pairs from arXiv papers, conducted extensive automatic and human evaluations on both generated and author-written captions, navigated the rapid rise of large language models (LLMs), launched annual challenges, and built interactive systems that help scientists write better captions. In this piece, we look back at the first five years of SciCap and summarize the key technical and methodological lessons we learned. We then outline five major unsolved challenges and propose directions for the next phase of research in scientific figure captioning.
format Preprint
id arxiv_https___arxiv_org_abs_2512_21789
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Five Years of SciCap: What We Learned and Future Directions for Scientific Figure Captioning
Huang, Ting-Hao 'Kenneth'
Rossi, Ryan A.
Kim, Sungchul
Yu, Tong
Hsu, Ting-Yao E.
Yin, Ho
Ng
Giles, C. Lee
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Human-Computer Interaction
Between 2021 and 2025, the SciCap project grew from a small seed-funded idea at The Pennsylvania State University (Penn State) into one of the central efforts shaping the scientific figure-captioning landscape. Supported by a Penn State seed grant, Adobe, and the Alfred P. Sloan Foundation, what began as our attempt to test whether domain-specific training, which was successful in text models like SciBERT, could also work for figure captions expanded into a multi-institution collaboration. Over these five years, we curated, released, and continually updated a large collection of figure-caption pairs from arXiv papers, conducted extensive automatic and human evaluations on both generated and author-written captions, navigated the rapid rise of large language models (LLMs), launched annual challenges, and built interactive systems that help scientists write better captions. In this piece, we look back at the first five years of SciCap and summarize the key technical and methodological lessons we learned. We then outline five major unsolved challenges and propose directions for the next phase of research in scientific figure captioning.
title Five Years of SciCap: What We Learned and Future Directions for Scientific Figure Captioning
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Human-Computer Interaction
url https://arxiv.org/abs/2512.21789