Five Years of SciCap: What We Learned and Future Directions for Scientific Figure Captioning
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Ting-Hao 'Kenneth', Rossi, Ryan A., Kim, Sungchul, Yu, Tong, Hsu, Ting-Yao E., Yin, Ho, Ng, Giles, C. Lee |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Do Large Multimodal Models Solve Caption Generation for Scientific Figures? Lessons Learned from SciCap Challenge 2023
by: Hsu, Ting-Yao E., et al.
Published: (2025)
by: Hsu, Ting-Yao E., et al.
Published: (2025)
SciCapenter: Supporting Caption Composition for Scientific Figures with Machine-Generated Captions and Ratings
by: Hsu, Ting-Yao, et al.
Published: (2024)
by: Hsu, Ting-Yao, et al.
Published: (2024)
LaMP-Cap: Personalized Figure Caption Generation With Multimodal Figure Profiles
by: Ng, Ho Yin 'Sam', et al.
Published: (2025)
by: Ng, Ho Yin 'Sam', et al.
Published: (2025)
Understanding How Paper Writers Use AI-Generated Captions in Figure Caption Writing
by: Yin, Ho, et al.
Published: (2025)
by: Yin, Ho, et al.
Published: (2025)
Multi-LLM Collaborative Caption Generation in Scientific Documents
by: Kim, Jaeyoung, et al.
Published: (2025)
by: Kim, Jaeyoung, et al.
Published: (2025)
Leveraging Author-Specific Context for Scientific Figure Caption Generation: 3rd SciCap Challenge
by: Timklaypachara, Watcharapong, et al.
Published: (2025)
by: Timklaypachara, Watcharapong, et al.
Published: (2025)
FigCaps-HF: A Figure-to-Caption Generative Framework and Benchmark with Human Feedback
by: Singh, Ashish, et al.
Published: (2023)
by: Singh, Ashish, et al.
Published: (2023)
What Color Scheme is More Effective in Assisting Readers to Locate Information in a Color-Coded Article?
by: Ng, Ho Yin, et al.
Published: (2024)
by: Ng, Ho Yin, et al.
Published: (2024)
Figuring out Figures: Using Textual References to Caption Scientific Figures
by: Cao, Stanley, et al.
Published: (2024)
by: Cao, Stanley, et al.
Published: (2024)
Federated Large Language Models: Current Progress and Future Directions
by: Yao, Yuhang, et al.
Published: (2024)
by: Yao, Yuhang, et al.
Published: (2024)
SuperCap: Multi-resolution Superpixel-based Image Captioning
by: Senior, Henry, et al.
Published: (2025)
by: Senior, Henry, et al.
Published: (2025)
Learning to Reduce: Optimal Representations of Structured Data in Prompting Large Language Models
by: Lee, Younghun, et al.
Published: (2024)
by: Lee, Younghun, et al.
Published: (2024)
Learning to Reduce: Towards Improving Performance of Large Language Models on Structured Data
by: Lee, Younghun, et al.
Published: (2024)
by: Lee, Younghun, et al.
Published: (2024)
Charts Are Not Images: On the Challenges of Scientific Chart Editing
by: Li, Shawn, et al.
Published: (2025)
by: Li, Shawn, et al.
Published: (2025)
Video ReCap: Recursive Captioning of Hour-Long Videos
by: Islam, Md Mohaiminul, et al.
Published: (2024)
by: Islam, Md Mohaiminul, et al.
Published: (2024)
Enhancing Scientific Figure Captioning Through Cross-modal Learning
by: Rojas, Mateo Alejandro, et al.
Published: (2024)
by: Rojas, Mateo Alejandro, et al.
Published: (2024)
Personalized Scientific Figure Caption Generation: An Empirical Study on Author-Specific Writing Style Transfer
by: Kim, Jaeyoung, et al.
Published: (2025)
by: Kim, Jaeyoung, et al.
Published: (2025)
Multi-Agent Collaborative Filtering: Orchestrating Users and Items for Agentic Recommendations
by: Xia, Yu, et al.
Published: (2025)
by: Xia, Yu, et al.
Published: (2025)
FigEx2: Visual-Conditioned Panel Detection and Captioning for Scientific Compound Figures
by: Song, Jifeng, et al.
Published: (2026)
by: Song, Jifeng, et al.
Published: (2026)
ChartCap: Mitigating Hallucination of Dense Chart Captioning
by: Lim, Junyoung, et al.
Published: (2025)
by: Lim, Junyoung, et al.
Published: (2025)
Scientific Hypothesis Generation and Validation: Methods, Datasets, and Future Directions
by: Kulkarni, Adithya, et al.
Published: (2025)
by: Kulkarni, Adithya, et al.
Published: (2025)
CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era
by: Cheng, Kanzhi, et al.
Published: (2025)
by: Cheng, Kanzhi, et al.
Published: (2025)
CapGeo: A Caption-Assisted Approach to Geometric Reasoning
by: Li, Yuying, et al.
Published: (2025)
by: Li, Yuying, et al.
Published: (2025)
SAND: Boosting LLM Agents with Self-Taught Action Deliberation
by: Xia, Yu, et al.
Published: (2025)
by: Xia, Yu, et al.
Published: (2025)
SciFIBench: Benchmarking Large Multimodal Models for Scientific Figure Interpretation
by: Roberts, Jonathan, et al.
Published: (2024)
by: Roberts, Jonathan, et al.
Published: (2024)
Diversify-verify-adapt: Efficient and Robust Retrieval-Augmented Ambiguous Question Answering
by: In, Yeonjun, et al.
Published: (2024)
by: In, Yeonjun, et al.
Published: (2024)
ControlCap: Controllable Region-level Captioning
by: Zhao, Yuzhong, et al.
Published: (2024)
by: Zhao, Yuzhong, et al.
Published: (2024)
The Solution for the ICCV 2023 1st Scientific Figure Captioning Challenge
by: Chao, Dian, et al.
Published: (2024)
by: Chao, Dian, et al.
Published: (2024)
Using Contextually Aligned Online Reviews to Measure LLMs' Performance Disparities Across Language Varieties
by: Tang, Zixin, et al.
Published: (2025)
by: Tang, Zixin, et al.
Published: (2025)
IF-VidCap: Can Video Caption Models Follow Instructions?
by: Li, Shihao, et al.
Published: (2025)
by: Li, Shihao, et al.
Published: (2025)
SciFigDetect: A Benchmark for AI-Generated Scientific Figure Detection
by: Hu, You, et al.
Published: (2026)
by: Hu, You, et al.
Published: (2026)
What Makes for Good Image Captions?
by: Chen, Delong, et al.
Published: (2024)
by: Chen, Delong, et al.
Published: (2024)
SciDA: Scientific Dynamic Assessor of LLMs
by: Zhou, Junting, et al.
Published: (2025)
by: Zhou, Junting, et al.
Published: (2025)
SnapCap: Efficient Snapshot Compressive Video Captioning
by: Sun, Jianqiao, et al.
Published: (2024)
by: Sun, Jianqiao, et al.
Published: (2024)
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing
by: Xing, Long, et al.
Published: (2025)
by: Xing, Long, et al.
Published: (2025)
CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning
by: Xing, Long, et al.
Published: (2025)
by: Xing, Long, et al.
Published: (2025)
GroundCap: A Visually Grounded Image Captioning Dataset
by: Oliveira, Daniel A. P., et al.
Published: (2025)
by: Oliveira, Daniel A. P., et al.
Published: (2025)
FinCap: Topic-Aligned Captions for Short-Form Financial YouTube Videos
by: Sukhani, Siddhant, et al.
Published: (2025)
by: Sukhani, Siddhant, et al.
Published: (2025)
SciClaimEval: Cross-modal Claim Verification in Scientific Papers
by: Ho, Xanh, et al.
Published: (2026)
by: Ho, Xanh, et al.
Published: (2026)
Unleashing Text-to-Image Diffusion Prior for Zero-Shot Image Captioning
by: Luo, Jianjie, et al.
Published: (2024)
by: Luo, Jianjie, et al.
Published: (2024)
Similar Items
-
Do Large Multimodal Models Solve Caption Generation for Scientific Figures? Lessons Learned from SciCap Challenge 2023
by: Hsu, Ting-Yao E., et al.
Published: (2025) -
SciCapenter: Supporting Caption Composition for Scientific Figures with Machine-Generated Captions and Ratings
by: Hsu, Ting-Yao, et al.
Published: (2024) -
LaMP-Cap: Personalized Figure Caption Generation With Multimodal Figure Profiles
by: Ng, Ho Yin 'Sam', et al.
Published: (2025) -
Understanding How Paper Writers Use AI-Generated Captions in Figure Caption Writing
by: Yin, Ho, et al.
Published: (2025) -
Multi-LLM Collaborative Caption Generation in Scientific Documents
by: Kim, Jaeyoung, et al.
Published: (2025)