Figuring out Figures: Using Textual References to Caption Scientific Figures
Fuente:
arXiv
Saved in:
| Main Authors: | Cao, Stanley, Liu, Kevin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LaMP-Cap: Personalized Figure Caption Generation With Multimodal Figure Profiles
by: Ng, Ho Yin 'Sam', et al.
Published: (2025)
by: Ng, Ho Yin 'Sam', et al.
Published: (2025)
FigEx2: Visual-Conditioned Panel Detection and Captioning for Scientific Compound Figures
by: Song, Jifeng, et al.
Published: (2026)
by: Song, Jifeng, et al.
Published: (2026)
Do Large Multimodal Models Solve Caption Generation for Scientific Figures? Lessons Learned from SciCap Challenge 2023
by: Hsu, Ting-Yao E., et al.
Published: (2025)
by: Hsu, Ting-Yao E., et al.
Published: (2025)
Five Years of SciCap: What We Learned and Future Directions for Scientific Figure Captioning
by: Huang, Ting-Hao 'Kenneth', et al.
Published: (2025)
by: Huang, Ting-Hao 'Kenneth', et al.
Published: (2025)
The Solution for the ICCV 2023 1st Scientific Figure Captioning Challenge
by: Chao, Dian, et al.
Published: (2024)
by: Chao, Dian, et al.
Published: (2024)
AutoFigure: Generating and Refining Publication-Ready Scientific Illustrations
by: Zhu, Minjun, et al.
Published: (2026)
by: Zhu, Minjun, et al.
Published: (2026)
Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs
by: Zhao, Haozhe, et al.
Published: (2026)
by: Zhao, Haozhe, et al.
Published: (2026)
Understanding Figurative Meaning through Explainable Visual Entailment
by: Saakyan, Arkadiy, et al.
Published: (2024)
by: Saakyan, Arkadiy, et al.
Published: (2024)
Every Part Matters: Integrity Verification of Scientific Figures Based on Multimodal Large Language Models
by: Shi, Xiang, et al.
Published: (2024)
by: Shi, Xiang, et al.
Published: (2024)
AutoFigure-Edit: Generating Editable Scientific Illustration
by: Lin, Zhen, et al.
Published: (2026)
by: Lin, Zhen, et al.
Published: (2026)
Personalized Scientific Figure Caption Generation: An Empirical Study on Author-Specific Writing Style Transfer
by: Kim, Jaeyoung, et al.
Published: (2025)
by: Kim, Jaeyoung, et al.
Published: (2025)
LAViTeR: Learning Aligned Visual and Textual Representations Assisted by Image and Caption Generation
by: Hashemi, Mohammad Abuzar, et al.
Published: (2021)
by: Hashemi, Mohammad Abuzar, et al.
Published: (2021)
FLEUR: An Explainable Reference-Free Evaluation Metric for Image Captioning Using a Large Multimodal Model
by: Lee, Yebin, et al.
Published: (2024)
by: Lee, Yebin, et al.
Published: (2024)
Predicting Winning Captions for Weekly New Yorker Comics
by: Cao, Stanley, et al.
Published: (2024)
by: Cao, Stanley, et al.
Published: (2024)
From Compound Figures to Composite Understanding: Developing a Multi-Modal LLM from Biomedical Literature with Medical Multiple-Image Benchmarking and Validation
by: Chen, Zhen, et al.
Published: (2025)
by: Chen, Zhen, et al.
Published: (2025)
Imagine How To Change: Explicit Procedure Modeling for Change Captioning
by: Sun, Jiayang, et al.
Published: (2026)
by: Sun, Jiayang, et al.
Published: (2026)
Enhancing Scientific Figure Captioning Through Cross-modal Learning
by: Rojas, Mateo Alejandro, et al.
Published: (2024)
by: Rojas, Mateo Alejandro, et al.
Published: (2024)
An Examination of the Robustness of Reference-Free Image Captioning Evaluation Metrics
by: Ahmadi, Saba, et al.
Published: (2023)
by: Ahmadi, Saba, et al.
Published: (2023)
Think Visually, Reason Textually: Vision-Language Synergy in ARC
by: Zhang, Beichen, et al.
Published: (2025)
by: Zhang, Beichen, et al.
Published: (2025)
VisioMath: Benchmarking Figure-based Mathematical Reasoning in LMMs
by: Li, Can, et al.
Published: (2025)
by: Li, Can, et al.
Published: (2025)
VFIG: Vectorizing Complex Figures in SVG with Vision-Language Models
by: He, Qijia, et al.
Published: (2026)
by: He, Qijia, et al.
Published: (2026)
DocumentCLIP: Linking Figures and Main Body Text in Reflowed Documents
by: Liu, Fuxiao, et al.
Published: (2023)
by: Liu, Fuxiao, et al.
Published: (2023)
CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning
by: Xing, Long, et al.
Published: (2025)
by: Xing, Long, et al.
Published: (2025)
Unveiling the Invisible: Captioning Videos with Metaphors
by: Kalarani, Abisek Rajakumar, et al.
Published: (2024)
by: Kalarani, Abisek Rajakumar, et al.
Published: (2024)
Text-only Synthesis for Image Captioning
by: Zhou, Qing, et al.
Published: (2024)
by: Zhou, Qing, et al.
Published: (2024)
The Role of Data Curation in Image Captioning
by: Li, Wenyan, et al.
Published: (2023)
by: Li, Wenyan, et al.
Published: (2023)
DeTikZify: Synthesizing Graphics Programs for Scientific Figures and Sketches with TikZ
by: Belouadi, Jonas, et al.
Published: (2024)
by: Belouadi, Jonas, et al.
Published: (2024)
G-VEval: A Versatile Metric for Evaluating Image and Video Captions Using GPT-4o
by: Tong, Tony Cheng, et al.
Published: (2024)
by: Tong, Tony Cheng, et al.
Published: (2024)
Leveraging Textual Compositional Reasoning for Robust Change Captioning
by: Park, Kyu Ri, et al.
Published: (2025)
by: Park, Kyu Ri, et al.
Published: (2025)
Decoding by Perturbation: Mitigating MLLM Hallucinations via Dynamic Textual Perturbation
by: Jia, Sihang, et al.
Published: (2026)
by: Jia, Sihang, et al.
Published: (2026)
Updating CLIP to Prefer Descriptions Over Captions
by: Zur, Amir, et al.
Published: (2024)
by: Zur, Amir, et al.
Published: (2024)
A bag of tricks for real-time Mitotic Figure detection
by: Marzahl, Christian, et al.
Published: (2025)
by: Marzahl, Christian, et al.
Published: (2025)
CAPEEN: Image Captioning with Early Exits and Knowledge Distillation
by: Bajpai, Divya Jyoti, et al.
Published: (2024)
by: Bajpai, Divya Jyoti, et al.
Published: (2024)
CIC: A Framework for Culturally-Aware Image Captioning
by: Yun, Youngsik, et al.
Published: (2024)
by: Yun, Youngsik, et al.
Published: (2024)
ChartCap: Mitigating Hallucination of Dense Chart Captioning
by: Lim, Junyoung, et al.
Published: (2025)
by: Lim, Junyoung, et al.
Published: (2025)
Video Summarization: Towards Entity-Aware Captions
by: Ayyubi, Hammad A., et al.
Published: (2023)
by: Ayyubi, Hammad A., et al.
Published: (2023)
FigCaps-HF: A Figure-to-Caption Generative Framework and Benchmark with Human Feedback
by: Singh, Ashish, et al.
Published: (2023)
by: Singh, Ashish, et al.
Published: (2023)
PatentLMM: Large Multimodal Model for Generating Descriptions for Patent Figures
by: Shukla, Shreya, et al.
Published: (2025)
by: Shukla, Shreya, et al.
Published: (2025)
CapGeo: A Caption-Assisted Approach to Geometric Reasoning
by: Li, Yuying, et al.
Published: (2025)
by: Li, Yuying, et al.
Published: (2025)
EXPERT: An Explainable Image Captioning Evaluation Metric with Structured Explanations
by: Kim, Hyunjong, et al.
Published: (2025)
by: Kim, Hyunjong, et al.
Published: (2025)
Similar Items
-
LaMP-Cap: Personalized Figure Caption Generation With Multimodal Figure Profiles
by: Ng, Ho Yin 'Sam', et al.
Published: (2025) -
FigEx2: Visual-Conditioned Panel Detection and Captioning for Scientific Compound Figures
by: Song, Jifeng, et al.
Published: (2026) -
Do Large Multimodal Models Solve Caption Generation for Scientific Figures? Lessons Learned from SciCap Challenge 2023
by: Hsu, Ting-Yao E., et al.
Published: (2025) -
Five Years of SciCap: What We Learned and Future Directions for Scientific Figure Captioning
by: Huang, Ting-Hao 'Kenneth', et al.
Published: (2025) -
The Solution for the ICCV 2023 1st Scientific Figure Captioning Challenge
by: Chao, Dian, et al.
Published: (2024)