FigEx2: Visual-Conditioned Panel Detection and Captioning for Scientific Compound Figures

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Song, Jifeng, Das, Arun, Wang, Pan, Ji, Hui, Zhao, Kun, Huang, Yufei
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917367924129792
author Song, Jifeng
Das, Arun
Wang, Pan
Ji, Hui
Zhao, Kun
Huang, Yufei
author_facet Song, Jifeng
Das, Arun
Wang, Pan
Ji, Hui
Zhao, Kun
Huang, Yufei
contents Scientific compound figures combine multiple labeled panels into a single image. However, in a PMC-scale crawl of 346,567 compound figures, 16.3% have no caption and 1.8% only have captions shorter than ten words, causing them to be discarded by existing caption-decomposition pipelines. We propose FigEx2, a visual-conditioned framework that localizes panels and generates panel-wise captions directly from the image, converting otherwise unusable figures into aligned panel-text pairs for downstream pretraining and retrieval. To mitigate linguistic variance in open-ended captioning, we introduce a noise-aware gated fusion module that adaptively controls how caption features condition the detection query space, and employ a staged SFT+RL strategy with CLIP-based alignment and BERTScore-based semantic rewards. To support high-quality supervision, we curate BioSci-Fig-Cap, a refined benchmark for panel-level grounding, alongside cross-disciplinary test suites in physics and chemistry. FigEx2 achieves 0.728 mAP@0.5:0.95 for detection, outperforms Qwen3-VL-8B by 0.44 in METEOR and 0.22 in BERTScore, and transfers zero-shot to out-of-distribution scientific domains without fine-tuning.
format Preprint
id arxiv_https___arxiv_org_abs_2601_08026
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle FigEx2: Visual-Conditioned Panel Detection and Captioning for Scientific Compound Figures
Song, Jifeng
Das, Arun
Wang, Pan
Ji, Hui
Zhao, Kun
Huang, Yufei
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Scientific compound figures combine multiple labeled panels into a single image. However, in a PMC-scale crawl of 346,567 compound figures, 16.3% have no caption and 1.8% only have captions shorter than ten words, causing them to be discarded by existing caption-decomposition pipelines. We propose FigEx2, a visual-conditioned framework that localizes panels and generates panel-wise captions directly from the image, converting otherwise unusable figures into aligned panel-text pairs for downstream pretraining and retrieval. To mitigate linguistic variance in open-ended captioning, we introduce a noise-aware gated fusion module that adaptively controls how caption features condition the detection query space, and employ a staged SFT+RL strategy with CLIP-based alignment and BERTScore-based semantic rewards. To support high-quality supervision, we curate BioSci-Fig-Cap, a refined benchmark for panel-level grounding, alongside cross-disciplinary test suites in physics and chemistry. FigEx2 achieves 0.728 mAP@0.5:0.95 for detection, outperforms Qwen3-VL-8B by 0.44 in METEOR and 0.22 in BERTScore, and transfers zero-shot to out-of-distribution scientific domains without fine-tuning.
title FigEx2: Visual-Conditioned Panel Detection and Captioning for Scientific Compound Figures
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2601.08026