SBS Figures: Pre-training Figure QA from Stage-by-Stage Synthesized Images
Fuente:
arXiv
Saved in:
| Main Authors: | Shinoda, Risa, Saito, Kuniaki, Tanaka, Shohei, Hirasawa, Tosho, Ushiku, Yoshitaka |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HalDec-Bench: Benchmarking Hallucination Detector in Image Captioning
by: Saito, Kuniaki, et al.
Published: (2026)
by: Saito, Kuniaki, et al.
Published: (2026)
HalDec-Bench: Benchmarking Hallucination Detector in Image Captioning
by: Saito, Kuniaki, et al.
Published: (2025)
by: Saito, Kuniaki, et al.
Published: (2025)
SciPostGen: Bridging the Gap between Scientific Papers and Poster Layouts
by: Inadumi, Shun, et al.
Published: (2025)
by: Inadumi, Shun, et al.
Published: (2025)
SciPostLayout: A Dataset for Layout Analysis and Layout Generation of Scientific Posters
by: Tanaka, Shohei, et al.
Published: (2024)
by: Tanaka, Shohei, et al.
Published: (2024)
SciPostLayoutTree: A Dataset for Structural Analysis of Scientific Posters
by: Tanaka, Shohei, et al.
Published: (2025)
by: Tanaka, Shohei, et al.
Published: (2025)
CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning
by: Saito, Kuniaki, et al.
Published: (2025)
by: Saito, Kuniaki, et al.
Published: (2025)
AgroBench: Vision-Language Model Benchmark in Agriculture
by: Shinoda, Risa, et al.
Published: (2025)
by: Shinoda, Risa, et al.
Published: (2025)
COM Kitchens: An Unedited Overhead-view Video Dataset as a Vision-Language Benchmark
by: Maeda, Koki, et al.
Published: (2024)
by: Maeda, Koki, et al.
Published: (2024)
BioVITA: Biological Dataset, Model, and Benchmark for Visual-Textual-Acoustic Alignment
by: Shinoda, Risa, et al.
Published: (2026)
by: Shinoda, Risa, et al.
Published: (2026)
Evaluating the Capability of Video Question Generation for Expert Knowledge Elicitation
by: Zhang, Huaying, et al.
Published: (2025)
by: Zhang, Huaying, et al.
Published: (2025)
PetFace: A Large-Scale Dataset and Benchmark for Animal Identification
by: Shinoda, Risa, et al.
Published: (2024)
by: Shinoda, Risa, et al.
Published: (2024)
OpenAnimalTracks: A Dataset for Animal Track Recognition
by: Shinoda, Risa, et al.
Published: (2024)
by: Shinoda, Risa, et al.
Published: (2024)
CAMOT: Camera Angle-aware Multi-Object Tracking
by: Limanta, Felix, et al.
Published: (2024)
by: Limanta, Felix, et al.
Published: (2024)
WarrantScore: Modeling Warrants between Claims and Evidence for Substantiation Evaluation in Peer Reviews
by: Mori, Kiyotada, et al.
Published: (2026)
by: Mori, Kiyotada, et al.
Published: (2026)
UniQA: Unified Vision-Language Pre-training for Image Quality and Aesthetic Assessment
by: Zhou, Hantao, et al.
Published: (2024)
by: Zhou, Hantao, et al.
Published: (2024)
GaussianPlant: Structure-aligned Gaussian Splatting for 3D Reconstruction of Plants
by: Yang, Yang, et al.
Published: (2025)
by: Yang, Yang, et al.
Published: (2025)
A Two-Stage Progressive Pre-training using Multi-Modal Contrastive Masked Autoencoders
by: Jamal, Muhammad Abdullah, et al.
Published: (2024)
by: Jamal, Muhammad Abdullah, et al.
Published: (2024)
Learning Contrastive Multimodal Fusion with Improved Modality Dropout for Disease Detection and Prediction
by: Gu, Yi, et al.
Published: (2025)
by: Gu, Yi, et al.
Published: (2025)
Towards Safer Mobile Agents: Scalable Generation and Evaluation of Diverse Scenarios for VLMs
by: Taniguchi, Takara, et al.
Published: (2026)
by: Taniguchi, Takara, et al.
Published: (2026)
Answering Questions in Stages: Prompt Chaining for Contract QA
by: Roegiest, Adam, et al.
Published: (2024)
by: Roegiest, Adam, et al.
Published: (2024)
AnimalClue: Recognizing Animals by their Traces
by: Shinoda, Risa, et al.
Published: (2025)
by: Shinoda, Risa, et al.
Published: (2025)
Exploring Disentangled and Controllable Human Image Synthesis: From End-to-End to Stage-by-Stage
by: Sun, Zhengwentai, et al.
Published: (2025)
by: Sun, Zhengwentai, et al.
Published: (2025)
A Two-Stage Globally-Diverse Adversarial Attack for Vision-Language Pre-training Models
by: Chen, Wutao, et al.
Published: (2026)
by: Chen, Wutao, et al.
Published: (2026)
Am I More Pointwise or Pairwise? Revealing Position Bias in Rubric-Based LLM-as-a-Judge
by: Xu, Yuzheng, et al.
Published: (2026)
by: Xu, Yuzheng, et al.
Published: (2026)
Weak-to-Strong Compositional Learning from Generative Models for Language-based Object Detection
by: Park, Kwanyong, et al.
Published: (2024)
by: Park, Kwanyong, et al.
Published: (2024)
TNF: Tri-branch Neural Fusion for Multimodal Medical Data Classification
by: Zheng, Tong, et al.
Published: (2024)
by: Zheng, Tong, et al.
Published: (2024)
Challenges and Lessons from MIDOG 2025: A Two-Stage Approach to Domain-Robust Mitotic Figure Detection
by: Song, Euiseop, et al.
Published: (2025)
by: Song, Euiseop, et al.
Published: (2025)
DeTikZify: Synthesizing Graphics Programs for Scientific Figures and Sketches with TikZ
by: Belouadi, Jonas, et al.
Published: (2024)
by: Belouadi, Jonas, et al.
Published: (2024)
VIFSS: View-Invariant and Figure Skating-Specific Pose Representation Learning for Temporal Action Segmentation
by: Tanaka, Ryota, et al.
Published: (2025)
by: Tanaka, Ryota, et al.
Published: (2025)
Recipe Generation from Unsegmented Cooking Videos
by: Nishimura, Taichi, et al.
Published: (2022)
by: Nishimura, Taichi, et al.
Published: (2022)
Should VLMs be Pre-trained with Image Data?
by: Keh, Sedrick, et al.
Published: (2025)
by: Keh, Sedrick, et al.
Published: (2025)
Dual-Stage Global and Local Feature Framework for Image Dehazing
by: Ali, Anas M., et al.
Published: (2025)
by: Ali, Anas M., et al.
Published: (2025)
Setting the Stage: Text-Driven Scene-Consistent Image Generation
by: Xie, Cong, et al.
Published: (2025)
by: Xie, Cong, et al.
Published: (2025)
User-in-the-Loop View Sampling with Error Peaking Visualization
by: Yasunaga, Ayaka, et al.
Published: (2025)
by: Yasunaga, Ayaka, et al.
Published: (2025)
Interaction-via-Actions: Cattle Interaction Detection with Joint Learning of Action-Interaction Latent Space
by: Nakagawa, Ren, et al.
Published: (2025)
by: Nakagawa, Ren, et al.
Published: (2025)
One-Stage-TFS: Thai One-Stage Fingerspelling Dataset for Fingerspelling Recognition Frameworks
by: Lata, Siriwiwat, et al.
Published: (2024)
by: Lata, Siriwiwat, et al.
Published: (2024)
Scalable Pre-training of Large Autoregressive Image Models
by: El-Nouby, Alaaeldin, et al.
Published: (2024)
by: El-Nouby, Alaaeldin, et al.
Published: (2024)
Beyond Isolated Facts: Synthesizing Narrative and Grounded Supervision for VideoQA
by: Liang, Jianxin, et al.
Published: (2025)
by: Liang, Jianxin, et al.
Published: (2025)
A Mountain-Shaped Single-Stage Network for Accurate Image Restoration
by: Gao, Hu, et al.
Published: (2023)
by: Gao, Hu, et al.
Published: (2023)
HCF: Hierarchical Cascade Framework for Distributed Multi-Stage Image Compression
by: Cai, Junhao, et al.
Published: (2025)
by: Cai, Junhao, et al.
Published: (2025)
Similar Items
-
HalDec-Bench: Benchmarking Hallucination Detector in Image Captioning
by: Saito, Kuniaki, et al.
Published: (2026) -
HalDec-Bench: Benchmarking Hallucination Detector in Image Captioning
by: Saito, Kuniaki, et al.
Published: (2025) -
SciPostGen: Bridging the Gap between Scientific Papers and Poster Layouts
by: Inadumi, Shun, et al.
Published: (2025) -
SciPostLayout: A Dataset for Layout Analysis and Layout Generation of Scientific Posters
by: Tanaka, Shohei, et al.
Published: (2024) -
SciPostLayoutTree: A Dataset for Structural Analysis of Scientific Posters
by: Tanaka, Shohei, et al.
Published: (2025)