Generalizable Geometric Image Caption Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | Xin, Yue, Wang, Wenyuan, Pan, Rui, Wang, Ruida, Meng, Howard, Pi, Renjie, Diao, Shizhe, Zhang, Tong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Instinctive Bias: Spurious Images lead to Illusion in MLLMs
by: Han, Tianyang, et al.
Published: (2024)
by: Han, Tianyang, et al.
Published: (2024)
GAR: Generative Adversarial Reinforcement Learning for Formal Theorem Proving
by: Wang, Ruida, et al.
Published: (2025)
by: Wang, Ruida, et al.
Published: (2025)
From Simple to Professional: A Combinatorial Controllable Image Captioning Agent
by: Wang, Xinran, et al.
Published: (2024)
by: Wang, Xinran, et al.
Published: (2024)
Aligning Forest and Trees in Images & Long Captions for Visually Grounded Understanding
by: Woo, Byeongju, et al.
Published: (2026)
by: Woo, Byeongju, et al.
Published: (2026)
Image Captions are Natural Prompts for Text-to-Image Models
by: Lei, Shiye, et al.
Published: (2023)
by: Lei, Shiye, et al.
Published: (2023)
Text-only Synthesis for Image Captioning
by: Zhou, Qing, et al.
Published: (2024)
by: Zhou, Qing, et al.
Published: (2024)
AlignGuard: Scalable Safety Alignment for Text-to-Image Generation
by: Liu, Runtao, et al.
Published: (2024)
by: Liu, Runtao, et al.
Published: (2024)
TheoremLlama: Transforming General-Purpose LLMs into Lean4 Experts
by: Wang, Ruida, et al.
Published: (2024)
by: Wang, Ruida, et al.
Published: (2024)
Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models
by: Lai, Zhengfeng, et al.
Published: (2024)
by: Lai, Zhengfeng, et al.
Published: (2024)
Gondola: Grounded Vision Language Planning for Generalizable Robotic Manipulation
by: Chen, Shizhe, et al.
Published: (2025)
by: Chen, Shizhe, et al.
Published: (2025)
Semi-Supervised Image Captioning Considering Wasserstein Graph Matching
by: Yang, Yang
Published: (2024)
by: Yang, Yang
Published: (2024)
VideoDPO: Omni-Preference Alignment for Video Diffusion Generation
by: Liu, Runtao, et al.
Published: (2024)
by: Liu, Runtao, et al.
Published: (2024)
Hyperdimensional Cross-Modal Alignment of Frozen Language and Image Models for Efficient Image Captioning
by: Dalvi, Abhishek, et al.
Published: (2026)
by: Dalvi, Abhishek, et al.
Published: (2026)
Top-Down Semantic Refinement for Image Captioning
by: Zhang, Jusheng, et al.
Published: (2025)
by: Zhang, Jusheng, et al.
Published: (2025)
Captioning Daily Activity Images in Early Childhood Education: Benchmark and Algorithm
by: Li, Sixing, et al.
Published: (2026)
by: Li, Sixing, et al.
Published: (2026)
Cross Modification Attention Based Deliberation Model for Image Captioning
by: Lian, Zheng, et al.
Published: (2021)
by: Lian, Zheng, et al.
Published: (2021)
LongVideoAgent: Multi-Agent Reasoning with Long Videos
by: Liu, Runtao, et al.
Published: (2025)
by: Liu, Runtao, et al.
Published: (2025)
RubiCap: Rubric-Guided Reinforcement Learning for Dense Image Captioning
by: Huang, Tzu-Heng, et al.
Published: (2026)
by: Huang, Tzu-Heng, et al.
Published: (2026)
Domain Generalizable Continual Learning
by: Yan, Hongwei, et al.
Published: (2025)
by: Yan, Hongwei, et al.
Published: (2025)
CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
by: Fiastre, Gabriel, et al.
Published: (2025)
by: Fiastre, Gabriel, et al.
Published: (2025)
Graph-Based Captioning: Enhancing Visual Descriptions by Interconnecting Region Captions
by: Hsieh, Yu-Guan, et al.
Published: (2024)
by: Hsieh, Yu-Guan, et al.
Published: (2024)
Captured by Captions: On Memorization and its Mitigation in CLIP Models
by: Wang, Wenhao, et al.
Published: (2025)
by: Wang, Wenhao, et al.
Published: (2025)
ComboStoc: Combinatorial Stochasticity for Diffusion Generative Models
by: Xu, Rui, et al.
Published: (2024)
by: Xu, Rui, et al.
Published: (2024)
CaptionFool: Universal Image Captioning Model Attacks
by: Parekh, Swapnil
Published: (2026)
by: Parekh, Swapnil
Published: (2026)
SynC: Synthetic Image Caption Dataset Refinement with One-to-many Mapping for Zero-shot Image Captioning
by: Kim, Si-Woo, et al.
Published: (2025)
by: Kim, Si-Woo, et al.
Published: (2025)
Curvature-Aware Captioning:Leveraging Geodesic Attention for 3D Scene Understanding
by: He, Ziyao, et al.
Published: (2026)
by: He, Ziyao, et al.
Published: (2026)
AGIC: Attention-Guided Image Captioning to Improve Caption Relevance
by: Teja, L. D. M. S. Sai, et al.
Published: (2025)
by: Teja, L. D. M. S. Sai, et al.
Published: (2025)
Generalizable Hyperparameter Optimization for Federated Learning on Non-IID Cancer Images
by: Ribeiro, Elisa Gonçalves, et al.
Published: (2026)
by: Ribeiro, Elisa Gonçalves, et al.
Published: (2026)
Adapt-Pruner: Adaptive Structural Pruning for Efficient Small Language Model Training
by: Pan, Rui, et al.
Published: (2025)
by: Pan, Rui, et al.
Published: (2025)
AutoGeo: Automating Geometric Image Dataset Creation for Enhanced Geometry Understanding
by: Huang, Zihan, et al.
Published: (2024)
by: Huang, Zihan, et al.
Published: (2024)
How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions
by: Brack, Manuel, et al.
Published: (2025)
by: Brack, Manuel, et al.
Published: (2025)
High-Resolution Image Synthesis via Next-Token Prediction
by: Chen, Dengsheng, et al.
Published: (2024)
by: Chen, Dengsheng, et al.
Published: (2024)
Reinforcement Learning with Generalizable Gaussian Splatting
by: Wang, Jiaxu, et al.
Published: (2024)
by: Wang, Jiaxu, et al.
Published: (2024)
VeCLIP: Improving CLIP Training via Visual-enriched Captions
by: Lai, Zhengfeng, et al.
Published: (2023)
by: Lai, Zhengfeng, et al.
Published: (2023)
MTA: Multimodal Task Alignment for BEV Perception and Captioning
by: Ma, Yunsheng, et al.
Published: (2024)
by: Ma, Yunsheng, et al.
Published: (2024)
Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M)
by: Merchant, Nicholas, et al.
Published: (2025)
by: Merchant, Nicholas, et al.
Published: (2025)
QG-CoC: Question-Guided Chain-of-Captions for Large Multimodal Models
by: Kao, Kuei-Chun, et al.
Published: (2025)
by: Kao, Kuei-Chun, et al.
Published: (2025)
Pix2Cap-COCO: Advancing Visual Comprehension via Pixel-Level Captioning
by: You, Zuyao, et al.
Published: (2025)
by: You, Zuyao, et al.
Published: (2025)
Is Your Text-to-Image Model Robust to Caption Noise?
by: Yu, Weichen, et al.
Published: (2024)
by: Yu, Weichen, et al.
Published: (2024)
Deep Multimodal Learning with Missing Modality: A Survey
by: Wu, Renjie, et al.
Published: (2024)
by: Wu, Renjie, et al.
Published: (2024)
Similar Items
-
The Instinctive Bias: Spurious Images lead to Illusion in MLLMs
by: Han, Tianyang, et al.
Published: (2024) -
GAR: Generative Adversarial Reinforcement Learning for Formal Theorem Proving
by: Wang, Ruida, et al.
Published: (2025) -
From Simple to Professional: A Combinatorial Controllable Image Captioning Agent
by: Wang, Xinran, et al.
Published: (2024) -
Aligning Forest and Trees in Images & Long Captions for Visually Grounded Understanding
by: Woo, Byeongju, et al.
Published: (2026) -
Image Captions are Natural Prompts for Text-to-Image Models
by: Lei, Shiye, et al.
Published: (2023)