Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M)
Fuente:
arXiv
Saved in:
| Main Authors: | Merchant, Nicholas, Borde, Haitz Sáez de Ocáriz, Popescu, Andrei Cristian, Suarez, Carlos Garcia Jurado |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Keep It Light! Simplifying Image Clustering Via Text-Free Adapters
by: Li, Yicen, et al.
Published: (2025)
by: Li, Yicen, et al.
Published: (2025)
Fine-Tuning Next-Scale Visual Autoregressive Models with Group Relative Policy Optimization
by: Gallici, Matteo, et al.
Published: (2025)
by: Gallici, Matteo, et al.
Published: (2025)
Neural Snowflakes: Universal Latent Graph Inference via Trainable Latent Geometries
by: Borde, Haitz Sáez de Ocáriz, et al.
Published: (2023)
by: Borde, Haitz Sáez de Ocáriz, et al.
Published: (2023)
Approximation Rates and VC-Dimension Bounds for (P)ReLU MLP Mixture of Experts
by: Kratsios, Anastasis, et al.
Published: (2024)
by: Kratsios, Anastasis, et al.
Published: (2024)
LoRA Fine-Tuning Without GPUs: A CPU-Efficient Meta-Generation Framework for LLMs
by: Arabpour, Reza, et al.
Published: (2025)
by: Arabpour, Reza, et al.
Published: (2025)
Beyond Parallelism: Synergistic Computational Graph Effects in Multi-Head Attention
by: Borde, Haitz Sáez de Ocáriz
Published: (2025)
by: Borde, Haitz Sáez de Ocáriz
Published: (2025)
Learning 3D Hypersonic Flow with Physics-Enhanced Neural Fields: A Case Study on the Orion Reentry Capsule
by: Borde, Haitz Sáez de Ocáriz, et al.
Published: (2026)
by: Borde, Haitz Sáez de Ocáriz, et al.
Published: (2026)
Mathematical Foundations of Geometric Deep Learning
by: Borde, Haitz Sáez de Ocáriz, et al.
Published: (2025)
by: Borde, Haitz Sáez de Ocáriz, et al.
Published: (2025)
Sharp Generalization Bounds for Foundation Models with Asymmetric Randomized Low-Rank Adapters
by: Kratsios, Anastasis, et al.
Published: (2025)
by: Kratsios, Anastasis, et al.
Published: (2025)
Neural Spacetimes for DAG Representation Learning
by: Borde, Haitz Sáez de Ocáriz, et al.
Published: (2024)
by: Borde, Haitz Sáez de Ocáriz, et al.
Published: (2024)
Text Data-Centric Image Captioning with Interactive Prompts
by: Wang, Yiyu, et al.
Published: (2024)
by: Wang, Yiyu, et al.
Published: (2024)
Improving Text Generation on Images with Synthetic Captions
by: Koh, Jun Young, et al.
Published: (2024)
by: Koh, Jun Young, et al.
Published: (2024)
Soro: A Lightweight Foundation Model and Chatbot for Tajik
by: Liashkov, Stanislav, et al.
Published: (2026)
by: Liashkov, Stanislav, et al.
Published: (2026)
Image Captions are Natural Prompts for Text-to-Image Models
by: Lei, Shiye, et al.
Published: (2023)
by: Lei, Shiye, et al.
Published: (2023)
The Stepwise Informativeness Assumption: Why are Entropy Dynamics and Reasoning Correlated in LLMs?
by: Català, Mar Gonzàlez I, et al.
Published: (2026)
by: Català, Mar Gonzàlez I, et al.
Published: (2026)
Every Feedforward Neural Network Definable in an o-Minimal Structure Has Finite Sample Complexity
by: Kratsios, Anastasis, et al.
Published: (2026)
by: Kratsios, Anastasis, et al.
Published: (2026)
Text-only Synthesis for Image Captioning
by: Zhou, Qing, et al.
Published: (2024)
by: Zhou, Qing, et al.
Published: (2024)
Closed-Form Diffusion Models
by: Scarvelis, Christopher, et al.
Published: (2023)
by: Scarvelis, Christopher, et al.
Published: (2023)
AGIC: Attention-Guided Image Captioning to Improve Caption Relevance
by: Teja, L. D. M. S. Sai, et al.
Published: (2025)
by: Teja, L. D. M. S. Sai, et al.
Published: (2025)
SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning
by: Zhang, Lin, et al.
Published: (2025)
by: Zhang, Lin, et al.
Published: (2025)
Removing Distributional Discrepancies in Captions Improves Image-Text Alignment
by: Li, Yuheng, et al.
Published: (2024)
by: Li, Yuheng, et al.
Published: (2024)
Continual Learning for Image Captioning through Improved Image-Text Alignment
by: Taetz, Bertram, et al.
Published: (2025)
by: Taetz, Bertram, et al.
Published: (2025)
ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image Captioning
by: Kim, Taewhan, et al.
Published: (2024)
by: Kim, Taewhan, et al.
Published: (2024)
Altogether: Image Captioning via Re-aligning Alt-text
by: Xu, Hu, et al.
Published: (2024)
by: Xu, Hu, et al.
Published: (2024)
Benchmarking and Improving Detail Image Caption
by: Dong, Hongyuan, et al.
Published: (2024)
by: Dong, Hongyuan, et al.
Published: (2024)
k-Maximum Inner Product Attention for Graph Transformers and the Expressive Power of GraphGPS
by: De Schouwer, Jonas, et al.
Published: (2026)
by: De Schouwer, Jonas, et al.
Published: (2026)
Panoptic Captioning: An Equivalence Bridge for Image and Text
by: Lin, Kun-Yu, et al.
Published: (2025)
by: Lin, Kun-Yu, et al.
Published: (2025)
Score Distillation via Reparametrized DDIM
by: Lukoianov, Artem, et al.
Published: (2024)
by: Lukoianov, Artem, et al.
Published: (2024)
VIXEN: Visual Text Comparison Network for Image Difference Captioning
by: Black, Alexander, et al.
Published: (2024)
by: Black, Alexander, et al.
Published: (2024)
CaptionQA: Is Your Caption as Useful as the Image Itself?
by: Yang, Shijia, et al.
Published: (2025)
by: Yang, Shijia, et al.
Published: (2025)
Image-Caption Encoding for Improving Zero-Shot Generalization
by: Yu, Eric Yang, et al.
Published: (2024)
by: Yu, Eric Yang, et al.
Published: (2024)
Improving Text-To-Audio Models with Synthetic Captions
by: Kong, Zhifeng, et al.
Published: (2024)
by: Kong, Zhifeng, et al.
Published: (2024)
RxnCaption: Reformulating Reaction Diagram Parsing as Visual Prompt Guided Captioning
by: Song, Jiahe, et al.
Published: (2025)
by: Song, Jiahe, et al.
Published: (2025)
Generating an Image From 1,000 Words: Enhancing Text-to-Image With Structured Captions
by: Gutflaish, Eyal, et al.
Published: (2025)
by: Gutflaish, Eyal, et al.
Published: (2025)
GeReA: Question-Aware Prompt Captions for Knowledge-based Visual Question Answering
by: Ma, Ziyu, et al.
Published: (2024)
by: Ma, Ziyu, et al.
Published: (2024)
Unleashing Text-to-Image Diffusion Prior for Zero-Shot Image Captioning
by: Luo, Jianjie, et al.
Published: (2024)
by: Luo, Jianjie, et al.
Published: (2024)
ANNA: Abstractive Text-to-Image Synthesis with Filtered News Captions
by: Ramakrishnan, Aashish Anantha, et al.
Published: (2023)
by: Ramakrishnan, Aashish Anantha, et al.
Published: (2023)
CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning
by: Saito, Kuniaki, et al.
Published: (2025)
by: Saito, Kuniaki, et al.
Published: (2025)
CaptionFool: Universal Image Captioning Model Attacks
by: Parekh, Swapnil
Published: (2026)
by: Parekh, Swapnil
Published: (2026)
LAION-SG: An Enhanced Large-Scale Dataset for Training Complex Image-Text Models with Structural Annotations
by: Li, Zejian, et al.
Published: (2024)
by: Li, Zejian, et al.
Published: (2024)
Similar Items
-
Keep It Light! Simplifying Image Clustering Via Text-Free Adapters
by: Li, Yicen, et al.
Published: (2025) -
Fine-Tuning Next-Scale Visual Autoregressive Models with Group Relative Policy Optimization
by: Gallici, Matteo, et al.
Published: (2025) -
Neural Snowflakes: Universal Latent Graph Inference via Trainable Latent Geometries
by: Borde, Haitz Sáez de Ocáriz, et al.
Published: (2023) -
Approximation Rates and VC-Dimension Bounds for (P)ReLU MLP Mixture of Experts
by: Kratsios, Anastasis, et al.
Published: (2024) -
LoRA Fine-Tuning Without GPUs: A CPU-Efficient Meta-Generation Framework for LLMs
by: Arabpour, Reza, et al.
Published: (2025)