Precision or Recall? An Analysis of Image Captions for Training Text-to-Image Generation Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cheng, Sheng, Patel, Maitreya, Yang, Yezhou |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
$λ$-ECLIPSE: Multi-Concept Personalized Text-to-Image Diffusion Models by Leveraging CLIP Latent Space
von: Patel, Maitreya, et al.
Veröffentlicht: (2024)
von: Patel, Maitreya, et al.
Veröffentlicht: (2024)
ConceptBed: Evaluating Concept Learning Abilities of Text-to-Image Diffusion Models
von: Patel, Maitreya, et al.
Veröffentlicht: (2023)
von: Patel, Maitreya, et al.
Veröffentlicht: (2023)
WOUAF: Weight Modulation for User Attribution and Fingerprinting in Text-to-Image Diffusion Models
von: Kim, Changhoon, et al.
Veröffentlicht: (2023)
von: Kim, Changhoon, et al.
Veröffentlicht: (2023)
TextInVision: Text and Prompt Complexity Driven Visual Text Generation Benchmark
von: Fallah, Forouzan, et al.
Veröffentlicht: (2025)
von: Fallah, Forouzan, et al.
Veröffentlicht: (2025)
AcT2I: Evaluating and Improving Action Depiction in Text-to-Image Models
von: Malaviya, Vatsal, et al.
Veröffentlicht: (2025)
von: Malaviya, Vatsal, et al.
Veröffentlicht: (2025)
TROPE: TRaining-Free Object-Part Enhancement for Seamlessly Improving Fine-Grained Zero-Shot Image Captioning
von: Feinglass, Joshua, et al.
Veröffentlicht: (2024)
von: Feinglass, Joshua, et al.
Veröffentlicht: (2024)
TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives
von: Patel, Maitreya, et al.
Veröffentlicht: (2024)
von: Patel, Maitreya, et al.
Veröffentlicht: (2024)
Steering Rectified Flow Models in the Vector Field for Controlled Image Generation
von: Patel, Maitreya, et al.
Veröffentlicht: (2024)
von: Patel, Maitreya, et al.
Veröffentlicht: (2024)
RefEdit: A Benchmark and Method for Improving Instruction-based Image Editing Model on Referring Expressions
von: Pathiraja, Bimsara, et al.
Veröffentlicht: (2025)
von: Pathiraja, Bimsara, et al.
Veröffentlicht: (2025)
VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations
von: Patel, Maitreya, et al.
Veröffentlicht: (2026)
von: Patel, Maitreya, et al.
Veröffentlicht: (2026)
Text-Only Training for Image Captioning with Retrieval Augmentation and Modality Gap Correction
von: Fonseca, Rui, et al.
Veröffentlicht: (2025)
von: Fonseca, Rui, et al.
Veröffentlicht: (2025)
The Devil is in the EOS: Sequence Training for Detailed Image Captioning
von: Mohamed, Abdelrahman, et al.
Veröffentlicht: (2025)
von: Mohamed, Abdelrahman, et al.
Veröffentlicht: (2025)
Help Me Identify: Is an LLM+VQA System All We Need to Identify Visual Concepts?
von: Sampat, Shailaja Keyur, et al.
Veröffentlicht: (2024)
von: Sampat, Shailaja Keyur, et al.
Veröffentlicht: (2024)
VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical Reasoning
von: Yilmaz, Nilay, et al.
Veröffentlicht: (2025)
von: Yilmaz, Nilay, et al.
Veröffentlicht: (2025)
VIXEN: Visual Text Comparison Network for Image Difference Captioning
von: Black, Alexander, et al.
Veröffentlicht: (2024)
von: Black, Alexander, et al.
Veröffentlicht: (2024)
Text-only Synthesis for Image Captioning
von: Zhou, Qing, et al.
Veröffentlicht: (2024)
von: Zhou, Qing, et al.
Veröffentlicht: (2024)
Image2Text2Image: A Novel Framework for Label-Free Evaluation of Image-to-Text Generation with Text-to-Image Diffusion Models
von: Huang, Jia-Hong, et al.
Veröffentlicht: (2024)
von: Huang, Jia-Hong, et al.
Veröffentlicht: (2024)
From Image Captioning to Visual Storytelling
von: Passadakis, Admitos, et al.
Veröffentlicht: (2025)
von: Passadakis, Admitos, et al.
Veröffentlicht: (2025)
Image-Caption Encoding for Improving Zero-Shot Generalization
von: Yu, Eric Yang, et al.
Veröffentlicht: (2024)
von: Yu, Eric Yang, et al.
Veröffentlicht: (2024)
Unleashing Text-to-Image Diffusion Prior for Zero-Shot Image Captioning
von: Luo, Jianjie, et al.
Veröffentlicht: (2024)
von: Luo, Jianjie, et al.
Veröffentlicht: (2024)
Image Captioning via Compact Bidirectional Architecture
von: Song, Zijie, et al.
Veröffentlicht: (2022)
von: Song, Zijie, et al.
Veröffentlicht: (2022)
MUNIChus: Multilingual News Image Captioning Benchmark
von: Chen, Yuji, et al.
Veröffentlicht: (2026)
von: Chen, Yuji, et al.
Veröffentlicht: (2026)
Distinctive Image Captioning: Leveraging Ground Truth Captions in CLIP Guided Reinforcement Learning
von: Chaffin, Antoine, et al.
Veröffentlicht: (2024)
von: Chaffin, Antoine, et al.
Veröffentlicht: (2024)
Fast Prompt Alignment for Text-to-Image Generation
von: Mrini, Khalil, et al.
Veröffentlicht: (2024)
von: Mrini, Khalil, et al.
Veröffentlicht: (2024)
CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era
von: Cheng, Kanzhi, et al.
Veröffentlicht: (2025)
von: Cheng, Kanzhi, et al.
Veröffentlicht: (2025)
Harnessing Synthetic Preference Data for Enhancing Temporal Understanding of Video-LLMs
von: Vani, Sameep, et al.
Veröffentlicht: (2025)
von: Vani, Sameep, et al.
Veröffentlicht: (2025)
Surveying the Landscape of Image Captioning Evaluation: A Comprehensive Taxonomy, Trends and Metrics Analysis
von: Berger, Uri, et al.
Veröffentlicht: (2024)
von: Berger, Uri, et al.
Veröffentlicht: (2024)
R.A.C.E.: Robust Adversarial Concept Erasure for Secure Text-to-Image Diffusion Model
von: Kim, Changhoon, et al.
Veröffentlicht: (2024)
von: Kim, Changhoon, et al.
Veröffentlicht: (2024)
Linguistically Informed Multimodal Fusion for Vietnamese Scene-Text Image Captioning: Dataset, Graph Framework, and Phonological Attention
von: Nguyen, Nhi Ngoc-Yen, et al.
Veröffentlicht: (2026)
von: Nguyen, Nhi Ngoc-Yen, et al.
Veröffentlicht: (2026)
Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning
von: Li, Wenyan, et al.
Veröffentlicht: (2024)
von: Li, Wenyan, et al.
Veröffentlicht: (2024)
Temporal Image Caption Retrieval Competition -- Description and Results
von: Pokrywka, Jakub, et al.
Veröffentlicht: (2024)
von: Pokrywka, Jakub, et al.
Veröffentlicht: (2024)
Diffusion-RSCC: Diffusion Probabilistic Model for Change Captioning in Remote Sensing Images
von: Yu, Xiaofei, et al.
Veröffentlicht: (2024)
von: Yu, Xiaofei, et al.
Veröffentlicht: (2024)
CEIDM: A Controlled Entity and Interaction Diffusion Model for Enhanced Text-to-Image Generation
von: Yang, Mingyue, et al.
Veröffentlicht: (2025)
von: Yang, Mingyue, et al.
Veröffentlicht: (2025)
Can Prompt Modifiers Control Bias? A Comparative Analysis of Text-to-Image Generative Models
von: Shin, Philip Wootaek, et al.
Veröffentlicht: (2024)
von: Shin, Philip Wootaek, et al.
Veröffentlicht: (2024)
ANNA: Abstractive Text-to-Image Synthesis with Filtered News Captions
von: Ramakrishnan, Aashish Anantha, et al.
Veröffentlicht: (2023)
von: Ramakrishnan, Aashish Anantha, et al.
Veröffentlicht: (2023)
Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M)
von: Merchant, Nicholas, et al.
Veröffentlicht: (2025)
von: Merchant, Nicholas, et al.
Veröffentlicht: (2025)
BACON: Improving Clarity of Image Captions via Bag-of-Concept Graphs
von: Yang, Zhantao, et al.
Veröffentlicht: (2024)
von: Yang, Zhantao, et al.
Veröffentlicht: (2024)
Optimizing Prompts for Text-to-Image Generation
von: Hao, Yaru, et al.
Veröffentlicht: (2022)
von: Hao, Yaru, et al.
Veröffentlicht: (2022)
Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation
von: Yang, Yue, et al.
Veröffentlicht: (2025)
von: Yang, Yue, et al.
Veröffentlicht: (2025)
Transformer based Multitask Learning for Image Captioning and Object Detection
von: Basak, Debolena, et al.
Veröffentlicht: (2024)
von: Basak, Debolena, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
$λ$-ECLIPSE: Multi-Concept Personalized Text-to-Image Diffusion Models by Leveraging CLIP Latent Space
von: Patel, Maitreya, et al.
Veröffentlicht: (2024) -
ConceptBed: Evaluating Concept Learning Abilities of Text-to-Image Diffusion Models
von: Patel, Maitreya, et al.
Veröffentlicht: (2023) -
WOUAF: Weight Modulation for User Attribution and Fingerprinting in Text-to-Image Diffusion Models
von: Kim, Changhoon, et al.
Veröffentlicht: (2023) -
TextInVision: Text and Prompt Complexity Driven Visual Text Generation Benchmark
von: Fallah, Forouzan, et al.
Veröffentlicht: (2025) -
AcT2I: Evaluating and Improving Action Depiction in Text-to-Image Models
von: Malaviya, Vatsal, et al.
Veröffentlicht: (2025)