Evaluating Image Caption via Cycle-consistent Text-to-Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Cui, Tianyu, Bai, Jinbin, Wang, Guo-Hua, Chen, Qing-Guo, Xu, Zhao, Luo, Weihua, Zhang, Kaifu, Shi, Ye |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CHATS: Combining Human-Aligned Optimization and Test-Time Sampling for Text-to-Image Generation
by: Fu, Minghao, et al.
Published: (2025)
by: Fu, Minghao, et al.
Published: (2025)
Diffusion-SDPO: Safeguarded Direct Preference Optimization for Diffusion Models
by: Fu, Minghao, et al.
Published: (2025)
by: Fu, Minghao, et al.
Published: (2025)
TeEFusion: Blending Text Embeddings to Distill Classifier-Free Guidance
by: Fu, Minghao, et al.
Published: (2025)
by: Fu, Minghao, et al.
Published: (2025)
UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image Generation
by: Duan, Lunhao, et al.
Published: (2024)
by: Duan, Lunhao, et al.
Published: (2024)
Deep But Reliable: Advancing Multi-turn Reasoning for Thinking with Images
by: Yang, Wenhao, et al.
Published: (2025)
by: Yang, Wenhao, et al.
Published: (2025)
Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation
by: Wang, Xinran, et al.
Published: (2025)
by: Wang, Xinran, et al.
Published: (2025)
Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
by: Zhao, Shanshan, et al.
Published: (2025)
by: Zhao, Shanshan, et al.
Published: (2025)
Text Data-Centric Image Captioning with Interactive Prompts
by: Wang, Yiyu, et al.
Published: (2024)
by: Wang, Yiyu, et al.
Published: (2024)
Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis
by: Bai, Jinbin, et al.
Published: (2024)
by: Bai, Jinbin, et al.
Published: (2024)
A Unified Agentic Framework for Evaluating Conditional Image Generation
by: Wang, Jifang, et al.
Published: (2025)
by: Wang, Jifang, et al.
Published: (2025)
Improving Text Generation on Images with Synthetic Captions
by: Koh, Jun Young, et al.
Published: (2024)
by: Koh, Jun Young, et al.
Published: (2024)
Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization
by: Yang, Wenhao, et al.
Published: (2026)
by: Yang, Wenhao, et al.
Published: (2026)
Omni-View: Unlocking How Generation Facilitates Understanding in Unified 3D Model based on Multiview images
by: Hu, JiaKui, et al.
Published: (2025)
by: Hu, JiaKui, et al.
Published: (2025)
Dynamic Prompt Optimizing for Text-to-Image Generation
by: Mo, Wenyi, et al.
Published: (2024)
by: Mo, Wenyi, et al.
Published: (2024)
FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any Granularity
by: Hua, Hang, et al.
Published: (2024)
by: Hua, Hang, et al.
Published: (2024)
Training-Free Image Editing with Visual Context Integration and Concept Alignment
by: Song, Rui, et al.
Published: (2026)
by: Song, Rui, et al.
Published: (2026)
Target-aware Image Editing via Cycle-consistent Constraints
by: Wang, Yanghao, et al.
Published: (2025)
by: Wang, Yanghao, et al.
Published: (2025)
Personalized Safety Alignment for Text-to-Image Diffusion Models
by: Lei, Yu, et al.
Published: (2025)
by: Lei, Yu, et al.
Published: (2025)
EmoAttack: Emotion-to-Image Diffusion Models for Emotional Backdoor Generation
by: Wei, Tianyu, et al.
Published: (2024)
by: Wei, Tianyu, et al.
Published: (2024)
Conditional Panoramic Image Generation via Masked Autoregressive Modeling
by: Wang, Chaoyang, et al.
Published: (2025)
by: Wang, Chaoyang, et al.
Published: (2025)
ITIScore: An Image-to-Text-to-Image Rating Framework for the Image Captioning Ability of MLLMs
by: Xu, Zitong, et al.
Published: (2026)
by: Xu, Zitong, et al.
Published: (2026)
Text-only Synthesis for Image Captioning
by: Zhou, Qing, et al.
Published: (2024)
by: Zhou, Qing, et al.
Published: (2024)
Integrating View Conditions for Image Synthesis
by: Bai, Jinbin, et al.
Published: (2023)
by: Bai, Jinbin, et al.
Published: (2023)
Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model
by: Shi, Qingyu, et al.
Published: (2025)
by: Shi, Qingyu, et al.
Published: (2025)
ImageDoctor: Diagnosing Text-to-Image Generation via Grounded Image Reasoning
by: Guo, Yuxiang, et al.
Published: (2025)
by: Guo, Yuxiang, et al.
Published: (2025)
Semi-supervised Chinese Poem-to-Painting Generation via Cycle-consistent Adversarial Networks
by: Lu, Zhengyang, et al.
Published: (2024)
by: Lu, Zhengyang, et al.
Published: (2024)
Generating an Image From 1,000 Words: Enhancing Text-to-Image With Structured Captions
by: Gutflaish, Eyal, et al.
Published: (2025)
by: Gutflaish, Eyal, et al.
Published: (2025)
MDP3: A Training-free Approach for List-wise Frame Selection in Video-LLMs
by: Sun, Hui, et al.
Published: (2025)
by: Sun, Hui, et al.
Published: (2025)
Ovis: Structural Embedding Alignment for Multimodal Large Language Model
by: Lu, Shiyin, et al.
Published: (2024)
by: Lu, Shiyin, et al.
Published: (2024)
Cycle Diffusion Model for Counterfactual Image Generation
by: Huang, Fangrui, et al.
Published: (2025)
by: Huang, Fangrui, et al.
Published: (2025)
Benchmarking and Improving Detail Image Caption
by: Dong, Hongyuan, et al.
Published: (2024)
by: Dong, Hongyuan, et al.
Published: (2024)
Panoptic Captioning: An Equivalence Bridge for Image and Text
by: Lin, Kun-Yu, et al.
Published: (2025)
by: Lin, Kun-Yu, et al.
Published: (2025)
RAPHAEL: Text-to-Image Generation via Large Mixture of Diffusion Paths
by: Xue, Zeyue, et al.
Published: (2023)
by: Xue, Zeyue, et al.
Published: (2023)
Enhancing Reward Models for High-quality Image Generation: Beyond Text-Image Alignment
by: Ba, Ying, et al.
Published: (2025)
by: Ba, Ying, et al.
Published: (2025)
Continual Learning for Image Captioning through Improved Image-Text Alignment
by: Taetz, Bertram, et al.
Published: (2025)
by: Taetz, Bertram, et al.
Published: (2025)
Image Captions are Natural Prompts for Text-to-Image Models
by: Lei, Shiye, et al.
Published: (2023)
by: Lei, Shiye, et al.
Published: (2023)
Precision or Recall? An Analysis of Image Captions for Training Text-to-Image Generation Model
by: Cheng, Sheng, et al.
Published: (2024)
by: Cheng, Sheng, et al.
Published: (2024)
Ovis-Image Technical Report
by: Wang, Guo-Hua, et al.
Published: (2025)
by: Wang, Guo-Hua, et al.
Published: (2025)
StyleInject: Parameter Efficient Tuning of Text-to-Image Diffusion Models
by: Zhou, Mohan, et al.
Published: (2024)
by: Zhou, Mohan, et al.
Published: (2024)
I4VGen: Image as Free Stepping Stone for Text-to-Video Generation
by: Guo, Xiefan, et al.
Published: (2024)
by: Guo, Xiefan, et al.
Published: (2024)
Similar Items
-
CHATS: Combining Human-Aligned Optimization and Test-Time Sampling for Text-to-Image Generation
by: Fu, Minghao, et al.
Published: (2025) -
Diffusion-SDPO: Safeguarded Direct Preference Optimization for Diffusion Models
by: Fu, Minghao, et al.
Published: (2025) -
TeEFusion: Blending Text Embeddings to Distill Classifier-Free Guidance
by: Fu, Minghao, et al.
Published: (2025) -
UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image Generation
by: Duan, Lunhao, et al.
Published: (2024) -
Deep But Reliable: Advancing Multi-turn Reasoning for Thinking with Images
by: Yang, Wenhao, et al.
Published: (2025)