Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Xinran, Diao, Muxi, Liu, Yuanzhi, Wang, Chunyu, Liang, Kongming, Ma, Zhanyu, Guo, Jun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Simple to Professional: A Combinatorial Controllable Image Captioning Agent
by: Wang, Xinran, et al.
Published: (2024)
by: Wang, Xinran, et al.
Published: (2024)
DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions
by: Wang, Xinran, et al.
Published: (2026)
by: Wang, Xinran, et al.
Published: (2026)
CineTechBench: A Benchmark for Cinematographic Technique Understanding and Generation
by: Wang, Xinran, et al.
Published: (2025)
by: Wang, Xinran, et al.
Published: (2025)
Detailed Object Description with Controllable Dimensions
by: Wang, Xinran, et al.
Published: (2024)
by: Wang, Xinran, et al.
Published: (2024)
Evaluating Attribute Comprehension in Large Vision-Language Models
by: Zhang, Haiwen, et al.
Published: (2024)
by: Zhang, Haiwen, et al.
Published: (2024)
DriveRX: A Vision-Language Reasoning Model for Cross-Task Autonomous Driving
by: Diao, Muxi, et al.
Published: (2025)
by: Diao, Muxi, et al.
Published: (2025)
Hepato-LLaVA: An Expert MLLM with Sparse Topo-Pack Attention for Hepatocellular Pathology Analysis on Whole Slide Images
by: Yang, Yuxuan, et al.
Published: (2026)
by: Yang, Yuxuan, et al.
Published: (2026)
Toward Generalizable Forgery Detection and Reasoning
by: Gao, Yueying, et al.
Published: (2025)
by: Gao, Yueying, et al.
Published: (2025)
Efficient Face Super-Resolution via Wavelet-based Feature Enhancement Network
by: Li, Wenjie, et al.
Published: (2024)
by: Li, Wenjie, et al.
Published: (2024)
Benchmarking Segmentation Models with Mask-Preserved Attribute Editing
by: Yin, Zijin, et al.
Published: (2024)
by: Yin, Zijin, et al.
Published: (2024)
RO-Bench: Large-scale robustness evaluation of MLLMs with text-driven counterfactual videos
by: Yang, Zixi, et al.
Published: (2025)
by: Yang, Zixi, et al.
Published: (2025)
Curriculum Group Policy Optimization: Adaptive Sampling for Unleashing the Potential of Text-to-Image Generation
by: Li, Baoteng, et al.
Published: (2026)
by: Li, Baoteng, et al.
Published: (2026)
MedReasoner: Reinforcement Learning Drives Reasoning Grounding from Clinical Thought to Pixel-Level Precision
by: Yan, Zhonghao, et al.
Published: (2025)
by: Yan, Zhonghao, et al.
Published: (2025)
FairHuman: Boosting Hand and Face Quality in Human Image Generation with Minimum Potential Delay Fairness in Diffusion Models
by: Wang, Yuxuan, et al.
Published: (2025)
by: Wang, Yuxuan, et al.
Published: (2025)
Benchmarking and Improving Detail Image Caption
by: Dong, Hongyuan, et al.
Published: (2024)
by: Dong, Hongyuan, et al.
Published: (2024)
Generative Visual Chain-of-Thought for Image Editing
by: Yin, Zijin, et al.
Published: (2026)
by: Yin, Zijin, et al.
Published: (2026)
PGP-SAM: Prototype-Guided Prompt Learning for Efficient Few-Shot Medical Image Segmentation
by: Yan, Zhonghao, et al.
Published: (2025)
by: Yan, Zhonghao, et al.
Published: (2025)
Benchmarking Semantic Segmentation Models via Appearance and Geometry Attribute Editing
by: Yin, Zijin, et al.
Published: (2026)
by: Yin, Zijin, et al.
Published: (2026)
SpatialLock: Precise Spatial Control in Text-to-Image Synthesis
by: Liu, Biao, et al.
Published: (2025)
by: Liu, Biao, et al.
Published: (2025)
ConMo: Controllable Motion Disentanglement and Recomposition for Zero-Shot Motion Transfer
by: Gao, Jiayi, et al.
Published: (2025)
by: Gao, Jiayi, et al.
Published: (2025)
Improving Text Generation on Images with Synthetic Captions
by: Koh, Jun Young, et al.
Published: (2024)
by: Koh, Jun Young, et al.
Published: (2024)
Evaluating Image Caption via Cycle-consistent Text-to-Image Generation
by: Cui, Tianyu, et al.
Published: (2025)
by: Cui, Tianyu, et al.
Published: (2025)
OmniEraser: Remove Objects and Their Effects in Images with Paired Video-Frame Data
by: Wei, Runpu, et al.
Published: (2025)
by: Wei, Runpu, et al.
Published: (2025)
Text Data-Centric Image Captioning with Interactive Prompts
by: Wang, Yiyu, et al.
Published: (2024)
by: Wang, Yiyu, et al.
Published: (2024)
Generating Accurate and Detailed Captions for High-Resolution Images
by: Lee, Hankyeol, et al.
Published: (2025)
by: Lee, Hankyeol, et al.
Published: (2025)
FourierSR: A Fourier Token-based Plugin for Efficient Image Super-Resolution
by: Li, Wenjie, et al.
Published: (2025)
by: Li, Wenjie, et al.
Published: (2025)
Describe Anything: Detailed Localized Image and Video Captioning
by: Lian, Long, et al.
Published: (2025)
by: Lian, Long, et al.
Published: (2025)
The Devil is in the EOS: Sequence Training for Detailed Image Captioning
by: Mohamed, Abdelrahman, et al.
Published: (2025)
by: Mohamed, Abdelrahman, et al.
Published: (2025)
Panoptic Captioning: An Equivalence Bridge for Image and Text
by: Lin, Kun-Yu, et al.
Published: (2025)
by: Lin, Kun-Yu, et al.
Published: (2025)
CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era
by: Cheng, Kanzhi, et al.
Published: (2025)
by: Cheng, Kanzhi, et al.
Published: (2025)
Visual Fact Checker: Enabling High-Fidelity Detailed Caption Generation
by: Ge, Yunhao, et al.
Published: (2024)
by: Ge, Yunhao, et al.
Published: (2024)
Painting with Words: Elevating Detailed Image Captioning with Benchmark and Alignment Learning
by: Ye, Qinghao, et al.
Published: (2025)
by: Ye, Qinghao, et al.
Published: (2025)
Polyp-E: Benchmarking the Robustness of Deep Segmentation Models via Polyp Editing
by: Wei, Runpu, et al.
Published: (2024)
by: Wei, Runpu, et al.
Published: (2024)
UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New Benchmarks
by: Wu, Peiran, et al.
Published: (2025)
by: Wu, Peiran, et al.
Published: (2025)
Evaluating Text-to-Image Generative Models: An Empirical Study on Human Image Synthesis
by: Chen, Muxi, et al.
Published: (2024)
by: Chen, Muxi, et al.
Published: (2024)
Geometric Image Editing via Effects-Sensitive In-Context Inpainting with Diffusion Transformers
by: Zhang, Shuo, et al.
Published: (2026)
by: Zhang, Shuo, et al.
Published: (2026)
VolumeDiffusion: Flexible Text-to-3D Generation with Efficient Volumetric Encoder
by: Tang, Zhicong, et al.
Published: (2023)
by: Tang, Zhicong, et al.
Published: (2023)
AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark
by: Chai, Wenhao, et al.
Published: (2024)
by: Chai, Wenhao, et al.
Published: (2024)
ReflectCAP: Detailed Image Captioning with Reflective Memory
by: Min, Kyungmin, et al.
Published: (2026)
by: Min, Kyungmin, et al.
Published: (2026)
Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed Perception
by: Ma, Ziyang, et al.
Published: (2025)
by: Ma, Ziyang, et al.
Published: (2025)
Similar Items
-
From Simple to Professional: A Combinatorial Controllable Image Captioning Agent
by: Wang, Xinran, et al.
Published: (2024) -
DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions
by: Wang, Xinran, et al.
Published: (2026) -
CineTechBench: A Benchmark for Cinematographic Technique Understanding and Generation
by: Wang, Xinran, et al.
Published: (2025) -
Detailed Object Description with Controllable Dimensions
by: Wang, Xinran, et al.
Published: (2024) -
Evaluating Attribute Comprehension in Large Vision-Language Models
by: Zhang, Haiwen, et al.
Published: (2024)