Painting with Words: Elevating Detailed Image Captioning with Benchmark and Alignment Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Ye, Qinghao, Zeng, Xianhan, Li, Fu, Li, Chunyuan, Fan, Haoqi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Benchmarking and Improving Detail Image Caption
por: Dong, Hongyuan, et al.
Publicado: (2024)
por: Dong, Hongyuan, et al.
Publicado: (2024)
LLaVA-Critic: Learning to Evaluate Multimodal Models
por: Xiong, Tianyi, et al.
Publicado: (2024)
por: Xiong, Tianyi, et al.
Publicado: (2024)
CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era
por: Cheng, Kanzhi, et al.
Publicado: (2025)
por: Cheng, Kanzhi, et al.
Publicado: (2025)
Classification Done Right for Vision-Language Pre-Training
por: Huang, Zilong, et al.
Publicado: (2024)
por: Huang, Zilong, et al.
Publicado: (2024)
UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New Benchmarks
por: Wu, Peiran, et al.
Publicado: (2025)
por: Wu, Peiran, et al.
Publicado: (2025)
SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning
por: Zhang, Lin, et al.
Publicado: (2025)
por: Zhang, Lin, et al.
Publicado: (2025)
DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions
por: Wang, Xinran, et al.
Publicado: (2026)
por: Wang, Xinran, et al.
Publicado: (2026)
Caption Generation for Dongba Paintings via Prompt Learning and Semantic Fusion
por: Qian, Shuangwu, et al.
Publicado: (2026)
por: Qian, Shuangwu, et al.
Publicado: (2026)
Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation
por: Wang, Xinran, et al.
Publicado: (2025)
por: Wang, Xinran, et al.
Publicado: (2025)
Describe Anything: Detailed Localized Image and Video Captioning
por: Lian, Long, et al.
Publicado: (2025)
por: Lian, Long, et al.
Publicado: (2025)
The Devil is in the EOS: Sequence Training for Detailed Image Captioning
por: Mohamed, Abdelrahman, et al.
Publicado: (2025)
por: Mohamed, Abdelrahman, et al.
Publicado: (2025)
AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark
por: Chai, Wenhao, et al.
Publicado: (2024)
por: Chai, Wenhao, et al.
Publicado: (2024)
Removing Distributional Discrepancies in Captions Improves Image-Text Alignment
por: Li, Yuheng, et al.
Publicado: (2024)
por: Li, Yuheng, et al.
Publicado: (2024)
Continual Learning for Image Captioning through Improved Image-Text Alignment
por: Taetz, Bertram, et al.
Publicado: (2025)
por: Taetz, Bertram, et al.
Publicado: (2025)
ImageInWords: Unlocking Hyper-Detailed Image Descriptions
por: Garg, Roopal, et al.
Publicado: (2024)
por: Garg, Roopal, et al.
Publicado: (2024)
Visual Fact Checker: Enabling High-Fidelity Detailed Caption Generation
por: Ge, Yunhao, et al.
Publicado: (2024)
por: Ge, Yunhao, et al.
Publicado: (2024)
Generating Accurate and Detailed Captions for High-Resolution Images
por: Lee, Hankyeol, et al.
Publicado: (2025)
por: Lee, Hankyeol, et al.
Publicado: (2025)
ReflectCAP: Detailed Image Captioning with Reflective Memory
por: Min, Kyungmin, et al.
Publicado: (2026)
por: Min, Kyungmin, et al.
Publicado: (2026)
No Detail Left Behind: Revisiting Self-Retrieval for Fine-Grained Image Captioning
por: Gaur, Manu, et al.
Publicado: (2024)
por: Gaur, Manu, et al.
Publicado: (2024)
Cockatiel: Ensembling Synthetic and Human Preferenced Training for Detailed Video Caption
por: Qin, Luozheng, et al.
Publicado: (2025)
por: Qin, Luozheng, et al.
Publicado: (2025)
PrefPaint: Aligning Image Inpainting Diffusion Model with Human Preference
por: Liu, Kendong, et al.
Publicado: (2024)
por: Liu, Kendong, et al.
Publicado: (2024)
Generating an Image From 1,000 Words: Enhancing Text-to-Image With Structured Captions
por: Gutflaish, Eyal, et al.
Publicado: (2025)
por: Gutflaish, Eyal, et al.
Publicado: (2025)
A Video is Worth 10,000 Words: Training and Benchmarking with Diverse Captions for Better Long Video Retrieval
por: Gwilliam, Matthew, et al.
Publicado: (2023)
por: Gwilliam, Matthew, et al.
Publicado: (2023)
DetailSemNet: Elevating Signature Verification through Detail-Semantic Integration
por: Shih, Meng-Cheng, et al.
Publicado: (2025)
por: Shih, Meng-Cheng, et al.
Publicado: (2025)
Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed Perception
por: Ma, Ziyang, et al.
Publicado: (2025)
por: Ma, Ziyang, et al.
Publicado: (2025)
When Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Chain-of-Thought
por: Zhou, Yiyang, et al.
Publicado: (2025)
por: Zhou, Yiyang, et al.
Publicado: (2025)
ViDA-UGC: Detailed Image Quality Analysis via Visual Distortion Assessment for UGC Images
por: Liao, Wenjie, et al.
Publicado: (2025)
por: Liao, Wenjie, et al.
Publicado: (2025)
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World
por: Li, Xiangtai, et al.
Publicado: (2025)
por: Li, Xiangtai, et al.
Publicado: (2025)
CaptionQA: Is Your Caption as Useful as the Image Itself?
por: Yang, Shijia, et al.
Publicado: (2025)
por: Yang, Shijia, et al.
Publicado: (2025)
MUNIChus: Multilingual News Image Captioning Benchmark
por: Chen, Yuji, et al.
Publicado: (2026)
por: Chen, Yuji, et al.
Publicado: (2026)
Seeing the Image: Prioritizing Visual Correlation by Contrastive Alignment
por: Xiao, Xin, et al.
Publicado: (2024)
por: Xiao, Xin, et al.
Publicado: (2024)
good4cir: Generating Detailed Synthetic Captions for Composed Image Retrieval
por: Kolouju, Pranavi, et al.
Publicado: (2025)
por: Kolouju, Pranavi, et al.
Publicado: (2025)
A Benchmark for Multi-Lingual Vision-Language Learning in Remote Sensing Image Captioning
por: Zhou, Qing, et al.
Publicado: (2025)
por: Zhou, Qing, et al.
Publicado: (2025)
Learning Inclusion Matching for Animation Paint Bucket Colorization
por: Dai, Yuekun, et al.
Publicado: (2024)
por: Dai, Yuekun, et al.
Publicado: (2024)
PaintFlow: A Unified Framework for Interactive Oil Paintings Editing and Generation
por: Hu, Zhangli, et al.
Publicado: (2025)
por: Hu, Zhangli, et al.
Publicado: (2025)
Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage
por: Lee, Saehyung, et al.
Publicado: (2024)
por: Lee, Saehyung, et al.
Publicado: (2024)
SynPO: Synergizing Descriptiveness and Preference Optimization for Video Detailed Captioning
por: Dang, Jisheng, et al.
Publicado: (2025)
por: Dang, Jisheng, et al.
Publicado: (2025)
HalDec-Bench: Benchmarking Hallucination Detector in Image Captioning
por: Saito, Kuniaki, et al.
Publicado: (2025)
por: Saito, Kuniaki, et al.
Publicado: (2025)
OwlCap: Harmonizing Motion-Detail for Video Captioning via HMD-270K and Caption Set Equivalence Reward
por: Zhong, Chunlin, et al.
Publicado: (2025)
por: Zhong, Chunlin, et al.
Publicado: (2025)
HalDec-Bench: Benchmarking Hallucination Detector in Image Captioning
por: Saito, Kuniaki, et al.
Publicado: (2026)
por: Saito, Kuniaki, et al.
Publicado: (2026)
Ejemplares similares
-
Benchmarking and Improving Detail Image Caption
por: Dong, Hongyuan, et al.
Publicado: (2024) -
LLaVA-Critic: Learning to Evaluate Multimodal Models
por: Xiong, Tianyi, et al.
Publicado: (2024) -
CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era
por: Cheng, Kanzhi, et al.
Publicado: (2025) -
Classification Done Right for Vision-Language Pre-Training
por: Huang, Zilong, et al.
Publicado: (2024) -
UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New Benchmarks
por: Wu, Peiran, et al.
Publicado: (2025)