Text Data-Centric Image Captioning with Interactive Prompts
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Yiyu, Luo, Hao, Xu, Jungang, Sun, Yingfei, Wang, Fan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation
di: Wang, Xinran, et al.
Pubblicazione: (2025)
di: Wang, Xinran, et al.
Pubblicazione: (2025)
Image Captions are Natural Prompts for Text-to-Image Models
di: Lei, Shiye, et al.
Pubblicazione: (2023)
di: Lei, Shiye, et al.
Pubblicazione: (2023)
Evaluating Image Caption via Cycle-consistent Text-to-Image Generation
di: Cui, Tianyu, et al.
Pubblicazione: (2025)
di: Cui, Tianyu, et al.
Pubblicazione: (2025)
EPIC: Efficient Prompt Interaction for Text-Image Classification
di: Yu, Xinyao, et al.
Pubblicazione: (2025)
di: Yu, Xinyao, et al.
Pubblicazione: (2025)
Player-Centric Multimodal Prompt Generation for Large Language Model Based Identity-Aware Basketball Video Captioning
di: Xi, Zeyu, et al.
Pubblicazione: (2025)
di: Xi, Zeyu, et al.
Pubblicazione: (2025)
RxnCaption: Reformulating Reaction Diagram Parsing as Visual Prompt Guided Captioning
di: Song, Jiahe, et al.
Pubblicazione: (2025)
di: Song, Jiahe, et al.
Pubblicazione: (2025)
Panoptic Captioning: An Equivalence Bridge for Image and Text
di: Lin, Kun-Yu, et al.
Pubblicazione: (2025)
di: Lin, Kun-Yu, et al.
Pubblicazione: (2025)
RACap: Relation-Aware Prompting for Lightweight Retrieval-Augmented Image Captioning
di: Long, Xiaosheng, et al.
Pubblicazione: (2025)
di: Long, Xiaosheng, et al.
Pubblicazione: (2025)
Memory-Inspired Temporal Prompt Interaction for Text-Image Classification
di: Yu, Xinyao, et al.
Pubblicazione: (2024)
di: Yu, Xinyao, et al.
Pubblicazione: (2024)
Uncovering the Text Embedding in Text-to-Image Diffusion Models
di: Yu, Hu, et al.
Pubblicazione: (2024)
di: Yu, Hu, et al.
Pubblicazione: (2024)
VIXEN: Visual Text Comparison Network for Image Difference Captioning
di: Black, Alexander, et al.
Pubblicazione: (2024)
di: Black, Alexander, et al.
Pubblicazione: (2024)
DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding
di: Wu, Hao, et al.
Pubblicazione: (2024)
di: Wu, Hao, et al.
Pubblicazione: (2024)
CaptionQA: Is Your Caption as Useful as the Image Itself?
di: Yang, Shijia, et al.
Pubblicazione: (2025)
di: Yang, Shijia, et al.
Pubblicazione: (2025)
ITIScore: An Image-to-Text-to-Image Rating Framework for the Image Captioning Ability of MLLMs
di: Xu, Zitong, et al.
Pubblicazione: (2026)
di: Xu, Zitong, et al.
Pubblicazione: (2026)
Text-only Synthesis for Image Captioning
di: Zhou, Qing, et al.
Pubblicazione: (2024)
di: Zhou, Qing, et al.
Pubblicazione: (2024)
Improving Text Generation on Images with Synthetic Captions
di: Koh, Jun Young, et al.
Pubblicazione: (2024)
di: Koh, Jun Young, et al.
Pubblicazione: (2024)
D-ORCA: Dialogue-Centric Optimization for Robust Audio-Visual Captioning
di: Tang, Changli, et al.
Pubblicazione: (2026)
di: Tang, Changli, et al.
Pubblicazione: (2026)
HICEScore: A Hierarchical Metric for Image Captioning Evaluation
di: Zeng, Zequn, et al.
Pubblicazione: (2024)
di: Zeng, Zequn, et al.
Pubblicazione: (2024)
Language Prompt for Autonomous Driving
di: Wu, Dongming, et al.
Pubblicazione: (2023)
di: Wu, Dongming, et al.
Pubblicazione: (2023)
RORPCap: Retrieval-based Objects and Relations Prompt for Image Captioning
di: Gu, Jinjing, et al.
Pubblicazione: (2025)
di: Gu, Jinjing, et al.
Pubblicazione: (2025)
TextCenGen: Attention-Guided Text-Centric Background Adaptation for Text-to-Image Generation
di: Liang, Tianyi, et al.
Pubblicazione: (2024)
di: Liang, Tianyi, et al.
Pubblicazione: (2024)
VISTAR:A User-Centric and Role-Driven Benchmark for Text-to-Image Evaluation
di: Jiang, Kaiyuan, et al.
Pubblicazione: (2025)
di: Jiang, Kaiyuan, et al.
Pubblicazione: (2025)
Towards Adaptable and Interactive Image Captioning with Data Augmentation and Episodic Memory
di: Anagnostopoulou, Aliki, et al.
Pubblicazione: (2023)
di: Anagnostopoulou, Aliki, et al.
Pubblicazione: (2023)
Prompt-Softbox-Prompt: A Free-Text Embedding Control for Image Editing
di: Yang, Yitong, et al.
Pubblicazione: (2024)
di: Yang, Yitong, et al.
Pubblicazione: (2024)
Effectively Enhancing Vision Language Large Models by Prompt Augmentation and Caption Utilization
di: Zhao, Minyi, et al.
Pubblicazione: (2024)
di: Zhao, Minyi, et al.
Pubblicazione: (2024)
VisualPrompter: Semantic-Aware Prompt Optimization with Visual Feedback for Text-to-Image Synthesis
di: Wu, Shiyu, et al.
Pubblicazione: (2025)
di: Wu, Shiyu, et al.
Pubblicazione: (2025)
Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting
di: Tang, Yunlong, et al.
Pubblicazione: (2025)
di: Tang, Yunlong, et al.
Pubblicazione: (2025)
TIT-Score: Evaluating Long-Prompt Based Text-to-Image Alignment via Text-to-Image-to-Text Consistency
di: Wang, Juntong, et al.
Pubblicazione: (2025)
di: Wang, Juntong, et al.
Pubblicazione: (2025)
Continual Learning for Image Captioning through Improved Image-Text Alignment
di: Taetz, Bertram, et al.
Pubblicazione: (2025)
di: Taetz, Bertram, et al.
Pubblicazione: (2025)
Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image Fusion
di: Yi, Xunpeng, et al.
Pubblicazione: (2024)
di: Yi, Xunpeng, et al.
Pubblicazione: (2024)
Optimizing Prompts for Text-to-Image Generation
di: Hao, Yaru, et al.
Pubblicazione: (2022)
di: Hao, Yaru, et al.
Pubblicazione: (2022)
TIPO: Text to Image with Text Presampling for Prompt Optimization
di: Yeh, Shih-Ying, et al.
Pubblicazione: (2024)
di: Yeh, Shih-Ying, et al.
Pubblicazione: (2024)
Is Your Text-to-Image Model Robust to Caption Noise?
di: Yu, Weichen, et al.
Pubblicazione: (2024)
di: Yu, Weichen, et al.
Pubblicazione: (2024)
Removing Distributional Discrepancies in Captions Improves Image-Text Alignment
di: Li, Yuheng, et al.
Pubblicazione: (2024)
di: Li, Yuheng, et al.
Pubblicazione: (2024)
Restore Text First, Enhance Image Later: Two-Stage Scene Text Image Super-Resolution with Glyph Structure Guidance
di: Luo, Minxing, et al.
Pubblicazione: (2025)
di: Luo, Minxing, et al.
Pubblicazione: (2025)
Prompt-Based Caption Generation for Single-Tooth Dental Images Using Vision-Language Models
di: Sukhanova, Anastasiia, et al.
Pubblicazione: (2026)
di: Sukhanova, Anastasiia, et al.
Pubblicazione: (2026)
Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M)
di: Merchant, Nicholas, et al.
Pubblicazione: (2025)
di: Merchant, Nicholas, et al.
Pubblicazione: (2025)
Shifting AI Efficiency From Model-Centric to Data-Centric Compression
di: Liu, Xuyang, et al.
Pubblicazione: (2025)
di: Liu, Xuyang, et al.
Pubblicazione: (2025)
Negative Entity Suppression for Zero-Shot Captioning with Synthetic Images
di: Lu, Zimao, et al.
Pubblicazione: (2025)
di: Lu, Zimao, et al.
Pubblicazione: (2025)
SPT: Sequence Prompt Transformer for Interactive Image Segmentation
di: Cheng, Senlin, et al.
Pubblicazione: (2024)
di: Cheng, Senlin, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation
di: Wang, Xinran, et al.
Pubblicazione: (2025) -
Image Captions are Natural Prompts for Text-to-Image Models
di: Lei, Shiye, et al.
Pubblicazione: (2023) -
Evaluating Image Caption via Cycle-consistent Text-to-Image Generation
di: Cui, Tianyu, et al.
Pubblicazione: (2025) -
EPIC: Efficient Prompt Interaction for Text-Image Classification
di: Yu, Xinyao, et al.
Pubblicazione: (2025) -
Player-Centric Multimodal Prompt Generation for Large Language Model Based Identity-Aware Basketball Video Captioning
di: Xi, Zeyu, et al.
Pubblicazione: (2025)