Improving Image Captioning by Mimicking Human Reformulation Feedback at Inference-time
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Berger, Uri, Abend, Omri, Frermann, Lea, Stanovsky, Gabriel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Surveying the Landscape of Image Captioning Evaluation: A Comprehensive Taxonomy, Trends and Metrics Analysis
von: Berger, Uri, et al.
Veröffentlicht: (2024)
von: Berger, Uri, et al.
Veröffentlicht: (2024)
Polos: Multimodal Metric Learning from Human Feedback for Image Captioning
von: Wada, Yuiga, et al.
Veröffentlicht: (2024)
von: Wada, Yuiga, et al.
Veröffentlicht: (2024)
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing
von: Xing, Long, et al.
Veröffentlicht: (2025)
von: Xing, Long, et al.
Veröffentlicht: (2025)
#PraCegoVer: A Large Dataset for Image Captioning in Portuguese
von: Santos, Gabriel Oliveira dos, et al.
Veröffentlicht: (2021)
von: Santos, Gabriel Oliveira dos, et al.
Veröffentlicht: (2021)
Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions
von: Shen, Yijun, et al.
Veröffentlicht: (2025)
von: Shen, Yijun, et al.
Veröffentlicht: (2025)
Measuring Pragmatic Influence in Large Language Model Instructions
von: Geng, Yilin, et al.
Veröffentlicht: (2026)
von: Geng, Yilin, et al.
Veröffentlicht: (2026)
BACON: Improving Clarity of Image Captions via Bag-of-Concept Graphs
von: Yang, Zhantao, et al.
Veröffentlicht: (2024)
von: Yang, Zhantao, et al.
Veröffentlicht: (2024)
Image-Caption Encoding for Improving Zero-Shot Generalization
von: Yu, Eric Yang, et al.
Veröffentlicht: (2024)
von: Yu, Eric Yang, et al.
Veröffentlicht: (2024)
FigCaps-HF: A Figure-to-Caption Generative Framework and Benchmark with Human Feedback
von: Singh, Ashish, et al.
Veröffentlicht: (2023)
von: Singh, Ashish, et al.
Veröffentlicht: (2023)
From Image Captioning to Visual Storytelling
von: Passadakis, Admitos, et al.
Veröffentlicht: (2025)
von: Passadakis, Admitos, et al.
Veröffentlicht: (2025)
Distinctive Image Captioning: Leveraging Ground Truth Captions in CLIP Guided Reinforcement Learning
von: Chaffin, Antoine, et al.
Veröffentlicht: (2024)
von: Chaffin, Antoine, et al.
Veröffentlicht: (2024)
Dynamic Relation Inference via Verb Embeddings
von: Suissa, Omri, et al.
Veröffentlicht: (2025)
von: Suissa, Omri, et al.
Veröffentlicht: (2025)
Image Captioning via Compact Bidirectional Architecture
von: Song, Zijie, et al.
Veröffentlicht: (2022)
von: Song, Zijie, et al.
Veröffentlicht: (2022)
MUNIChus: Multilingual News Image Captioning Benchmark
von: Chen, Yuji, et al.
Veröffentlicht: (2026)
von: Chen, Yuji, et al.
Veröffentlicht: (2026)
TROPE: TRaining-Free Object-Part Enhancement for Seamlessly Improving Fine-Grained Zero-Shot Image Captioning
von: Feinglass, Joshua, et al.
Veröffentlicht: (2024)
von: Feinglass, Joshua, et al.
Veröffentlicht: (2024)
The Devil is in the EOS: Sequence Training for Detailed Image Captioning
von: Mohamed, Abdelrahman, et al.
Veröffentlicht: (2025)
von: Mohamed, Abdelrahman, et al.
Veröffentlicht: (2025)
Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning
von: Li, Wenyan, et al.
Veröffentlicht: (2024)
von: Li, Wenyan, et al.
Veröffentlicht: (2024)
Temporal Image Caption Retrieval Competition -- Description and Results
von: Pokrywka, Jakub, et al.
Veröffentlicht: (2024)
von: Pokrywka, Jakub, et al.
Veröffentlicht: (2024)
HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
Precision or Recall? An Analysis of Image Captions for Training Text-to-Image Generation Model
von: Cheng, Sheng, et al.
Veröffentlicht: (2024)
von: Cheng, Sheng, et al.
Veröffentlicht: (2024)
VIXEN: Visual Text Comparison Network for Image Difference Captioning
von: Black, Alexander, et al.
Veröffentlicht: (2024)
von: Black, Alexander, et al.
Veröffentlicht: (2024)
Transformer based Multitask Learning for Image Captioning and Object Detection
von: Basak, Debolena, et al.
Veröffentlicht: (2024)
von: Basak, Debolena, et al.
Veröffentlicht: (2024)
Altogether: Image Captioning via Re-aligning Alt-text
von: Xu, Hu, et al.
Veröffentlicht: (2024)
von: Xu, Hu, et al.
Veröffentlicht: (2024)
RxnCaption: Reformulating Reaction Diagram Parsing as Visual Prompt Guided Captioning
von: Song, Jiahe, et al.
Veröffentlicht: (2025)
von: Song, Jiahe, et al.
Veröffentlicht: (2025)
OmniCaptioner: One Captioner to Rule Them All
von: Lu, Yiting, et al.
Veröffentlicht: (2025)
von: Lu, Yiting, et al.
Veröffentlicht: (2025)
LLMs as Bridges: Reformulating Grounded Multimodal Named Entity Recognition
von: Li, Jinyuan, et al.
Veröffentlicht: (2024)
von: Li, Jinyuan, et al.
Veröffentlicht: (2024)
See or Guess: Counterfactually Regularized Image Captioning
von: Cao, Qian, et al.
Veröffentlicht: (2024)
von: Cao, Qian, et al.
Veröffentlicht: (2024)
Unifying Vision-Language Latents for Zero-label Image Caption Enhancement
von: Byun, Sanghyun, et al.
Veröffentlicht: (2025)
von: Byun, Sanghyun, et al.
Veröffentlicht: (2025)
CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era
von: Cheng, Kanzhi, et al.
Veröffentlicht: (2025)
von: Cheng, Kanzhi, et al.
Veröffentlicht: (2025)
Discovering Meaningful Units with Visually Grounded Semantics from Image Captions
von: Behjati, Melika, et al.
Veröffentlicht: (2025)
von: Behjati, Melika, et al.
Veröffentlicht: (2025)
SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation
von: Chen, Xiaofu, et al.
Veröffentlicht: (2025)
von: Chen, Xiaofu, et al.
Veröffentlicht: (2025)
Towards Adaptable and Interactive Image Captioning with Data Augmentation and Episodic Memory
von: Anagnostopoulou, Aliki, et al.
Veröffentlicht: (2023)
von: Anagnostopoulou, Aliki, et al.
Veröffentlicht: (2023)
Regional Attention-Enhanced Swin Transformer for Clinically Relevant Medical Image Captioning
von: Naz, Zubia, et al.
Veröffentlicht: (2025)
von: Naz, Zubia, et al.
Veröffentlicht: (2025)
Text-Only Training for Image Captioning with Retrieval Augmentation and Modality Gap Correction
von: Fonseca, Rui, et al.
Veröffentlicht: (2025)
von: Fonseca, Rui, et al.
Veröffentlicht: (2025)
EAMA : Entity-Aware Multimodal Alignment Based Approach for News Image Captioning
von: Zhang, Junzhe, et al.
Veröffentlicht: (2024)
von: Zhang, Junzhe, et al.
Veröffentlicht: (2024)
Diffusion-RSCC: Diffusion Probabilistic Model for Change Captioning in Remote Sensing Images
von: Yu, Xiaofei, et al.
Veröffentlicht: (2024)
von: Yu, Xiaofei, et al.
Veröffentlicht: (2024)
Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M)
von: Merchant, Nicholas, et al.
Veröffentlicht: (2025)
von: Merchant, Nicholas, et al.
Veröffentlicht: (2025)
Text-only Synthesis for Image Captioning
von: Zhou, Qing, et al.
Veröffentlicht: (2024)
von: Zhou, Qing, et al.
Veröffentlicht: (2024)
The Role of Data Curation in Image Captioning
von: Li, Wenyan, et al.
Veröffentlicht: (2023)
von: Li, Wenyan, et al.
Veröffentlicht: (2023)
CultureCLIP: Empowering CLIP with Cultural Awareness through Synthetic Images and Contextualized Captions
von: Huang, Yuchen, et al.
Veröffentlicht: (2025)
von: Huang, Yuchen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Surveying the Landscape of Image Captioning Evaluation: A Comprehensive Taxonomy, Trends and Metrics Analysis
von: Berger, Uri, et al.
Veröffentlicht: (2024) -
Polos: Multimodal Metric Learning from Human Feedback for Image Captioning
von: Wada, Yuiga, et al.
Veröffentlicht: (2024) -
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing
von: Xing, Long, et al.
Veröffentlicht: (2025) -
#PraCegoVer: A Large Dataset for Image Captioning in Portuguese
von: Santos, Gabriel Oliveira dos, et al.
Veröffentlicht: (2021) -
Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions
von: Shen, Yijun, et al.
Veröffentlicht: (2025)