The Devil is in the EOS: Sequence Training for Detailed Image Captioning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mohamed, Abdelrahman, Kementchedjhieva, Yova |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation
von: Chen, Xiaofu, et al.
Veröffentlicht: (2025)
von: Chen, Xiaofu, et al.
Veröffentlicht: (2025)
LLMs Can Compensate for Deficiencies in Visual Representations
von: Takishita, Sho, et al.
Veröffentlicht: (2025)
von: Takishita, Sho, et al.
Veröffentlicht: (2025)
Long Story Short: Disentangling Compositionality and Long-Caption Understanding in Contrastive VLMs
von: Salazar, Israfel, et al.
Veröffentlicht: (2025)
von: Salazar, Israfel, et al.
Veröffentlicht: (2025)
EFSA: Episodic Few-Shot Adaptation for Text-to-Image Retrieval
von: Huzaifa, Muhammad, et al.
Veröffentlicht: (2024)
von: Huzaifa, Muhammad, et al.
Veröffentlicht: (2024)
A Simple Data Augmentation Strategy for Text-in-Image Scientific VQA
von: Shoer, Belal, et al.
Veröffentlicht: (2025)
von: Shoer, Belal, et al.
Veröffentlicht: (2025)
VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors
von: Shahgir, Haz Sameen, et al.
Veröffentlicht: (2026)
von: Shahgir, Haz Sameen, et al.
Veröffentlicht: (2026)
Do Vision and Language Models Share Concepts? A Vector Space Alignment Study
von: Li, Jiaang, et al.
Veröffentlicht: (2023)
von: Li, Jiaang, et al.
Veröffentlicht: (2023)
LinguDistill: Recovering Linguistic Ability in Vision-Language Models via Selective Cross-Modal Distillation
von: Irawan, Patrick Amadeus, et al.
Veröffentlicht: (2026)
von: Irawan, Patrick Amadeus, et al.
Veröffentlicht: (2026)
CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era
von: Cheng, Kanzhi, et al.
Veröffentlicht: (2025)
von: Cheng, Kanzhi, et al.
Veröffentlicht: (2025)
DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions
von: Wang, Xinran, et al.
Veröffentlicht: (2026)
von: Wang, Xinran, et al.
Veröffentlicht: (2026)
Precision or Recall? An Analysis of Image Captions for Training Text-to-Image Generation Model
von: Cheng, Sheng, et al.
Veröffentlicht: (2024)
von: Cheng, Sheng, et al.
Veröffentlicht: (2024)
CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning
von: Ibrahim, George, et al.
Veröffentlicht: (2025)
von: Ibrahim, George, et al.
Veröffentlicht: (2025)
Less is More: Mitigating Multimodal Hallucination from an EOS Decision Perspective
von: Yue, Zihao, et al.
Veröffentlicht: (2024)
von: Yue, Zihao, et al.
Veröffentlicht: (2024)
Text-Only Training for Image Captioning with Retrieval Augmentation and Modality Gap Correction
von: Fonseca, Rui, et al.
Veröffentlicht: (2025)
von: Fonseca, Rui, et al.
Veröffentlicht: (2025)
Unblocking Fine-Grained Evaluation of Detailed Captions: An Explaining AutoRater and Critic-and-Revise Pipeline
von: Gordon, Brian, et al.
Veröffentlicht: (2025)
von: Gordon, Brian, et al.
Veröffentlicht: (2025)
Differentiable JPEG: The Devil is in the Details
von: Reich, Christoph, et al.
Veröffentlicht: (2023)
von: Reich, Christoph, et al.
Veröffentlicht: (2023)
Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed Perception
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
Benchmarking and Improving Detail Image Caption
von: Dong, Hongyuan, et al.
Veröffentlicht: (2024)
von: Dong, Hongyuan, et al.
Veröffentlicht: (2024)
The Devil is in the Details: Simple Remedies for Image-to-LiDAR Representation Learning
von: Jo, Wonjun, et al.
Veröffentlicht: (2025)
von: Jo, Wonjun, et al.
Veröffentlicht: (2025)
From Image Captioning to Visual Storytelling
von: Passadakis, Admitos, et al.
Veröffentlicht: (2025)
von: Passadakis, Admitos, et al.
Veröffentlicht: (2025)
Noise is an Efficient Learner for Zero-Shot Vision-Language Models
von: Imam, Raza, et al.
Veröffentlicht: (2025)
von: Imam, Raza, et al.
Veröffentlicht: (2025)
Distinctive Image Captioning: Leveraging Ground Truth Captions in CLIP Guided Reinforcement Learning
von: Chaffin, Antoine, et al.
Veröffentlicht: (2024)
von: Chaffin, Antoine, et al.
Veröffentlicht: (2024)
ImageInWords: Unlocking Hyper-Detailed Image Descriptions
von: Garg, Roopal, et al.
Veröffentlicht: (2024)
von: Garg, Roopal, et al.
Veröffentlicht: (2024)
Image Captioning via Compact Bidirectional Architecture
von: Song, Zijie, et al.
Veröffentlicht: (2022)
von: Song, Zijie, et al.
Veröffentlicht: (2022)
MUNIChus: Multilingual News Image Captioning Benchmark
von: Chen, Yuji, et al.
Veröffentlicht: (2026)
von: Chen, Yuji, et al.
Veröffentlicht: (2026)
The Devil is in the Details: StyleFeatureEditor for Detail-Rich StyleGAN Inversion and High Quality Image Editing
von: Bobkov, Denis, et al.
Veröffentlicht: (2024)
von: Bobkov, Denis, et al.
Veröffentlicht: (2024)
Devil is in the Detail: Towards Injecting Fine Details of Image Prompt in Image Generation via Conflict-free Guidance and Stratified Attention
von: Jo, Kyungmin, et al.
Veröffentlicht: (2025)
von: Jo, Kyungmin, et al.
Veröffentlicht: (2025)
Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning
von: Li, Wenyan, et al.
Veröffentlicht: (2024)
von: Li, Wenyan, et al.
Veröffentlicht: (2024)
Temporal Image Caption Retrieval Competition -- Description and Results
von: Pokrywka, Jakub, et al.
Veröffentlicht: (2024)
von: Pokrywka, Jakub, et al.
Veröffentlicht: (2024)
Image Textualization: An Automatic Framework for Creating Accurate and Detailed Image Descriptions
von: Pi, Renjie, et al.
Veröffentlicht: (2024)
von: Pi, Renjie, et al.
Veröffentlicht: (2024)
The Devil is in the Prompts: Retrieval-Augmented Prompt Optimization for Text-to-Video Generation
von: Gao, Bingjie, et al.
Veröffentlicht: (2025)
von: Gao, Bingjie, et al.
Veröffentlicht: (2025)
VIXEN: Visual Text Comparison Network for Image Difference Captioning
von: Black, Alexander, et al.
Veröffentlicht: (2024)
von: Black, Alexander, et al.
Veröffentlicht: (2024)
Transformer based Multitask Learning for Image Captioning and Object Detection
von: Basak, Debolena, et al.
Veröffentlicht: (2024)
von: Basak, Debolena, et al.
Veröffentlicht: (2024)
Altogether: Image Captioning via Re-aligning Alt-text
von: Xu, Hu, et al.
Veröffentlicht: (2024)
von: Xu, Hu, et al.
Veröffentlicht: (2024)
OmniCaptioner: One Captioner to Rule Them All
von: Lu, Yiting, et al.
Veröffentlicht: (2025)
von: Lu, Yiting, et al.
Veröffentlicht: (2025)
See or Guess: Counterfactually Regularized Image Captioning
von: Cao, Qian, et al.
Veröffentlicht: (2024)
von: Cao, Qian, et al.
Veröffentlicht: (2024)
Unifying Vision-Language Latents for Zero-label Image Caption Enhancement
von: Byun, Sanghyun, et al.
Veröffentlicht: (2025)
von: Byun, Sanghyun, et al.
Veröffentlicht: (2025)
Improving Image Captioning by Mimicking Human Reformulation Feedback at Inference-time
von: Berger, Uri, et al.
Veröffentlicht: (2025)
von: Berger, Uri, et al.
Veröffentlicht: (2025)
Discovering Meaningful Units with Visually Grounded Semantics from Image Captions
von: Behjati, Melika, et al.
Veröffentlicht: (2025)
von: Behjati, Melika, et al.
Veröffentlicht: (2025)
#PraCegoVer: A Large Dataset for Image Captioning in Portuguese
von: Santos, Gabriel Oliveira dos, et al.
Veröffentlicht: (2021)
von: Santos, Gabriel Oliveira dos, et al.
Veröffentlicht: (2021)
Ähnliche Einträge
-
SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation
von: Chen, Xiaofu, et al.
Veröffentlicht: (2025) -
LLMs Can Compensate for Deficiencies in Visual Representations
von: Takishita, Sho, et al.
Veröffentlicht: (2025) -
Long Story Short: Disentangling Compositionality and Long-Caption Understanding in Contrastive VLMs
von: Salazar, Israfel, et al.
Veröffentlicht: (2025) -
EFSA: Episodic Few-Shot Adaptation for Text-to-Image Retrieval
von: Huzaifa, Muhammad, et al.
Veröffentlicht: (2024) -
A Simple Data Augmentation Strategy for Text-in-Image Scientific VQA
von: Shoer, Belal, et al.
Veröffentlicht: (2025)