What Makes for Good Image Captions?
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Delong, Cahyawijaya, Samuel, Ishii, Etsuko, Chan, Ho Shu, Bang, Yejin, Fung, Pascale |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Subobject-level Image Tokenization
by: Chen, Delong, et al.
Published: (2024)
by: Chen, Delong, et al.
Published: (2024)
WorldPrediction: A Benchmark for High-level World Modeling and Long-horizon Procedural Planning
by: Chen, Delong, et al.
Published: (2025)
by: Chen, Delong, et al.
Published: (2025)
LLM Internal States Reveal Hallucination Risk Faced With a Query
by: Ji, Ziwei, et al.
Published: (2024)
by: Ji, Ziwei, et al.
Published: (2024)
Negative Object Presence Evaluation (NOPE) to Measure Object Hallucination in Vision-Language Models
by: Lovenia, Holy, et al.
Published: (2023)
by: Lovenia, Holy, et al.
Published: (2023)
Action100M: A Large-scale Video Action Dataset
by: Chen, Delong, et al.
Published: (2026)
by: Chen, Delong, et al.
Published: (2026)
High-Dimension Human Value Representation in Large Language Models
by: Cahyawijaya, Samuel, et al.
Published: (2024)
by: Cahyawijaya, Samuel, et al.
Published: (2024)
VL-JEPA: Joint Embedding Predictive Architecture for Vision-language
by: Chen, Delong, et al.
Published: (2025)
by: Chen, Delong, et al.
Published: (2025)
What Makes for a Good Stereoscopic Image?
by: Tamir, Netanel Y., et al.
Published: (2024)
by: Tamir, Netanel Y., et al.
Published: (2024)
Making Large Vision Language Models to be Good Few-shot Learners
by: Liu, Fan, et al.
Published: (2024)
by: Liu, Fan, et al.
Published: (2024)
Belief Revision: The Adaptability of Large Language Models Reasoning
by: Wilie, Bryan, et al.
Published: (2024)
by: Wilie, Bryan, et al.
Published: (2024)
Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions
by: Shen, Yijun, et al.
Published: (2025)
by: Shen, Yijun, et al.
Published: (2025)
What Makes a Good Dataset for Knowledge Distillation?
by: Frank, Logan, et al.
Published: (2024)
by: Frank, Logan, et al.
Published: (2024)
Measuring Political Bias in Large Language Models: What Is Said and How It Is Said
by: Bang, Yejin, et al.
Published: (2024)
by: Bang, Yejin, et al.
Published: (2024)
What Makes a Good Generated Image? Investigating Human and Multimodal LLM Image Preference Alignment
by: Parthasarathy, Rishab, et al.
Published: (2025)
by: Parthasarathy, Rishab, et al.
Published: (2025)
User-Aware Prefix-Tuning is a Good Learner for Personalized Image Captioning
by: Wang, Xuan, et al.
Published: (2023)
by: Wang, Xuan, et al.
Published: (2023)
What Makes Good Few-shot Examples for Vision-Language Models?
by: Guo, Zhaojun, et al.
Published: (2024)
by: Guo, Zhaojun, et al.
Published: (2024)
MedM-VL: What Makes a Good Medical LVLM?
by: Shi, Yiming, et al.
Published: (2025)
by: Shi, Yiming, et al.
Published: (2025)
Image Understanding Makes for A Good Tokenizer for Image Generation
by: Wang, Luting, et al.
Published: (2024)
by: Wang, Luting, et al.
Published: (2024)
What Makes Good Synthetic Training Data for Zero-Shot Stereo Matching?
by: Yan, David, et al.
Published: (2025)
by: Yan, David, et al.
Published: (2025)
Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor
by: Lin, Ethan, et al.
Published: (2025)
by: Lin, Ethan, et al.
Published: (2025)
VideoRoPE: What Makes for Good Video Rotary Position Embedding?
by: Wei, Xilin, et al.
Published: (2025)
by: Wei, Xilin, et al.
Published: (2025)
SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning
by: Zhang, Lin, et al.
Published: (2025)
by: Zhang, Lin, et al.
Published: (2025)
CultureCLIP: Empowering CLIP with Cultural Awareness through Synthetic Images and Contextualized Captions
by: Huang, Yuchen, et al.
Published: (2025)
by: Huang, Yuchen, et al.
Published: (2025)
Good Seed Makes a Good Crop: Discovering Secret Seeds in Text-to-Image Diffusion Models
by: Xu, Katherine, et al.
Published: (2024)
by: Xu, Katherine, et al.
Published: (2024)
CaptionQA: Is Your Caption as Useful as the Image Itself?
by: Yang, Shijia, et al.
Published: (2025)
by: Yang, Shijia, et al.
Published: (2025)
How to Make Cross Encoder a Good Teacher for Efficient Image-Text Retrieval?
by: Chen, Yuxin, et al.
Published: (2024)
by: Chen, Yuxin, et al.
Published: (2024)
Towards Efficient and Robust VQA-NLE Data Generation with Large Vision-Language Models
by: Irawan, Patrick Amadeus, et al.
Published: (2024)
by: Irawan, Patrick Amadeus, et al.
Published: (2024)
CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning
by: Saito, Kuniaki, et al.
Published: (2025)
by: Saito, Kuniaki, et al.
Published: (2025)
Good Noise Makes Good Edits: A Training-Free Diffusion-Based Video Editing with Image and Text Prompts
by: Choi, Saemee, et al.
Published: (2025)
by: Choi, Saemee, et al.
Published: (2025)
Group-based Distinctive Image Captioning with Memory Difference Encoding and Attention
by: Wang, Jiuniu, et al.
Published: (2025)
by: Wang, Jiuniu, et al.
Published: (2025)
Latent Denoising Makes Good Tokenizers
by: Yang, Jiawei, et al.
Published: (2025)
by: Yang, Jiawei, et al.
Published: (2025)
Exploring Diverse In-Context Configurations for Image Captioning
by: Yang, Xu, et al.
Published: (2023)
by: Yang, Xu, et al.
Published: (2023)
Style Transfer Dataset: What Makes A Good Stylization?
by: Kitov, Victor, et al.
Published: (2024)
by: Kitov, Victor, et al.
Published: (2024)
What Makes Synthetic Data Effective in Image Segmentation
by: Zhang, Jinjin, et al.
Published: (2026)
by: Zhang, Jinjin, et al.
Published: (2026)
What Makes Good Collaborative Views? Contrastive Mutual Information Maximization for Multi-Agent Perception
by: Su, Wanfang, et al.
Published: (2024)
by: Su, Wanfang, et al.
Published: (2024)
MeaCap: Memory-Augmented Zero-shot Image Captioning
by: Zeng, Zequn, et al.
Published: (2024)
by: Zeng, Zequn, et al.
Published: (2024)
Benchmarking and Improving Detail Image Caption
by: Dong, Hongyuan, et al.
Published: (2024)
by: Dong, Hongyuan, et al.
Published: (2024)
What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
by: Du, Yifan, et al.
Published: (2023)
by: Du, Yifan, et al.
Published: (2023)
HICEScore: A Hierarchical Metric for Image Captioning Evaluation
by: Zeng, Zequn, et al.
Published: (2024)
by: Zeng, Zequn, et al.
Published: (2024)
Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning
by: Jeon, MinJu, et al.
Published: (2025)
by: Jeon, MinJu, et al.
Published: (2025)
Similar Items
-
Subobject-level Image Tokenization
by: Chen, Delong, et al.
Published: (2024) -
WorldPrediction: A Benchmark for High-level World Modeling and Long-horizon Procedural Planning
by: Chen, Delong, et al.
Published: (2025) -
LLM Internal States Reveal Hallucination Risk Faced With a Query
by: Ji, Ziwei, et al.
Published: (2024) -
Negative Object Presence Evaluation (NOPE) to Measure Object Hallucination in Vision-Language Models
by: Lovenia, Holy, et al.
Published: (2023) -
Action100M: A Large-scale Video Action Dataset
by: Chen, Delong, et al.
Published: (2026)