Unblocking Fine-Grained Evaluation of Detailed Captions: An Explaining AutoRater and Critic-and-Revise Pipeline
Fuente:
arXiv
Saved in:
| Main Authors: | Gordon, Brian, Bitton, Yonatan, Marzoca, Andreea, Onoe, Yasumasa, Wang, Xiao, Cohen-Or, Daniel, Szpektor, Idan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bridging the Visual Gap: Fine-Tuning Multimodal Models with Knowledge-Adapted Captions
by: Yanuka, Moran, et al.
Published: (2024)
by: Yanuka, Moran, et al.
Published: (2024)
Mismatch Quest: Visual and Textual Feedback for Image-Text Misalignment
by: Gordon, Brian, et al.
Published: (2023)
by: Gordon, Brian, et al.
Published: (2023)
TALC: Time-Aligned Captions for Multi-Scene Text-to-Video Generation
by: Bansal, Hritik, et al.
Published: (2024)
by: Bansal, Hritik, et al.
Published: (2024)
Contrastive Sequential-Diffusion Learning: Non-linear and Multi-Scene Instructional Video Synthesis
by: Ramos, Vasco, et al.
Published: (2024)
by: Ramos, Vasco, et al.
Published: (2024)
ImageInWords: Unlocking Hyper-Detailed Image Descriptions
by: Garg, Roopal, et al.
Published: (2024)
by: Garg, Roopal, et al.
Published: (2024)
Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?
by: Zhang, Yue, et al.
Published: (2026)
by: Zhang, Yue, et al.
Published: (2026)
Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision
by: Zohar, Orr, et al.
Published: (2024)
by: Zohar, Orr, et al.
Published: (2024)
RefVNLI: Towards Scalable Evaluation of Subject-driven Text-to-image Generation
by: Slobodkin, Aviv, et al.
Published: (2025)
by: Slobodkin, Aviv, et al.
Published: (2025)
Visual Riddles: a Commonsense and World Knowledge Challenge for Large Vision and Language Models
by: Bitton-Guetta, Nitzan, et al.
Published: (2024)
by: Bitton-Guetta, Nitzan, et al.
Published: (2024)
Error-Driven Scene Editing for 3D Grounding in Large Language Models
by: Zhang, Yue, et al.
Published: (2025)
by: Zhang, Yue, et al.
Published: (2025)
Distinguishing Ignorance from Error in LLM Hallucinations
by: Simhi, Adi, et al.
Published: (2024)
by: Simhi, Adi, et al.
Published: (2024)
Constructing Benchmarks and Interventions for Combating Hallucinations in LLMs
by: Simhi, Adi, et al.
Published: (2024)
by: Simhi, Adi, et al.
Published: (2024)
Structsum Generation for Faster Text Comprehension
by: Jain, Parag, et al.
Published: (2024)
by: Jain, Parag, et al.
Published: (2024)
Generating Coherent Sequences of Visual Illustrations for Real-World Manual Tasks
by: Bordalo, João, et al.
Published: (2024)
by: Bordalo, João, et al.
Published: (2024)
Beyond the Noise: Aligning Prompts with Latent Representations in Diffusion Models
by: Ramos, Vasco, et al.
Published: (2025)
by: Ramos, Vasco, et al.
Published: (2025)
ParallelPARC: A Scalable Pipeline for Generating Natural-Language Analogies
by: Sultan, Oren, et al.
Published: (2024)
by: Sultan, Oren, et al.
Published: (2024)
3DLLM-Mem: Long-Term Spatial-Temporal Memory for Embodied 3D Large Language Model
by: Hu, Wenbo, et al.
Published: (2025)
by: Hu, Wenbo, et al.
Published: (2025)
MDCure: A Scalable Pipeline for Multi-Document Instruction-Following
by: Liu, Gabrielle Kaili-May, et al.
Published: (2024)
by: Liu, Gabrielle Kaili-May, et al.
Published: (2024)
No Detail Left Behind: Revisiting Self-Retrieval for Fine-Grained Image Captioning
by: Gaur, Manu, et al.
Published: (2024)
by: Gaur, Manu, et al.
Published: (2024)
ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
by: Simhi, Adi, et al.
Published: (2025)
by: Simhi, Adi, et al.
Published: (2025)
Latent Beam Diffusion Models for Generating Visual Sequences
by: Fernandes, Guilherme, et al.
Published: (2025)
by: Fernandes, Guilherme, et al.
Published: (2025)
Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed Perception
by: Ma, Ziyang, et al.
Published: (2025)
by: Ma, Ziyang, et al.
Published: (2025)
DOCCI: Descriptions of Connected and Contrasting Images
by: Onoe, Yasumasa, et al.
Published: (2024)
by: Onoe, Yasumasa, et al.
Published: (2024)
TECCI: Tricky Edits of Collected and Curated Images
by: Agrawal, Aishwarya, et al.
Published: (2026)
by: Agrawal, Aishwarya, et al.
Published: (2026)
TANQ: An open domain dataset of table answered questions
by: Akhtar, Mubashara, et al.
Published: (2024)
by: Akhtar, Mubashara, et al.
Published: (2024)
Editor's Choice: Evaluating Abstract Intent in Image Editing through Atomic Entity Analysis
by: Ventura, Mor, et al.
Published: (2026)
by: Ventura, Mor, et al.
Published: (2026)
Inside-Out: Hidden Factual Knowledge in LLMs
by: Gekhman, Zorik, et al.
Published: (2025)
by: Gekhman, Zorik, et al.
Published: (2025)
DetailCLIP: Detail-Oriented CLIP for Fine-Grained Tasks
by: Monsefi, Amin Karimi, et al.
Published: (2024)
by: Monsefi, Amin Karimi, et al.
Published: (2024)
LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
by: Orgad, Hadas, et al.
Published: (2024)
by: Orgad, Hadas, et al.
Published: (2024)
Are LLMs Better than Reported? Detecting Label Errors and Mitigating Their Effect on Model Performance
by: Nahum, Omer, et al.
Published: (2024)
by: Nahum, Omer, et al.
Published: (2024)
Beneath the Surface of Consistency: Exploring Cross-lingual Knowledge Representation Sharing in LLMs
by: Ifergan, Maxim, et al.
Published: (2024)
by: Ifergan, Maxim, et al.
Published: (2024)
The Devil is in the EOS: Sequence Training for Detailed Image Captioning
by: Mohamed, Abdelrahman, et al.
Published: (2025)
by: Mohamed, Abdelrahman, et al.
Published: (2025)
Benchmarking and Improving Detail Image Caption
by: Dong, Hongyuan, et al.
Published: (2024)
by: Dong, Hongyuan, et al.
Published: (2024)
Unblocking Placement and Routing in Rearrangeable Multi‐Stage Networks
by: Caio Von Rondow Morais, et al.
Published: (2025)
by: Caio Von Rondow Morais, et al.
Published: (2025)
Laurel: Unblocking Automated Verification with Large Language Models
by: Mugnier, Eric, et al.
Published: (2024)
by: Mugnier, Eric, et al.
Published: (2024)
Towards Fine-Grained Human Motion Video Captioning
by: Song, Guorui, et al.
Published: (2025)
by: Song, Guorui, et al.
Published: (2025)
EditInspector: A Benchmark for Evaluation of Text-Guided Image Edits
by: Yosef, Ron, et al.
Published: (2025)
by: Yosef, Ron, et al.
Published: (2025)
Detecting Stylistic Fingerprints of Large Language Models
by: Bitton, Yehonatan, et al.
Published: (2025)
by: Bitton, Yehonatan, et al.
Published: (2025)
DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions
by: Wang, Xinran, et al.
Published: (2026)
by: Wang, Xinran, et al.
Published: (2026)
CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era
by: Cheng, Kanzhi, et al.
Published: (2025)
by: Cheng, Kanzhi, et al.
Published: (2025)
Similar Items
-
Bridging the Visual Gap: Fine-Tuning Multimodal Models with Knowledge-Adapted Captions
by: Yanuka, Moran, et al.
Published: (2024) -
Mismatch Quest: Visual and Textual Feedback for Image-Text Misalignment
by: Gordon, Brian, et al.
Published: (2023) -
TALC: Time-Aligned Captions for Multi-Scene Text-to-Video Generation
by: Bansal, Hritik, et al.
Published: (2024) -
Contrastive Sequential-Diffusion Learning: Non-linear and Multi-Scene Instructional Video Synthesis
by: Ramos, Vasco, et al.
Published: (2024) -
ImageInWords: Unlocking Hyper-Detailed Image Descriptions
by: Garg, Roopal, et al.
Published: (2024)