BLUEX Revisited: Enhancing Benchmark Coverage with Automatic Captioning
Fuente:
arXiv
Saved in:
| Main Authors: | Santos, João Guilherme Alves, Bonás, Giovana Kerche, Almeida, Thales Sales |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TiEBe: Tracking Language Model Recall of Notable Worldwide Events Through Time
by: Almeida, Thales Sales, et al.
Published: (2025)
by: Almeida, Thales Sales, et al.
Published: (2025)
BRoverbs -- Measuring how much LLMs understand Portuguese proverbs
by: Almeida, Thales Sales, et al.
Published: (2025)
by: Almeida, Thales Sales, et al.
Published: (2025)
Ticket-Bench: A Kickoff for Multilingual and Regionalized Agent Evaluation
by: Almeida, Thales Sales, et al.
Published: (2025)
by: Almeida, Thales Sales, et al.
Published: (2025)
Sabiá-3 Technical Report
by: Abonizio, Hugo, et al.
Published: (2024)
by: Abonizio, Hugo, et al.
Published: (2024)
Prosa: Rubric-Based Evaluation of LLMs on Real User Chats in Brazilian Portuguese
by: Junior, Roseval Malaquias, et al.
Published: (2026)
by: Junior, Roseval Malaquias, et al.
Published: (2026)
Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks
by: Pires, Ramon, et al.
Published: (2026)
by: Pires, Ramon, et al.
Published: (2026)
MARCA: A Checklist-Based Benchmark for Multilingual Web Search
by: Almeida, Thales Sales, et al.
Published: (2026)
by: Almeida, Thales Sales, et al.
Published: (2026)
CAPITU: A Benchmark for Evaluating Instruction-Following in Brazilian Portuguese with Literary Context
by: Bonás, Giovana Kerche, et al.
Published: (2026)
by: Bonás, Giovana Kerche, et al.
Published: (2026)
Sabiá-2: A New Generation of Portuguese Large Language Models
by: Almeida, Thales Sales, et al.
Published: (2024)
by: Almeida, Thales Sales, et al.
Published: (2024)
Sabiá-4 Technical Report
by: Laitz, Thiago, et al.
Published: (2026)
by: Laitz, Thiago, et al.
Published: (2026)
Measuring Opinion Bias and Sycophancy via LLM-based Persuasion
by: Nogueira, Rodrigo, et al.
Published: (2026)
by: Nogueira, Rodrigo, et al.
Published: (2026)
LLM-Based Persuasion Enables Guardrail Override in Frontier LLMs
by: Nogueira, Rodrigo, et al.
Published: (2026)
by: Nogueira, Rodrigo, et al.
Published: (2026)
The interplay between domain specialization and model size
by: Junior, Roseval Malaquias, et al.
Published: (2025)
by: Junior, Roseval Malaquias, et al.
Published: (2025)
MolCap-Arena: A Comprehensive Captioning Benchmark on Language-Enhanced Molecular Property Prediction
by: Edwards, Carl, et al.
Published: (2024)
by: Edwards, Carl, et al.
Published: (2024)
DENEB: A Hallucination-Robust Automatic Evaluation Metric for Image Captioning
by: Matsuda, Kazuki, et al.
Published: (2024)
by: Matsuda, Kazuki, et al.
Published: (2024)
C2-Faith: Benchmarking LLM Judges for Causal and Coverage Faithfulness in Chain-of-Thought Reasoning
by: Mittal, Avni, et al.
Published: (2026)
by: Mittal, Avni, et al.
Published: (2026)
Rule-driven News Captioning
by: Xu, Ning, et al.
Published: (2024)
by: Xu, Ning, et al.
Published: (2024)
Fine-Tuning MedGemma for Clinical Captioning to Enhance Multimodal RAG over Malaysia CPGs
by: Zun, Lee Qi, et al.
Published: (2025)
by: Zun, Lee Qi, et al.
Published: (2025)
PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
by: Wang, Yidong, et al.
Published: (2023)
by: Wang, Yidong, et al.
Published: (2023)
Revisiting Image Captioning Training Paradigm via Direct CLIP-based Optimization
by: Moratelli, Nicholas, et al.
Published: (2024)
by: Moratelli, Nicholas, et al.
Published: (2024)
Humor in AI: Massive Scale Crowd-Sourced Preferences and Benchmarks for Cartoon Captioning
by: Zhang, Jifan, et al.
Published: (2024)
by: Zhang, Jifan, et al.
Published: (2024)
Proof2Hybrid: Automatic Mathematical Benchmark Synthesis for Proof-Centric Problems
by: Peng, Yebo, et al.
Published: (2025)
by: Peng, Yebo, et al.
Published: (2025)
TestAgent: Automatic Benchmarking and Exploratory Interaction for Evaluating LLMs in Vertical Domains
by: Wang, Wanying, et al.
Published: (2024)
by: Wang, Wanying, et al.
Published: (2024)
A New Benchmark for Evaluating Automatic Speech Recognition in the Arabic Call Domain
by: Obaidah, Qusai Abo, et al.
Published: (2024)
by: Obaidah, Qusai Abo, et al.
Published: (2024)
MIRAGE-Bench: Automatic Multilingual Benchmark Arena for Retrieval-Augmented Generation Systems
by: Thakur, Nandan, et al.
Published: (2024)
by: Thakur, Nandan, et al.
Published: (2024)
SGSimEval: A Comprehensive Multifaceted and Similarity-Enhanced Benchmark for Automatic Survey Generation Systems
by: Guo, Beichen, et al.
Published: (2025)
by: Guo, Beichen, et al.
Published: (2025)
Zero-shot Benchmarking: A Framework for Flexible and Scalable Automatic Evaluation of Language Models
by: Pombal, José, et al.
Published: (2025)
by: Pombal, José, et al.
Published: (2025)
Controlling Multimodal Conversational Agents with Coverage-Enhanced Latent Actions
by: Li, Yongqi, et al.
Published: (2026)
by: Li, Yongqi, et al.
Published: (2026)
Coverage-based Fairness in Multi-document Summarization
by: Li, Haoyuan, et al.
Published: (2024)
by: Li, Haoyuan, et al.
Published: (2024)
A Data-Driven Guided Decoding Mechanism for Diagnostic Captioning
by: Kaliosis, Panagiotis, et al.
Published: (2024)
by: Kaliosis, Panagiotis, et al.
Published: (2024)
Understanding How Paper Writers Use AI-Generated Captions in Figure Caption Writing
by: Yin, Ho, et al.
Published: (2025)
by: Yin, Ho, et al.
Published: (2025)
Towards Fairness Assessment of Dutch Hate Speech Detection
by: Bauer, Julie, et al.
Published: (2025)
by: Bauer, Julie, et al.
Published: (2025)
Unveiling Effective In-Context Configurations for Image Captioning: An External & Internal Analysis
by: Li, Li, et al.
Published: (2025)
by: Li, Li, et al.
Published: (2025)
AFRICAPTION: Establishing a New Paradigm for Image Captioning in African Languages
by: Oduwole, Mardiyyah, et al.
Published: (2025)
by: Oduwole, Mardiyyah, et al.
Published: (2025)
How to Understand Named Entities: Using Common Sense for News Captioning
by: Xu, Ning, et al.
Published: (2024)
by: Xu, Ning, et al.
Published: (2024)
From Policy to Logic for Efficient and Interpretable Coverage Assessment
by: Pokharel, Rhitabrat, et al.
Published: (2026)
by: Pokharel, Rhitabrat, et al.
Published: (2026)
No Culture Left Behind: ArtELingo-28, a Benchmark of WikiArt with Captions in 28 Languages
by: Mohamed, Youssef, et al.
Published: (2024)
by: Mohamed, Youssef, et al.
Published: (2024)
Rate, Explain and Cite (REC): Enhanced Explanation and Attribution in Automatic Evaluation by Large Language Models
by: Hsu, Aliyah R., et al.
Published: (2024)
by: Hsu, Aliyah R., et al.
Published: (2024)
Spectrum Tuning: Post-Training for Distributional Coverage and In-Context Steerability
by: Sorensen, Taylor, et al.
Published: (2025)
by: Sorensen, Taylor, et al.
Published: (2025)
Evaluation of Multilingual Image Captioning: How far can we get with CLIP models?
by: Gomes, Gonçalo, et al.
Published: (2025)
by: Gomes, Gonçalo, et al.
Published: (2025)
Similar Items
-
TiEBe: Tracking Language Model Recall of Notable Worldwide Events Through Time
by: Almeida, Thales Sales, et al.
Published: (2025) -
BRoverbs -- Measuring how much LLMs understand Portuguese proverbs
by: Almeida, Thales Sales, et al.
Published: (2025) -
Ticket-Bench: A Kickoff for Multilingual and Regionalized Agent Evaluation
by: Almeida, Thales Sales, et al.
Published: (2025) -
Sabiá-3 Technical Report
by: Abonizio, Hugo, et al.
Published: (2024) -
Prosa: Rubric-Based Evaluation of LLMs on Real User Chats in Brazilian Portuguese
by: Junior, Roseval Malaquias, et al.
Published: (2026)