Diagnosing Translated Benchmarks: An Automated Quality Assurance Study of the EU20 Benchmark Suite
Fuente:
arXiv
Salvato in:
| Autori principali: | Thellmann, Klaudia, Stadler, Bernhard, Färber, Michael |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Quantifying the Impact of Translation Errors on Multilingual LLM Evaluation
di: Thellmann, Klaudia-Doris, et al.
Pubblicazione: (2026)
di: Thellmann, Klaudia-Doris, et al.
Pubblicazione: (2026)
LogEval: A Comprehensive Benchmark Suite for Large Language Models In Log Analysis
di: Cui, Tianyu, et al.
Pubblicazione: (2024)
di: Cui, Tianyu, et al.
Pubblicazione: (2024)
A Benchmark Suite of Reddit-Derived Datasets for Mental Health Detection
di: Hasan, Khalid, et al.
Pubblicazione: (2026)
di: Hasan, Khalid, et al.
Pubblicazione: (2026)
AIR-Bench: Automated Heterogeneous Information Retrieval Benchmark
di: Chen, Jianlyu, et al.
Pubblicazione: (2024)
di: Chen, Jianlyu, et al.
Pubblicazione: (2024)
Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference
di: Kummer, Cornelius, et al.
Pubblicazione: (2026)
di: Kummer, Cornelius, et al.
Pubblicazione: (2026)
SQuAI: Scientific Question-Answering with Multi-Agent Retrieval-Augmented Generation
di: Besrour, Ines, et al.
Pubblicazione: (2025)
di: Besrour, Ines, et al.
Pubblicazione: (2025)
HyperPIE: Hyperparameter Information Extraction from Scientific Publications
di: Saier, Tarek, et al.
Pubblicazione: (2023)
di: Saier, Tarek, et al.
Pubblicazione: (2023)
Revisiting Projection-based Data Transfer for Cross-Lingual Named Entity Recognition in Low-Resource Languages
di: Politov, Andrei, et al.
Pubblicazione: (2025)
di: Politov, Andrei, et al.
Pubblicazione: (2025)
Frustratingly Simple Retrieval Improves Challenging, Reasoning-Intensive Benchmarks
di: Lyu, Xinxi, et al.
Pubblicazione: (2025)
di: Lyu, Xinxi, et al.
Pubblicazione: (2025)
The Effects of Hallucinations in Synthetic Training Data for Relation Extraction
di: Rogulsky, Steven, et al.
Pubblicazione: (2024)
di: Rogulsky, Steven, et al.
Pubblicazione: (2024)
RAR-b: Reasoning as Retrieval Benchmark
di: Xiao, Chenghao, et al.
Pubblicazione: (2024)
di: Xiao, Chenghao, et al.
Pubblicazione: (2024)
Benchmarking LLM-based Relevance Judgment Methods
di: Arabzadeh, Negar, et al.
Pubblicazione: (2025)
di: Arabzadeh, Negar, et al.
Pubblicazione: (2025)
Paths to Causality: Finding Informative Subgraphs Within Knowledge Graphs for Knowledge-Based Causal Discovery
di: Susanti, Yuni, et al.
Pubblicazione: (2025)
di: Susanti, Yuni, et al.
Pubblicazione: (2025)
JFinTEB: Japanese Financial Text Embedding Benchmark
di: Suzuki, Masahiro, et al.
Pubblicazione: (2026)
di: Suzuki, Masahiro, et al.
Pubblicazione: (2026)
SAGE: Benchmarking and Improving Retrieval for Deep Research Agents
di: Hu, Tiansheng, et al.
Pubblicazione: (2026)
di: Hu, Tiansheng, et al.
Pubblicazione: (2026)
ScholarSearch: Benchmarking Scholar Searching Ability of LLMs
di: Zhou, Junting, et al.
Pubblicazione: (2025)
di: Zhou, Junting, et al.
Pubblicazione: (2025)
Building Russian Benchmark for Evaluation of Information Retrieval Models
di: Kovalev, Grigory, et al.
Pubblicazione: (2025)
di: Kovalev, Grigory, et al.
Pubblicazione: (2025)
FinMTEB: Finance Massive Text Embedding Benchmark
di: Tang, Yixuan, et al.
Pubblicazione: (2025)
di: Tang, Yixuan, et al.
Pubblicazione: (2025)
DAPR: A Benchmark on Document-Aware Passage Retrieval
di: Wang, Kexin, et al.
Pubblicazione: (2023)
di: Wang, Kexin, et al.
Pubblicazione: (2023)
BERGEN: A Benchmarking Library for Retrieval-Augmented Generation
di: Rau, David, et al.
Pubblicazione: (2024)
di: Rau, David, et al.
Pubblicazione: (2024)
Optimizing Small Transformer-Based Language Models for Multi-Label Sentiment Analysis in Short Texts
di: Neumann, Julius, et al.
Pubblicazione: (2025)
di: Neumann, Julius, et al.
Pubblicazione: (2025)
FollowTable: A Benchmark for Instruction-Following Table Retrieval
di: Jin, Rihui, et al.
Pubblicazione: (2026)
di: Jin, Rihui, et al.
Pubblicazione: (2026)
PARSE: An Open-Domain Reasoning Question Answering Benchmark for Persian
di: Mozafari, Jamshid, et al.
Pubblicazione: (2026)
di: Mozafari, Jamshid, et al.
Pubblicazione: (2026)
TabEmbed: Benchmarking and Learning Generalist Embeddings for Tabular Understanding
di: Qiang, Minjie, et al.
Pubblicazione: (2026)
di: Qiang, Minjie, et al.
Pubblicazione: (2026)
PJB: A Reasoning-Aware Benchmark for Person-Job Retrieval
di: Wang, Guangzhi, et al.
Pubblicazione: (2026)
di: Wang, Guangzhi, et al.
Pubblicazione: (2026)
PosIR: Position-Aware Heterogeneous Information Retrieval Benchmark
di: Zeng, Ziyang, et al.
Pubblicazione: (2026)
di: Zeng, Ziyang, et al.
Pubblicazione: (2026)
exHarmony: Authorship and Citations for Benchmarking the Reviewer Assignment Problem
di: Ebrahimi, Sajad, et al.
Pubblicazione: (2025)
di: Ebrahimi, Sajad, et al.
Pubblicazione: (2025)
MCiteBench: A Multimodal Benchmark for Generating Text with Citations
di: Hu, Caiyu, et al.
Pubblicazione: (2025)
di: Hu, Caiyu, et al.
Pubblicazione: (2025)
Hindi-BEIR : A Large Scale Retrieval Benchmark in Hindi
di: Acharya, Arkadeep, et al.
Pubblicazione: (2024)
di: Acharya, Arkadeep, et al.
Pubblicazione: (2024)
ORBIT -- Open Recommendation Benchmark for Reproducible Research with Hidden Tests
di: He, Jingyuan, et al.
Pubblicazione: (2025)
di: He, Jingyuan, et al.
Pubblicazione: (2025)
CORAL: Benchmarking Multi-turn Conversational Retrieval-Augmentation Generation
di: Cheng, Yiruo, et al.
Pubblicazione: (2024)
di: Cheng, Yiruo, et al.
Pubblicazione: (2024)
FAB-Bench: A Framework for Adaptive RAG Benchmarking in Semiconductor Manufacturing
di: Qian, Jingbin, et al.
Pubblicazione: (2026)
di: Qian, Jingbin, et al.
Pubblicazione: (2026)
JUÁ -- A Benchmark for Information Retrieval in Brazilian Legal Text Collections
di: Pereira, Jayr, et al.
Pubblicazione: (2026)
di: Pereira, Jayr, et al.
Pubblicazione: (2026)
ASTRA-QA: A Benchmark for Abstract Question Answering over Documents
di: Wang, Shu, et al.
Pubblicazione: (2026)
di: Wang, Shu, et al.
Pubblicazione: (2026)
Wikipedia-based Datasets in Russian Information Retrieval Benchmark RusBEIR
di: Kovalev, Grigory, et al.
Pubblicazione: (2025)
di: Kovalev, Grigory, et al.
Pubblicazione: (2025)
Generating Diverse Q&A Benchmarks for RAG Evaluation with DataMorgana
di: Filice, Simone, et al.
Pubblicazione: (2025)
di: Filice, Simone, et al.
Pubblicazione: (2025)
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
di: Du, Mingxuan, et al.
Pubblicazione: (2025)
di: Du, Mingxuan, et al.
Pubblicazione: (2025)
RankMamba: Benchmarking Mamba's Document Ranking Performance in the Era of Transformers
di: Xu, Zhichao
Pubblicazione: (2024)
di: Xu, Zhichao
Pubblicazione: (2024)
CoIR: A Comprehensive Benchmark for Code Information Retrieval Models
di: Li, Xiangyang, et al.
Pubblicazione: (2024)
di: Li, Xiangyang, et al.
Pubblicazione: (2024)
Diagnosing and Repairing Citation Failures in Generative Engine Optimization
di: Tian, Zhihua, et al.
Pubblicazione: (2026)
di: Tian, Zhihua, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Quantifying the Impact of Translation Errors on Multilingual LLM Evaluation
di: Thellmann, Klaudia-Doris, et al.
Pubblicazione: (2026) -
LogEval: A Comprehensive Benchmark Suite for Large Language Models In Log Analysis
di: Cui, Tianyu, et al.
Pubblicazione: (2024) -
A Benchmark Suite of Reddit-Derived Datasets for Mental Health Detection
di: Hasan, Khalid, et al.
Pubblicazione: (2026) -
AIR-Bench: Automated Heterogeneous Information Retrieval Benchmark
di: Chen, Jianlyu, et al.
Pubblicazione: (2024) -
Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference
di: Kummer, Cornelius, et al.
Pubblicazione: (2026)