Charting the European LLM Benchmarking Landscape: A New Taxonomy and a Set of Best Practices
Fuente:
arXiv
Saved in:
| Main Authors: | Vintar, Špela, Pungeršek, Taja Kuzman, Brglez, Mojca, Ljubešić, Nikola |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Polyester Girlfriends to Blind Mice: Creating the First Pragmatics Understanding Benchmarks for Slovene
by: Brglez, Mojca, et al.
Published: (2025)
by: Brglez, Mojca, et al.
Published: (2025)
Supercharging Agenda Setting Research: The ParlaCAP Dataset of 28 European Parliaments and a Scalable Multilingual LLM-Based Classification
by: Pungeršek, Taja Kuzman, et al.
Published: (2026)
by: Pungeršek, Taja Kuzman, et al.
Published: (2026)
State of the Art in Text Classification for South Slavic Languages: Fine-Tuning or Prompting?
by: Pungeršek, Taja Kuzman, et al.
Published: (2025)
by: Pungeršek, Taja Kuzman, et al.
Published: (2025)
LLM Teacher-Student Framework for Text Classification With No Manually Annotated Data: A Case Study in IPTC News Topic Classification
by: Kuzman, Taja, et al.
Published: (2024)
by: Kuzman, Taja, et al.
Published: (2024)
CLASSLA-web: Comparable Web Corpora of South Slavic Languages Enriched with Linguistic and Genre Annotation
by: Ljubešić, Nikola, et al.
Published: (2024)
by: Ljubešić, Nikola, et al.
Published: (2024)
ParlaSpeech 3.0: Richly Annotated Spoken Parliamentary Corpora of Croatian, Czech, Polish, and Serbian
by: Ljubešić, Nikola, et al.
Published: (2025)
by: Ljubešić, Nikola, et al.
Published: (2025)
The Growing Gains and Pains of Iterative Web Corpora Crawling: Insights from South Slavic CLASSLA-web 2.0 Corpora
by: Pungeršek, Taja Kuzman, et al.
Published: (2026)
by: Pungeršek, Taja Kuzman, et al.
Published: (2026)
CLASSLA-Express: a Train of CLARIN.SI Workshops on Language Resources and Tools with Easily Expanding Route
by: Ljubešić, Nikola, et al.
Published: (2024)
by: Ljubešić, Nikola, et al.
Published: (2024)
Language Models on a Diet: Cost-Efficient Development of Encoders for Closely-Related Languages via Additional Pretraining
by: Ljubešić, Nikola, et al.
Published: (2024)
by: Ljubešić, Nikola, et al.
Published: (2024)
The truth is no diaper: Human and AI-generated associations to emotional words
by: Vintar, Špela, et al.
Published: (2025)
by: Vintar, Špela, et al.
Published: (2025)
Do Language Models Care About Text Quality? Evaluating Web-Crawled Corpora Across 11 Languages
by: van Noord, Rik, et al.
Published: (2024)
by: van Noord, Rik, et al.
Published: (2024)
Chart-HQA: A Benchmark for Hypothetical Question Answering in Charts
by: Chen, Xiangnan, et al.
Published: (2025)
by: Chen, Xiangnan, et al.
Published: (2025)
ErrorMap and ErrorAtlas: Charting the Failure Landscape of Large Language Models
by: Ashury-Tahan, Shir, et al.
Published: (2026)
by: Ashury-Tahan, Shir, et al.
Published: (2026)
Towards Best Practices for Open Datasets for LLM Training
by: Baack, Stefan, et al.
Published: (2025)
by: Baack, Stefan, et al.
Published: (2025)
A Computational Analysis of the Dehumanisation of Migrants from Syria and Ukraine in Slovene News Media
by: Caporusso, Jaya, et al.
Published: (2024)
by: Caporusso, Jaya, et al.
Published: (2024)
ChartCards: A Chart-Metadata Generation Framework for Multi-Task Chart Understanding
by: Wu, Yifan, et al.
Published: (2025)
by: Wu, Yifan, et al.
Published: (2025)
Enhancing Retrieval-Augmented Generation: A Study of Best Practices
by: Li, Siran, et al.
Published: (2025)
by: Li, Siran, et al.
Published: (2025)
A Taxonomy of Prompt Defects in LLM Systems
by: Tian, Haoye, et al.
Published: (2025)
by: Tian, Haoye, et al.
Published: (2025)
Automated Benchmark Generation from Domain Guidelines Informed by Bloom's Taxonomy
by: Chen, Si, et al.
Published: (2026)
by: Chen, Si, et al.
Published: (2026)
Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems
by: Cui, Tianyu, et al.
Published: (2024)
by: Cui, Tianyu, et al.
Published: (2024)
Overhearing LLM Agents: A Survey, Taxonomy, and Roadmap
by: Zhu, Andrew, et al.
Published: (2025)
by: Zhu, Andrew, et al.
Published: (2025)
Beyond SELECT: A Comprehensive Taxonomy-Guided Benchmark for Real-World Text-to-SQL Translation
by: Wang, Hao, et al.
Published: (2025)
by: Wang, Hao, et al.
Published: (2025)
Unmasking Deceptive Visuals: Benchmarking Multimodal Large Language Models on Misleading Chart Question Answering
by: Chen, Zixin, et al.
Published: (2025)
by: Chen, Zixin, et al.
Published: (2025)
Maintaining Journalistic Integrity in the Digital Age: A Comprehensive NLP Framework for Evaluating Online News Content
by: Bojic, Ljubisa, et al.
Published: (2024)
by: Bojic, Ljubisa, et al.
Published: (2024)
Combining the Best of Both Worlds: A Method for Hybrid NMT and LLM Translation
by: Wu, Zhanglin, et al.
Published: (2025)
by: Wu, Zhanglin, et al.
Published: (2025)
Consistent and Distinctive: LLM Benchmark Efficiency via Maximum Independent Set Prompt Selection on Similarity Graphs
by: Kjorvezir, Denica, et al.
Published: (2026)
by: Kjorvezir, Denica, et al.
Published: (2026)
Exploring Multimodal Challenges in Toxic Chinese Detection: Taxonomy, Benchmark, and Findings
by: Yang, Shujian, et al.
Published: (2025)
by: Yang, Shujian, et al.
Published: (2025)
ChartCitor: Multi-Agent Framework for Fine-Grained Chart Visual Attribution
by: Goswami, Kanika, et al.
Published: (2025)
by: Goswami, Kanika, et al.
Published: (2025)
COMPL-AI Framework: A Technical Interpretation and LLM Benchmarking Suite for the EU Artificial Intelligence Act
by: Guldimann, Philipp, et al.
Published: (2024)
by: Guldimann, Philipp, et al.
Published: (2024)
LELA: An End-to-end LLM-based Entity Linking Framework with Zero-shot Domain Adaptation
by: Haffoudhi, Samy, et al.
Published: (2026)
by: Haffoudhi, Samy, et al.
Published: (2026)
LiveFact: A Dynamic, Time-Aware Benchmark for LLM-Driven Fake News Detection
by: Xu, Cheng, et al.
Published: (2026)
by: Xu, Cheng, et al.
Published: (2026)
CHARTOM: A Visual Theory-of-Mind Benchmark for LLMs on Misleading Charts
by: Bharti, Shubham, et al.
Published: (2024)
by: Bharti, Shubham, et al.
Published: (2024)
ChartEditBench: Evaluating Grounded Multi-Turn Chart Editing in Multimodal Language Models
by: Kapadnis, Manav Nitin, et al.
Published: (2026)
by: Kapadnis, Manav Nitin, et al.
Published: (2026)
LLM-Generated Negative News Headlines Dataset: Creation and Benchmarking Against Real Journalism
by: Babalola, Olusola, et al.
Published: (2025)
by: Babalola, Olusola, et al.
Published: (2025)
Benchmark Profiling: Mechanistic Diagnosis of LLM Benchmarks
by: Kim, Dongjun, et al.
Published: (2025)
by: Kim, Dongjun, et al.
Published: (2025)
RAGTurk: Best Practices for Retrieval Augmented Generation in Turkish
by: Köse, Süha Kağan, et al.
Published: (2026)
by: Köse, Süha Kağan, et al.
Published: (2026)
A Geometric Taxonomy of Hallucinations in LLMs
by: Marín, Javier
Published: (2026)
by: Marín, Javier
Published: (2026)
WikiMixQA: A Multimodal Benchmark for Question Answering over Tables and Charts
by: Foroutan, Negar, et al.
Published: (2025)
by: Foroutan, Negar, et al.
Published: (2025)
Who Benchmarks the Benchmarks? A Case Study of LLM Evaluation in Icelandic
by: Ingimundarson, Finnur Ágúst, et al.
Published: (2026)
by: Ingimundarson, Finnur Ágúst, et al.
Published: (2026)
$C^2$: Scalable Auto-Feedback for LLM-based Chart Generation
by: Koh, Woosung, et al.
Published: (2024)
by: Koh, Woosung, et al.
Published: (2024)
Similar Items
-
From Polyester Girlfriends to Blind Mice: Creating the First Pragmatics Understanding Benchmarks for Slovene
by: Brglez, Mojca, et al.
Published: (2025) -
Supercharging Agenda Setting Research: The ParlaCAP Dataset of 28 European Parliaments and a Scalable Multilingual LLM-Based Classification
by: Pungeršek, Taja Kuzman, et al.
Published: (2026) -
State of the Art in Text Classification for South Slavic Languages: Fine-Tuning or Prompting?
by: Pungeršek, Taja Kuzman, et al.
Published: (2025) -
LLM Teacher-Student Framework for Text Classification With No Manually Annotated Data: A Case Study in IPTC News Topic Classification
by: Kuzman, Taja, et al.
Published: (2024) -
CLASSLA-web: Comparable Web Corpora of South Slavic Languages Enriched with Linguistic and Genre Annotation
by: Ljubešić, Nikola, et al.
Published: (2024)