Charting the European LLM Benchmarking Landscape: A New Taxonomy and a Set of Best Practices
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Vintar, Špela, Pungeršek, Taja Kuzman, Brglez, Mojca, Ljubešić, Nikola |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
From Polyester Girlfriends to Blind Mice: Creating the First Pragmatics Understanding Benchmarks for Slovene
par: Brglez, Mojca, et autres
Publié: (2025)
par: Brglez, Mojca, et autres
Publié: (2025)
Supercharging Agenda Setting Research: The ParlaCAP Dataset of 28 European Parliaments and a Scalable Multilingual LLM-Based Classification
par: Pungeršek, Taja Kuzman, et autres
Publié: (2026)
par: Pungeršek, Taja Kuzman, et autres
Publié: (2026)
State of the Art in Text Classification for South Slavic Languages: Fine-Tuning or Prompting?
par: Pungeršek, Taja Kuzman, et autres
Publié: (2025)
par: Pungeršek, Taja Kuzman, et autres
Publié: (2025)
LLM Teacher-Student Framework for Text Classification With No Manually Annotated Data: A Case Study in IPTC News Topic Classification
par: Kuzman, Taja, et autres
Publié: (2024)
par: Kuzman, Taja, et autres
Publié: (2024)
CLASSLA-web: Comparable Web Corpora of South Slavic Languages Enriched with Linguistic and Genre Annotation
par: Ljubešić, Nikola, et autres
Publié: (2024)
par: Ljubešić, Nikola, et autres
Publié: (2024)
ParlaSpeech 3.0: Richly Annotated Spoken Parliamentary Corpora of Croatian, Czech, Polish, and Serbian
par: Ljubešić, Nikola, et autres
Publié: (2025)
par: Ljubešić, Nikola, et autres
Publié: (2025)
The Growing Gains and Pains of Iterative Web Corpora Crawling: Insights from South Slavic CLASSLA-web 2.0 Corpora
par: Pungeršek, Taja Kuzman, et autres
Publié: (2026)
par: Pungeršek, Taja Kuzman, et autres
Publié: (2026)
CLASSLA-Express: a Train of CLARIN.SI Workshops on Language Resources and Tools with Easily Expanding Route
par: Ljubešić, Nikola, et autres
Publié: (2024)
par: Ljubešić, Nikola, et autres
Publié: (2024)
Language Models on a Diet: Cost-Efficient Development of Encoders for Closely-Related Languages via Additional Pretraining
par: Ljubešić, Nikola, et autres
Publié: (2024)
par: Ljubešić, Nikola, et autres
Publié: (2024)
The truth is no diaper: Human and AI-generated associations to emotional words
par: Vintar, Špela, et autres
Publié: (2025)
par: Vintar, Špela, et autres
Publié: (2025)
Do Language Models Care About Text Quality? Evaluating Web-Crawled Corpora Across 11 Languages
par: van Noord, Rik, et autres
Publié: (2024)
par: van Noord, Rik, et autres
Publié: (2024)
Chart-HQA: A Benchmark for Hypothetical Question Answering in Charts
par: Chen, Xiangnan, et autres
Publié: (2025)
par: Chen, Xiangnan, et autres
Publié: (2025)
ErrorMap and ErrorAtlas: Charting the Failure Landscape of Large Language Models
par: Ashury-Tahan, Shir, et autres
Publié: (2026)
par: Ashury-Tahan, Shir, et autres
Publié: (2026)
Towards Best Practices for Open Datasets for LLM Training
par: Baack, Stefan, et autres
Publié: (2025)
par: Baack, Stefan, et autres
Publié: (2025)
A Computational Analysis of the Dehumanisation of Migrants from Syria and Ukraine in Slovene News Media
par: Caporusso, Jaya, et autres
Publié: (2024)
par: Caporusso, Jaya, et autres
Publié: (2024)
ChartCards: A Chart-Metadata Generation Framework for Multi-Task Chart Understanding
par: Wu, Yifan, et autres
Publié: (2025)
par: Wu, Yifan, et autres
Publié: (2025)
Enhancing Retrieval-Augmented Generation: A Study of Best Practices
par: Li, Siran, et autres
Publié: (2025)
par: Li, Siran, et autres
Publié: (2025)
A Taxonomy of Prompt Defects in LLM Systems
par: Tian, Haoye, et autres
Publié: (2025)
par: Tian, Haoye, et autres
Publié: (2025)
Automated Benchmark Generation from Domain Guidelines Informed by Bloom's Taxonomy
par: Chen, Si, et autres
Publié: (2026)
par: Chen, Si, et autres
Publié: (2026)
Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems
par: Cui, Tianyu, et autres
Publié: (2024)
par: Cui, Tianyu, et autres
Publié: (2024)
Overhearing LLM Agents: A Survey, Taxonomy, and Roadmap
par: Zhu, Andrew, et autres
Publié: (2025)
par: Zhu, Andrew, et autres
Publié: (2025)
Beyond SELECT: A Comprehensive Taxonomy-Guided Benchmark for Real-World Text-to-SQL Translation
par: Wang, Hao, et autres
Publié: (2025)
par: Wang, Hao, et autres
Publié: (2025)
Unmasking Deceptive Visuals: Benchmarking Multimodal Large Language Models on Misleading Chart Question Answering
par: Chen, Zixin, et autres
Publié: (2025)
par: Chen, Zixin, et autres
Publié: (2025)
Maintaining Journalistic Integrity in the Digital Age: A Comprehensive NLP Framework for Evaluating Online News Content
par: Bojic, Ljubisa, et autres
Publié: (2024)
par: Bojic, Ljubisa, et autres
Publié: (2024)
Combining the Best of Both Worlds: A Method for Hybrid NMT and LLM Translation
par: Wu, Zhanglin, et autres
Publié: (2025)
par: Wu, Zhanglin, et autres
Publié: (2025)
Consistent and Distinctive: LLM Benchmark Efficiency via Maximum Independent Set Prompt Selection on Similarity Graphs
par: Kjorvezir, Denica, et autres
Publié: (2026)
par: Kjorvezir, Denica, et autres
Publié: (2026)
Exploring Multimodal Challenges in Toxic Chinese Detection: Taxonomy, Benchmark, and Findings
par: Yang, Shujian, et autres
Publié: (2025)
par: Yang, Shujian, et autres
Publié: (2025)
ChartCitor: Multi-Agent Framework for Fine-Grained Chart Visual Attribution
par: Goswami, Kanika, et autres
Publié: (2025)
par: Goswami, Kanika, et autres
Publié: (2025)
COMPL-AI Framework: A Technical Interpretation and LLM Benchmarking Suite for the EU Artificial Intelligence Act
par: Guldimann, Philipp, et autres
Publié: (2024)
par: Guldimann, Philipp, et autres
Publié: (2024)
LELA: An End-to-end LLM-based Entity Linking Framework with Zero-shot Domain Adaptation
par: Haffoudhi, Samy, et autres
Publié: (2026)
par: Haffoudhi, Samy, et autres
Publié: (2026)
LiveFact: A Dynamic, Time-Aware Benchmark for LLM-Driven Fake News Detection
par: Xu, Cheng, et autres
Publié: (2026)
par: Xu, Cheng, et autres
Publié: (2026)
CHARTOM: A Visual Theory-of-Mind Benchmark for LLMs on Misleading Charts
par: Bharti, Shubham, et autres
Publié: (2024)
par: Bharti, Shubham, et autres
Publié: (2024)
ChartEditBench: Evaluating Grounded Multi-Turn Chart Editing in Multimodal Language Models
par: Kapadnis, Manav Nitin, et autres
Publié: (2026)
par: Kapadnis, Manav Nitin, et autres
Publié: (2026)
LLM-Generated Negative News Headlines Dataset: Creation and Benchmarking Against Real Journalism
par: Babalola, Olusola, et autres
Publié: (2025)
par: Babalola, Olusola, et autres
Publié: (2025)
Benchmark Profiling: Mechanistic Diagnosis of LLM Benchmarks
par: Kim, Dongjun, et autres
Publié: (2025)
par: Kim, Dongjun, et autres
Publié: (2025)
RAGTurk: Best Practices for Retrieval Augmented Generation in Turkish
par: Köse, Süha Kağan, et autres
Publié: (2026)
par: Köse, Süha Kağan, et autres
Publié: (2026)
A Geometric Taxonomy of Hallucinations in LLMs
par: Marín, Javier
Publié: (2026)
par: Marín, Javier
Publié: (2026)
WikiMixQA: A Multimodal Benchmark for Question Answering over Tables and Charts
par: Foroutan, Negar, et autres
Publié: (2025)
par: Foroutan, Negar, et autres
Publié: (2025)
Who Benchmarks the Benchmarks? A Case Study of LLM Evaluation in Icelandic
par: Ingimundarson, Finnur Ágúst, et autres
Publié: (2026)
par: Ingimundarson, Finnur Ágúst, et autres
Publié: (2026)
$C^2$: Scalable Auto-Feedback for LLM-based Chart Generation
par: Koh, Woosung, et autres
Publié: (2024)
par: Koh, Woosung, et autres
Publié: (2024)
Documents similaires
-
From Polyester Girlfriends to Blind Mice: Creating the First Pragmatics Understanding Benchmarks for Slovene
par: Brglez, Mojca, et autres
Publié: (2025) -
Supercharging Agenda Setting Research: The ParlaCAP Dataset of 28 European Parliaments and a Scalable Multilingual LLM-Based Classification
par: Pungeršek, Taja Kuzman, et autres
Publié: (2026) -
State of the Art in Text Classification for South Slavic Languages: Fine-Tuning or Prompting?
par: Pungeršek, Taja Kuzman, et autres
Publié: (2025) -
LLM Teacher-Student Framework for Text Classification With No Manually Annotated Data: A Case Study in IPTC News Topic Classification
par: Kuzman, Taja, et autres
Publié: (2024) -
CLASSLA-web: Comparable Web Corpora of South Slavic Languages Enriched with Linguistic and Genre Annotation
par: Ljubešić, Nikola, et autres
Publié: (2024)