Test Set Quality in Multilingual LLM Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kranti, Chalamalasetti, Bernier-Colborne, Gabriel, Gauthier, Yvan, Vajjala, Sowmya |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MetricalARGS: A Taxonomy for Studying Metrical Poetry with LLMs
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2025)
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2025)
MATA: Mindful Assessment of the Telugu Abilities of Large Language Models
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2025)
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2025)
Annotation Errors and NER: A Study with OntoNotes 5.0
von: Bernier-Colborne, Gabriel, et al.
Veröffentlicht: (2024)
von: Bernier-Colborne, Gabriel, et al.
Veröffentlicht: (2024)
Mind the Gap: Evaluating LLM Understanding of Human-Taught Road Safety Principles
von: Kranti, Chalamalasetti
Veröffentlicht: (2025)
von: Kranti, Chalamalasetti
Veröffentlicht: (2025)
IndicGEC: Powerful Models, or a Measurement Mirage?
von: Vajjala, Sowmya
Veröffentlicht: (2025)
von: Vajjala, Sowmya
Veröffentlicht: (2025)
The Problem with Safety Classification is not just the Models
von: Vajjala, Sowmya
Veröffentlicht: (2025)
von: Vajjala, Sowmya
Veröffentlicht: (2025)
clem:todd: A Framework for the Systematic Benchmarking of LLM-Based Task-Oriented Dialogue System Realisations
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2025)
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2025)
Text Classification in the LLM Era -- Where do we stand?
von: Vajjala, Sowmya, et al.
Veröffentlicht: (2025)
von: Vajjala, Sowmya, et al.
Veröffentlicht: (2025)
A Third Paradigm for LLM Evaluation: Dialogue Game-Based Evaluation using clembench
von: Schlangen, David, et al.
Veröffentlicht: (2025)
von: Schlangen, David, et al.
Veröffentlicht: (2025)
From Templates to Natural Language: Generalization Challenges in Instruction-Tuned LLMs for Spatial Reasoning
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2025)
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2025)
Retrieval-Augmented Code Generation for Situated Action Generation: A Case Study on Minecraft
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2024)
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2024)
Multi-Turn Multi-Agent Dialogue for Collaborative Reconstruction Improves VLM Performance on Spatial Reasoning, But Only Barely
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2026)
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2026)
Towards No-Code Programming of Cobots: Experiments with Code Synthesis by Large Code Models for Conversational Programming
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2024)
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2024)
Does Synthetic Data Help Named Entity Recognition for Low-Resource Languages?
von: Kamath, Gaurav, et al.
Veröffentlicht: (2025)
von: Kamath, Gaurav, et al.
Veröffentlicht: (2025)
Dravidian language family through Universal Dependencies lens
von: Rama, Taraka, et al.
Veröffentlicht: (2024)
von: Rama, Taraka, et al.
Veröffentlicht: (2024)
clembench-2024: A Challenging, Dynamic, Complementary, Multilingual Benchmark and Underlying Flexible Framework for LLMs as Multi-Action Agents
von: Beyer, Anne, et al.
Veröffentlicht: (2024)
von: Beyer, Anne, et al.
Veröffentlicht: (2024)
Scope Ambiguities in Large Language Models
von: Kamath, Gaurav, et al.
Veröffentlicht: (2024)
von: Kamath, Gaurav, et al.
Veröffentlicht: (2024)
Opportunities and Challenges of LLMs in Education: An NLP Perspective
von: Vajjala, Sowmya, et al.
Veröffentlicht: (2025)
von: Vajjala, Sowmya, et al.
Veröffentlicht: (2025)
LLMs in Education: Novel Perspectives, Challenges, and Opportunities
von: Alhafni, Bashar, et al.
Veröffentlicht: (2024)
von: Alhafni, Bashar, et al.
Veröffentlicht: (2024)
LLM Judges Inconsistently Disagree Across Safety Criteria and Harm Categories
von: Vishnubhotla, Krishnapriya, et al.
Veröffentlicht: (2026)
von: Vishnubhotla, Krishnapriya, et al.
Veröffentlicht: (2026)
On the evolution of research in hypersonics: application of natural language processing and machine learning
von: Ebadi, Ashkan, et al.
Veröffentlicht: (2022)
von: Ebadi, Ashkan, et al.
Veröffentlicht: (2022)
RDF-Based Structured Quality Assessment Representation of Multilingual LLM Evaluations
von: Gwozdz, Jonas, et al.
Veröffentlicht: (2025)
von: Gwozdz, Jonas, et al.
Veröffentlicht: (2025)
UniversalCEFR: Enabling Open Multilingual Research on Language Proficiency Assessment
von: Imperial, Joseph Marvin, et al.
Veröffentlicht: (2025)
von: Imperial, Joseph Marvin, et al.
Veröffentlicht: (2025)
Exploring the Multilingual NLG Evaluation Abilities of LLM-Based Evaluators
von: Chang, Jiayi, et al.
Veröffentlicht: (2025)
von: Chang, Jiayi, et al.
Veröffentlicht: (2025)
Quantifying the Impact of Translation Errors on Multilingual LLM Evaluation
von: Thellmann, Klaudia-Doris, et al.
Veröffentlicht: (2026)
von: Thellmann, Klaudia-Doris, et al.
Veröffentlicht: (2026)
Multilingual Blending: LLM Safety Alignment Evaluation with Language Mixture
von: Song, Jiayang, et al.
Veröffentlicht: (2024)
von: Song, Jiayang, et al.
Veröffentlicht: (2024)
HardTests: Synthesizing High-Quality Test Cases for LLM Coding
von: He, Zhongmou, et al.
Veröffentlicht: (2025)
von: He, Zhongmou, et al.
Veröffentlicht: (2025)
RoSE: Round-robin Synthetic Data Evaluation for Selecting LLM Generators without Human Test Sets
von: Cegin, Jan, et al.
Veröffentlicht: (2025)
von: Cegin, Jan, et al.
Veröffentlicht: (2025)
Found in Translation: Measuring Multilingual LLM Consistency as Simple as Translate then Evaluate
von: Gupta, Ashim, et al.
Veröffentlicht: (2025)
von: Gupta, Ashim, et al.
Veröffentlicht: (2025)
PerQ: Efficient Evaluation of Multilingual Text Personalization Quality
von: Macko, Dominik, et al.
Veröffentlicht: (2025)
von: Macko, Dominik, et al.
Veröffentlicht: (2025)
MTQ-Eval: Multilingual Text Quality Evaluation for Language Models
von: Pokharel, Rhitabrat, et al.
Veröffentlicht: (2025)
von: Pokharel, Rhitabrat, et al.
Veröffentlicht: (2025)
M-RewardBench: Evaluating Reward Models in Multilingual Settings
von: Gureja, Srishti, et al.
Veröffentlicht: (2024)
von: Gureja, Srishti, et al.
Veröffentlicht: (2024)
MultiHal: Multilingual Dataset for Knowledge-Graph Grounded Evaluation of LLM Hallucinations
von: Lavrinovics, Ernests, et al.
Veröffentlicht: (2025)
von: Lavrinovics, Ernests, et al.
Veröffentlicht: (2025)
EthioLLM: Multilingual Large Language Models for Ethiopian Languages with Task Evaluation
von: Tonja, Atnafu Lambebo, et al.
Veröffentlicht: (2024)
von: Tonja, Atnafu Lambebo, et al.
Veröffentlicht: (2024)
GlotEval: A Test Suite for Massively Multilingual Evaluation of Large Language Models
von: Luo, Hengyu, et al.
Veröffentlicht: (2025)
von: Luo, Hengyu, et al.
Veröffentlicht: (2025)
Towards Multilingual LLM Evaluation for European Languages
von: Thellmann, Klaudia, et al.
Veröffentlicht: (2024)
von: Thellmann, Klaudia, et al.
Veröffentlicht: (2024)
Scaling Crowdsourced Election Monitoring: Construction and Evaluation of Classification Models for Multilingual and Cross-Domain Classification Settings
von: Magomere, Jabez, et al.
Veröffentlicht: (2025)
von: Magomere, Jabez, et al.
Veröffentlicht: (2025)
Déjà Vu: Multilingual LLM Evaluation through the Lens of Machine Translation Evaluation
von: Kreutzer, Julia, et al.
Veröffentlicht: (2025)
von: Kreutzer, Julia, et al.
Veröffentlicht: (2025)
Are LLM-based Evaluators Confusing NLG Quality Criteria?
von: Hu, Xinyu, et al.
Veröffentlicht: (2024)
von: Hu, Xinyu, et al.
Veröffentlicht: (2024)
Africa-Centric Self-Supervised Pre-Training for Multilingual Speech Representation in a Sub-Saharan Context
von: Caubrière, Antoine, et al.
Veröffentlicht: (2024)
von: Caubrière, Antoine, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MetricalARGS: A Taxonomy for Studying Metrical Poetry with LLMs
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2025) -
MATA: Mindful Assessment of the Telugu Abilities of Large Language Models
von: Kranti, Chalamalasetti, et al.
Veröffentlicht: (2025) -
Annotation Errors and NER: A Study with OntoNotes 5.0
von: Bernier-Colborne, Gabriel, et al.
Veröffentlicht: (2024) -
Mind the Gap: Evaluating LLM Understanding of Human-Taught Road Safety Principles
von: Kranti, Chalamalasetti
Veröffentlicht: (2025) -
IndicGEC: Powerful Models, or a Measurement Mirage?
von: Vajjala, Sowmya
Veröffentlicht: (2025)