Beyond statistical significance: Quantifying uncertainty and statistical variability in multilingual and multitask NLP evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Sälevä, Jonne, Ataman, Duygu, Lignos, Constantine |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Effectiveness of Morphology-aware Segmentation in Low-Resource Neural Machine Translation
by: Sälevä, Jonne, et al.
Published: (2021)
by: Sälevä, Jonne, et al.
Published: (2021)
ParaNames 1.0: Creating an Entity Name Corpus for 400+ Languages using Wikidata
by: Sälevä, Jonne, et al.
Published: (2024)
by: Sälevä, Jonne, et al.
Published: (2024)
OpenNER 1.0: Standardized Open-Access Named Entity Recognition Datasets in 50+ Languages
by: Palen-Michel, Chester, et al.
Published: (2024)
by: Palen-Michel, Chester, et al.
Published: (2024)
Evaluating Morphological Compositional Generalization in Large Language Models
by: Ismayilzada, Mete, et al.
Published: (2024)
by: Ismayilzada, Mete, et al.
Published: (2024)
Comparing Approaches to Automatic Summarization in Less-Resourced Languages
by: Palen-Michel, Chester, et al.
Published: (2025)
by: Palen-Michel, Chester, et al.
Published: (2025)
The Ouroboros of Benchmarking: Reasoning Evaluation in an Era of Saturation
by: Deveci, İbrahim Ethem, et al.
Published: (2025)
by: Deveci, İbrahim Ethem, et al.
Published: (2025)
Generalization Measures for Zero-Shot Cross-Lingual Transfer
by: Bassi, Saksham, et al.
Published: (2024)
by: Bassi, Saksham, et al.
Published: (2024)
CoNLL#: Fine-grained Error Analysis and a Corrected Test Set for CoNLL-03 English
by: Rueda, Andrew, et al.
Published: (2024)
by: Rueda, Andrew, et al.
Published: (2024)
Halluverse-M^3: A multitask multilingual benchmark for hallucination in LLMs
by: Abdaljalil, Samir, et al.
Published: (2026)
by: Abdaljalil, Samir, et al.
Published: (2026)
Overview of ADoBo at IberLEF 2025: Automatic Detection of Anglicisms in Spanish
by: Alvarez-Mellado, Elena, et al.
Published: (2025)
by: Alvarez-Mellado, Elena, et al.
Published: (2025)
QueryNER: Segmentation of E-commerce Queries
by: Palen-Michel, Chester, et al.
Published: (2024)
by: Palen-Michel, Chester, et al.
Published: (2024)
Open foundation models for Azerbaijani language
by: Isbarov, Jafar, et al.
Published: (2024)
by: Isbarov, Jafar, et al.
Published: (2024)
D-NLP at SemEval-2024 Task 2: Evaluating Clinical Inference Capabilities of Large Language Models
by: Altinok, Duygu
Published: (2024)
by: Altinok, Duygu
Published: (2024)
A statistically consistent measure of semantic uncertainty using Language Models
by: Liu, Yi
Published: (2025)
by: Liu, Yi
Published: (2025)
CMMLU: Measuring massive multitask language understanding in Chinese
by: Li, Haonan, et al.
Published: (2023)
by: Li, Haonan, et al.
Published: (2023)
Collaboration or Corporate Capture? Quantifying NLP's Reliance on Industry Artifacts and Contributions
by: Aitken, Will, et al.
Published: (2023)
by: Aitken, Will, et al.
Published: (2023)
A multitask learning framework for leveraging subjectivity of annotators to identify misogyny
by: Angel, Jason, et al.
Published: (2024)
by: Angel, Jason, et al.
Published: (2024)
A multitask transformer to sign language translation using motion gesture primitives
by: López, Fredy Alejandro Mendoza, et al.
Published: (2025)
by: López, Fredy Alejandro Mendoza, et al.
Published: (2025)
Beyond Consensus: Perspectivist Modeling and Evaluation of Annotator Disagreement in NLP
by: Xu, Yinuo, et al.
Published: (2026)
by: Xu, Yinuo, et al.
Published: (2026)
The statistical advantage of automatic NLG metrics at the system level
by: Wei, Johnny Tian-Zheng, et al.
Published: (2021)
by: Wei, Johnny Tian-Zheng, et al.
Published: (2021)
Beyond Hate Speech: NLP's Challenges and Opportunities in Uncovering Dehumanizing Language
by: Saffari, Hamidreza, et al.
Published: (2024)
by: Saffari, Hamidreza, et al.
Published: (2024)
Beyond speculation: Measuring the growing presence of LLM-generated texts in multilingual disinformation
by: Macko, Dominik, et al.
Published: (2025)
by: Macko, Dominik, et al.
Published: (2025)
Can LLMs reason over extended multilingual contexts? Towards long-context evaluation beyond retrieval and haystacks
by: Hengle, Amey, et al.
Published: (2025)
by: Hengle, Amey, et al.
Published: (2025)
WorldMedQA-V: a multilingual, multimodal medical examination dataset for multimodal language models evaluation
by: Matos, João, et al.
Published: (2024)
by: Matos, João, et al.
Published: (2024)
Universal statistical laws governing culinary design
by: Bagler, Ganesh, et al.
Published: (2026)
by: Bagler, Ganesh, et al.
Published: (2026)
TurkicNLP: An NLP Toolkit for Turkic Languages
by: Hakimov, Sherzod
Published: (2026)
by: Hakimov, Sherzod
Published: (2026)
The Nature of NLP: Analyzing Contributions in NLP Papers
by: Pramanick, Aniket, et al.
Published: (2024)
by: Pramanick, Aniket, et al.
Published: (2024)
Hallucinations are inevitable but can be made statistically negligible
by: Suzuki, Atsushi, et al.
Published: (2025)
by: Suzuki, Atsushi, et al.
Published: (2025)
Batayan: A Filipino NLP benchmark for evaluating Large Language Models
by: Montalan, Jann Railey, et al.
Published: (2025)
by: Montalan, Jann Railey, et al.
Published: (2025)
A study of Vietnamese readability assessing through semantic and statistical features
by: Le, Hung Tuan, et al.
Published: (2024)
by: Le, Hung Tuan, et al.
Published: (2024)
Morphosyntactic probing of multilingual BERT models
by: Acs, Judit, et al.
Published: (2023)
by: Acs, Judit, et al.
Published: (2023)
Beyond Performance: Quantifying and Mitigating Label Bias in LLMs
by: Reif, Yuval, et al.
Published: (2024)
by: Reif, Yuval, et al.
Published: (2024)
Language statistics at different spatial, temporal, and grammatical scales
by: Sánchez-Puig, Fernanda, et al.
Published: (2022)
by: Sánchez-Puig, Fernanda, et al.
Published: (2022)
Beyond Majority Voting: Agreement-Based Clustering to Model Annotator Perspectives in Subjective NLP Tasks
by: Belay, Tadesse Destaw, et al.
Published: (2026)
by: Belay, Tadesse Destaw, et al.
Published: (2026)
Continuous sentiment scores for literary and multilingual contexts
by: Lyngbaek, Laurits, et al.
Published: (2025)
by: Lyngbaek, Laurits, et al.
Published: (2025)
Differential contributions of machine learning and statistical analysis to language and cognitive sciences
by: Sun, Kun, et al.
Published: (2024)
by: Sun, Kun, et al.
Published: (2024)
Introducing TrGLUE and SentiTurca: A Comprehensive Benchmark for Turkish General Language Understanding and Sentiment Analysis
by: Altinok, Duygu
Published: (2025)
by: Altinok, Duygu
Published: (2025)
Whispering Context: Distilling Syntax and Semantics for Long Speech Transcripts
by: Altinok, Duygu
Published: (2025)
by: Altinok, Duygu
Published: (2025)
Optimal Turkish Subword Strategies at Scale: Systematic Evaluation of Data, Vocabulary, Morphology Interplay
by: Altinok, Duygu
Published: (2026)
by: Altinok, Duygu
Published: (2026)
Probing the statistical properties of enriched co-occurrence networks
by: Amancio, Diego R., et al.
Published: (2024)
by: Amancio, Diego R., et al.
Published: (2024)
Similar Items
-
The Effectiveness of Morphology-aware Segmentation in Low-Resource Neural Machine Translation
by: Sälevä, Jonne, et al.
Published: (2021) -
ParaNames 1.0: Creating an Entity Name Corpus for 400+ Languages using Wikidata
by: Sälevä, Jonne, et al.
Published: (2024) -
OpenNER 1.0: Standardized Open-Access Named Entity Recognition Datasets in 50+ Languages
by: Palen-Michel, Chester, et al.
Published: (2024) -
Evaluating Morphological Compositional Generalization in Large Language Models
by: Ismayilzada, Mete, et al.
Published: (2024) -
Comparing Approaches to Automatic Summarization in Less-Resourced Languages
by: Palen-Michel, Chester, et al.
Published: (2025)