Y-NQ: English-Yorùbá Evaluation dataset for Open-Book Reading Comprehension and Text Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Costa-jussà, Marta R., Chen, Joy, Adebara, Ifeoluwanimi, Chuang, Joe, Ropers, Christophe, Sánchez, Eduardo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BOUQuET: dataset, Benchmark and Open initiative for Universal Quality Evaluation in Translation
by: The Omnilingual MT Team, et al.
Published: (2025)
by: The Omnilingual MT Team, et al.
Published: (2025)
2M-BELEBELE: Highly Multilingual Speech and American Sign Language Comprehension Dataset
by: Costa-jussà, Marta R., et al.
Published: (2024)
by: Costa-jussà, Marta R., et al.
Published: (2024)
LCFO: Long Context and Long Form Output Dataset and Benchmarking
by: Costa-jussà, Marta R., et al.
Published: (2024)
by: Costa-jussà, Marta R., et al.
Published: (2024)
Towards Massive Multilingual Holistic Bias
by: Tan, Xiaoqing Ellen, et al.
Published: (2024)
by: Tan, Xiaoqing Ellen, et al.
Published: (2024)
MuTox: Universal MUltilingual Audio-based TOXicity Dataset and Zero-shot Detector
by: Costa-jussà, Marta R., et al.
Published: (2024)
by: Costa-jussà, Marta R., et al.
Published: (2024)
Towards Red Teaming in Multimodal and Multilingual Translation
by: Ropers, Christophe, et al.
Published: (2024)
by: Ropers, Christophe, et al.
Published: (2024)
Comprehensiveness Metrics for Automatic Evaluation of Factual Recall in Text Generation
by: Dejl, Adam, et al.
Published: (2025)
by: Dejl, Adam, et al.
Published: (2025)
Omnilingual MT: Machine Translation for 1,600 Languages
by: Omnilingual MT Team, et al.
Published: (2026)
by: Omnilingual MT Team, et al.
Published: (2026)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
by: Oketunji, Abiodun Finbarrs
Published: (2023)
by: Oketunji, Abiodun Finbarrs
Published: (2023)
RTI-Bench: A Structured Dataset for Indian Right-to-Information Decision Analysis
by: Bose, Joy
Published: (2026)
by: Bose, Joy
Published: (2026)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
by: Peters, Sydney, et al.
Published: (2025)
by: Peters, Sydney, et al.
Published: (2025)
Rubrik's Cube: Testing a New Rubric for Evaluating Explanations on the CUBE dataset
by: Galvan-Sosa, Diana, et al.
Published: (2025)
by: Galvan-Sosa, Diana, et al.
Published: (2025)
Unstructured Text Enhanced Open-domain Dialogue System: A Systematic Survey
by: Ma, Longxuan, et al.
Published: (2024)
by: Ma, Longxuan, et al.
Published: (2024)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
by: Ashuach, Tomer, et al.
Published: (2025)
by: Ashuach, Tomer, et al.
Published: (2025)
Toward Architecture-Aware Evaluation Metrics for LLM Agents
by: Souza, Débora, et al.
Published: (2026)
by: Souza, Débora, et al.
Published: (2026)
Hidden Failures in Robustness: Why Supervised Uncertainty Quantification Needs Better Evaluation
by: Stacey, Joe, et al.
Published: (2026)
by: Stacey, Joe, et al.
Published: (2026)
Beyond Rating: A Comprehensive Evaluation and Benchmark for AI Reviews
by: Li, Bowen, et al.
Published: (2026)
by: Li, Bowen, et al.
Published: (2026)
RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors
by: Dugan, Liam, et al.
Published: (2024)
by: Dugan, Liam, et al.
Published: (2024)
TrustAI at SemEval-2024 Task 8: A Comprehensive Analysis of Multi-domain Machine Generated Text Detection Techniques
by: Urlana, Ashok, et al.
Published: (2024)
by: Urlana, Ashok, et al.
Published: (2024)
OpenFactCheck: Building, Benchmarking Customized Fact-Checking Systems and Evaluating the Factuality of Claims and LLMs
by: Wang, Yuxia, et al.
Published: (2024)
by: Wang, Yuxia, et al.
Published: (2024)
Large Language Models for Persian $ \leftrightarrow $ English Idiom Translation
by: Rezaeimanesh, Sara, et al.
Published: (2024)
by: Rezaeimanesh, Sara, et al.
Published: (2024)
Qomhra: A Bilingual Irish and English Large Language Model
by: McInerney, Joseph, et al.
Published: (2025)
by: McInerney, Joseph, et al.
Published: (2025)
LLMBridge: An LLM Pipeline for End-to-end Referential Bridging Resolution in English
by: Levine, Lauren, et al.
Published: (2026)
by: Levine, Lauren, et al.
Published: (2026)
Unifying the Scope of Bridging Anaphora Types in English: Bridging Annotations in ARRAU and GUM
by: Levine, Lauren, et al.
Published: (2024)
by: Levine, Lauren, et al.
Published: (2024)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
by: Saji, Alan, et al.
Published: (2025)
by: Saji, Alan, et al.
Published: (2025)
Robustness of Large Language Models to Perturbations in Text
by: Singh, Ayush, et al.
Published: (2024)
by: Singh, Ayush, et al.
Published: (2024)
Are Non-English Papers Reviewed Fairly? Language-of-Study Bias in NLP Peer Reviews
by: Barkhordar, Ehsan, et al.
Published: (2026)
by: Barkhordar, Ehsan, et al.
Published: (2026)
Text Summarization With Graph Attention Networks
by: Ardestani, Mohammadreza, et al.
Published: (2026)
by: Ardestani, Mohammadreza, et al.
Published: (2026)
Evaluating GenAI for Simplifying Texts for Education: Improving Accuracy and Consistency for Enhanced Readability
by: Day, Stephanie L., et al.
Published: (2025)
by: Day, Stephanie L., et al.
Published: (2025)
Investigating the Impact of Text Summarization on Topic Modeling
by: Khandelwal, Trishia
Published: (2024)
by: Khandelwal, Trishia
Published: (2024)
Active Few-Shot Learning for Text Classification
by: Ahmadnia, Saeed, et al.
Published: (2025)
by: Ahmadnia, Saeed, et al.
Published: (2025)
Normalization of Lithuanian Text Using Regular Expressions
by: Kasparaitis, Pijus
Published: (2023)
by: Kasparaitis, Pijus
Published: (2023)
PARAPHRASUS : A Comprehensive Benchmark for Evaluating Paraphrase Detection Models
by: Michail, Andrianos, et al.
Published: (2024)
by: Michail, Andrianos, et al.
Published: (2024)
Recent Trends in Linear Text Segmentation: a Survey
by: Ghinassi, Iacopo, et al.
Published: (2024)
by: Ghinassi, Iacopo, et al.
Published: (2024)
Spotlights and Blindspots: Evaluating Machine-Generated Text Detection
by: Stowe, Kevin, et al.
Published: (2026)
by: Stowe, Kevin, et al.
Published: (2026)
Atomic Inference for NLI with Generated Facts as Atoms
by: Stacey, Joe, et al.
Published: (2023)
by: Stacey, Joe, et al.
Published: (2023)
VertAttack: Taking advantage of Text Classifiers' horizontal vision
by: Rusert, Jonathan
Published: (2024)
by: Rusert, Jonathan
Published: (2024)
UM_FHS at the CLEF 2025 SimpleText Track: Comparing No-Context and Fine-Tune Approaches for GPT-4.1 Models in Sentence and Document-Level Text Simplification
by: Kocbek, Primoz, et al.
Published: (2025)
by: Kocbek, Primoz, et al.
Published: (2025)
Improving the OOD Performance of Closed-Source LLMs on NLI Through Strategic Data Selection
by: Stacey, Joe, et al.
Published: (2025)
by: Stacey, Joe, et al.
Published: (2025)
Extracting Structured Insights from Financial News: An Augmented LLM Driven Approach
by: Dolphin, Rian, et al.
Published: (2024)
by: Dolphin, Rian, et al.
Published: (2024)
Similar Items
-
BOUQuET: dataset, Benchmark and Open initiative for Universal Quality Evaluation in Translation
by: The Omnilingual MT Team, et al.
Published: (2025) -
2M-BELEBELE: Highly Multilingual Speech and American Sign Language Comprehension Dataset
by: Costa-jussà, Marta R., et al.
Published: (2024) -
LCFO: Long Context and Long Form Output Dataset and Benchmarking
by: Costa-jussà, Marta R., et al.
Published: (2024) -
Towards Massive Multilingual Holistic Bias
by: Tan, Xiaoqing Ellen, et al.
Published: (2024) -
MuTox: Universal MUltilingual Audio-based TOXicity Dataset and Zero-shot Detector
by: Costa-jussà, Marta R., et al.
Published: (2024)