Tgea: An error-annotated dataset and benchmark tasks for text generation from pretrained language models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | He, Jie, Peng, Bo, Liao, Yi, Liu, Qun, Xiong, Deyi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Box is in the Pen: Evaluating Commonsense Reasoning in Neural Machine Translation
von: He, Jie, et al.
Veröffentlicht: (2025)
von: He, Jie, et al.
Veröffentlicht: (2025)
Large language models struggle with ethnographic text annotation
von: Goodall, Leonardo S., et al.
Veröffentlicht: (2026)
von: Goodall, Leonardo S., et al.
Veröffentlicht: (2026)
ks-lit-3m: A 3.1 million word kashmiri text dataset for large language model pretraining
von: Malik, Haq Nawaz
Veröffentlicht: (2026)
von: Malik, Haq Nawaz
Veröffentlicht: (2026)
GLAP: General contrastive audio-text pretraining across domains and languages
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2025)
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2025)
LFED: A Literary Fiction Evaluation Dataset for Large Language Models
von: Yu, Linhao, et al.
Veröffentlicht: (2024)
von: Yu, Linhao, et al.
Veröffentlicht: (2024)
Meta-RTL: Reinforcement-Based Meta-Transfer Learning for Low-Resource Commonsense Reasoning
von: Fu, Yu, et al.
Veröffentlicht: (2024)
von: Fu, Yu, et al.
Veröffentlicht: (2024)
A benchmark dataset for evaluating Syndrome Differentiation and Treatment in large language models
von: Li, Kunning, et al.
Veröffentlicht: (2025)
von: Li, Kunning, et al.
Veröffentlicht: (2025)
Are generative AI text annotations systematically biased?
von: Stolwijk, Sjoerd B., et al.
Veröffentlicht: (2025)
von: Stolwijk, Sjoerd B., et al.
Veröffentlicht: (2025)
Empirical study of pretrained multilingual language models for zero-shot cross-lingual knowledge transfer in generation
von: Chirkova, Nadezhda, et al.
Veröffentlicht: (2023)
von: Chirkova, Nadezhda, et al.
Veröffentlicht: (2023)
Evaluating the fairness of task-adaptive pretraining on unlabeled test data before few-shot text classification
von: Dubey, Kush
Veröffentlicht: (2024)
von: Dubey, Kush
Veröffentlicht: (2024)
Evaluating Discourse Cohesion in Pre-trained Language Models
von: He, Jie, et al.
Veröffentlicht: (2025)
von: He, Jie, et al.
Veröffentlicht: (2025)
A large-scale image-text dataset benchmark for farmland segmentation
von: Tao, Chao, et al.
Veröffentlicht: (2025)
von: Tao, Chao, et al.
Veröffentlicht: (2025)
A thorough benchmark of automatic text classification: From traditional approaches to large language models
von: Cunha, Washington, et al.
Veröffentlicht: (2025)
von: Cunha, Washington, et al.
Veröffentlicht: (2025)
ProText: A benchmark dataset for measuring (mis)gendering in long-form texts
von: Kotek, Hadas, et al.
Veröffentlicht: (2026)
von: Kotek, Hadas, et al.
Veröffentlicht: (2026)
Machine-generated text detection prevents language model collapse
von: Drayson, George, et al.
Veröffentlicht: (2025)
von: Drayson, George, et al.
Veröffentlicht: (2025)
TAGLAS: An atlas of text-attributed graph datasets in the era of large graph and language models
von: Feng, Jiarui, et al.
Veröffentlicht: (2024)
von: Feng, Jiarui, et al.
Veröffentlicht: (2024)
Natural language guidance of high-fidelity text-to-speech with synthetic annotations
von: Lyth, Dan, et al.
Veröffentlicht: (2024)
von: Lyth, Dan, et al.
Veröffentlicht: (2024)
A dataset and benchmark for hospital course summarization with adapted large language models
von: Aali, Asad, et al.
Veröffentlicht: (2024)
von: Aali, Asad, et al.
Veröffentlicht: (2024)
Prompting open-source and commercial language models for grammatical error correction of English learner text
von: Davis, Christopher, et al.
Veröffentlicht: (2024)
von: Davis, Christopher, et al.
Veröffentlicht: (2024)
When marine radar target detection meets pretrained large language models
von: Hu, Qiying, et al.
Veröffentlicht: (2025)
von: Hu, Qiying, et al.
Veröffentlicht: (2025)
AlleNoise: large-scale text classification benchmark dataset with real-world label noise
von: Rączkowska, Alicja, et al.
Veröffentlicht: (2024)
von: Rączkowska, Alicja, et al.
Veröffentlicht: (2024)
BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs
von: Zhang, Sheng, et al.
Veröffentlicht: (2023)
von: Zhang, Sheng, et al.
Veröffentlicht: (2023)
Designing large language model prompts to extract scores from messy text: A shared dataset and challenge
von: Thelwall, Mike
Veröffentlicht: (2026)
von: Thelwall, Mike
Veröffentlicht: (2026)
Taec: a Manually annotated text dataset for trait and phenotype extraction and entity linking in wheat breeding literature
von: Nédellec, Claire, et al.
Veröffentlicht: (2024)
von: Nédellec, Claire, et al.
Veröffentlicht: (2024)
Analyzing the relationships between pretraining language, phonetic, tonal, and speaker information in self-supervised speech models
von: Gubian, Michele, et al.
Veröffentlicht: (2025)
von: Gubian, Michele, et al.
Veröffentlicht: (2025)
Differentially-private text generation degrades output language quality
von: Çano, Erion, et al.
Veröffentlicht: (2025)
von: Çano, Erion, et al.
Veröffentlicht: (2025)
CRiskEval: A Chinese Multi-Level Risk Evaluation Benchmark Dataset for Large Language Models
von: Shi, Ling, et al.
Veröffentlicht: (2024)
von: Shi, Ling, et al.
Veröffentlicht: (2024)
LANDeRMT: Detecting and Routing Language-Aware Neurons for Selectively Finetuning LLMs to Machine Translation
von: Zhu, Shaolin, et al.
Veröffentlicht: (2024)
von: Zhu, Shaolin, et al.
Veröffentlicht: (2024)
GRASP: A novel benchmark for evaluating language GRounding And Situated Physics understanding in multimodal language models
von: Jassim, Serwan, et al.
Veröffentlicht: (2023)
von: Jassim, Serwan, et al.
Veröffentlicht: (2023)
MedicalBERT: enhancing biomedical natural language processing using pretrained BERT-based model
von: Reddy, K. Sahit, et al.
Veröffentlicht: (2025)
von: Reddy, K. Sahit, et al.
Veröffentlicht: (2025)
Large corpora and large language models: a replicable method for automating grammatical annotation
von: Morin, Cameron, et al.
Veröffentlicht: (2024)
von: Morin, Cameron, et al.
Veröffentlicht: (2024)
An Empirical Study on the Robustness of Massively Multilingual Neural Machine Translation
von: Supryadi, et al.
Veröffentlicht: (2024)
von: Supryadi, et al.
Veröffentlicht: (2024)
AIDBench: A benchmark for evaluating the authorship identification capability of large language models
von: Wen, Zichen, et al.
Veröffentlicht: (2024)
von: Wen, Zichen, et al.
Veröffentlicht: (2024)
Synthetic bootstrapped pretraining
von: Yang, Zitong, et al.
Veröffentlicht: (2025)
von: Yang, Zitong, et al.
Veröffentlicht: (2025)
Synthetically generated text for supervised text analysis
von: Halterman, Andrew
Veröffentlicht: (2023)
von: Halterman, Andrew
Veröffentlicht: (2023)
Towards Understanding Multi-Task Learning (Generalization) of LLMs via Detecting and Exploring Task-Specific Neurons
von: Leng, Yongqi, et al.
Veröffentlicht: (2024)
von: Leng, Yongqi, et al.
Veröffentlicht: (2024)
N2C2: Nearest Neighbor Enhanced Confidence Calibration for Cross-Lingual In-Context Learning
von: He, Jie, et al.
Veröffentlicht: (2025)
von: He, Jie, et al.
Veröffentlicht: (2025)
The language of time: a language model perspective on time-series foundation models
von: Xie, Yi, et al.
Veröffentlicht: (2025)
von: Xie, Yi, et al.
Veröffentlicht: (2025)
FuxiMT: Sparsifying Large Language Models for Chinese-Centric Multilingual Machine Translation
von: Zhu, Shaolin, et al.
Veröffentlicht: (2025)
von: Zhu, Shaolin, et al.
Veröffentlicht: (2025)
Mix-MoE: Improving Multilingual Machine Translation of Large Language Models through Mixed MoEs
von: Li, Bo, et al.
Veröffentlicht: (2026)
von: Li, Bo, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
The Box is in the Pen: Evaluating Commonsense Reasoning in Neural Machine Translation
von: He, Jie, et al.
Veröffentlicht: (2025) -
Large language models struggle with ethnographic text annotation
von: Goodall, Leonardo S., et al.
Veröffentlicht: (2026) -
ks-lit-3m: A 3.1 million word kashmiri text dataset for large language model pretraining
von: Malik, Haq Nawaz
Veröffentlicht: (2026) -
GLAP: General contrastive audio-text pretraining across domains and languages
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2025) -
LFED: A Literary Fiction Evaluation Dataset for Large Language Models
von: Yu, Linhao, et al.
Veröffentlicht: (2024)