Introducing OmniGEC: A Silver Multilingual Dataset for Grammatical Error Correction
Fuente:
arXiv
Saved in:
| Main Authors: | Kovalchuk, Roman, Romanyshyn, Mariana, Ivaniuk, Petro |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RoLegalGEC: Legal Domain Grammatical Error Detection and Correction Dataset for Romanian
by: Timpuriu, Mircea, et al.
Published: (2026)
by: Timpuriu, Mircea, et al.
Published: (2026)
KoGEC : Korean Grammatical Error Correction with Pre-trained Translation Models
by: Kim, Taeeun, et al.
Published: (2025)
by: Kim, Taeeun, et al.
Published: (2025)
LLMCL-GEC: Advancing Grammatical Error Correction with LLM-Driven Curriculum Learning
by: Fang, Tao, et al.
Published: (2024)
by: Fang, Tao, et al.
Published: (2024)
CL$^2$GEC: A Multi-Discipline Benchmark for Continual Learning in Chinese Literature Grammatical Error Correction
by: Qin, Shang, et al.
Published: (2025)
by: Qin, Shang, et al.
Published: (2025)
GPT-3.5 for Grammatical Error Correction
by: Katinskaia, Anisia, et al.
Published: (2024)
by: Katinskaia, Anisia, et al.
Published: (2024)
ArbESC+: Arabic Enhanced Edit Selection System Combination for Grammatical Error Correction Resolving conflict and improving system combination in Arabic GEC
by: Alrehili, Ahlam, et al.
Published: (2025)
by: Alrehili, Ahlam, et al.
Published: (2025)
Morpheme Boundary Detection & Grammatical Feature Prediction for Gujarati : Dataset & Model
by: Baxi, Jatayu, et al.
Published: (2021)
by: Baxi, Jatayu, et al.
Published: (2021)
Tagengo: A Multilingual Chat Dataset
by: Devine, Peter
Published: (2024)
by: Devine, Peter
Published: (2024)
EUROPA: A Legal Multilingual Keyphrase Generation Dataset
by: Salaün, Olivier, et al.
Published: (2024)
by: Salaün, Olivier, et al.
Published: (2024)
Neural Grammatical Error Correction for Romanian
by: Cotet, Teodor-Mihai, et al.
Published: (2026)
by: Cotet, Teodor-Mihai, et al.
Published: (2026)
Introducing L2M3, A Multilingual Medical Large Language Model to Advance Health Equity in Low-Resource Regions
by: Gangavarapu, Agasthya
Published: (2024)
by: Gangavarapu, Agasthya
Published: (2024)
Adapting LLMs for Minimal-edit Grammatical Error Correction
by: Staruch, Ryszard, et al.
Published: (2025)
by: Staruch, Ryszard, et al.
Published: (2025)
Improving Multilingual Instruction Finetuning via Linguistically Natural and Diverse Datasets
by: Indurthi, Sathish Reddy, et al.
Published: (2024)
by: Indurthi, Sathish Reddy, et al.
Published: (2024)
MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes
by: Abacha, Asma Ben, et al.
Published: (2024)
by: Abacha, Asma Ben, et al.
Published: (2024)
Corrections Meet Explanations: A Unified Framework for Explainable Grammatical Error Correction
by: Ye, Jingheng, et al.
Published: (2025)
by: Ye, Jingheng, et al.
Published: (2025)
MedRECT: A Medical Reasoning Benchmark for Error Correction in Clinical Texts
by: Iwase, Naoto, et al.
Published: (2025)
by: Iwase, Naoto, et al.
Published: (2025)
MedAidDialog: A Multilingual Multi-Turn Medical Dialogue Dataset for Accessible Healthcare
by: Nigam, Shubham Kumar, et al.
Published: (2026)
by: Nigam, Shubham Kumar, et al.
Published: (2026)
Improving Grammatical Error Correction via Contextual Data Augmentation
by: Wang, Yixuan, et al.
Published: (2024)
by: Wang, Yixuan, et al.
Published: (2024)
Loss-Aware Curriculum Learning for Chinese Grammatical Error Correction
by: Zhang, Ding, et al.
Published: (2024)
by: Zhang, Ding, et al.
Published: (2024)
Evaluating Mathematical Reasoning of Large Language Models: A Focus on Error Identification and Correction
by: Li, Xiaoyuan, et al.
Published: (2024)
by: Li, Xiaoyuan, et al.
Published: (2024)
CKnowEdit: A New Chinese Knowledge Editing Dataset for Linguistics, Facts, and Logic Error Correction in LLMs
by: Fang, Jizhan, et al.
Published: (2024)
by: Fang, Jizhan, et al.
Published: (2024)
Introducing MeMo: A Multimodal Dataset for Memory Modelling in Multiparty Conversations
by: Tsfasman, Maria, et al.
Published: (2024)
by: Tsfasman, Maria, et al.
Published: (2024)
mCSQA: Multilingual Commonsense Reasoning Dataset with Unified Creation Strategy by Language Models and Humans
by: Sakai, Yusuke, et al.
Published: (2024)
by: Sakai, Yusuke, et al.
Published: (2024)
Synthesizing and Adapting Error Correction Data for Mobile Large Language Model Applications
by: Zhang, Yanxiang, et al.
Published: (2025)
by: Zhang, Yanxiang, et al.
Published: (2025)
COLA-GEC: A Bidirectional Framework for Enhancing Grammatical Acceptability and Error Correction
by: Yang, Xiangyu, et al.
Published: (2025)
by: Yang, Xiangyu, et al.
Published: (2025)
AR-Omni: A Unified Autoregressive Model for Any-to-Any Generation
by: Cheng, Dongjie, et al.
Published: (2026)
by: Cheng, Dongjie, et al.
Published: (2026)
Harnessing Rule-Based Reinforcement Learning for Enhanced Grammatical Error Correction
by: Li, Yilin, et al.
Published: (2025)
by: Li, Yilin, et al.
Published: (2025)
Organic Data-Driven Approach for Turkish Grammatical Error Correction and LLMs
by: Ersoy, Asım, et al.
Published: (2024)
by: Ersoy, Asım, et al.
Published: (2024)
Multilingual Instruction Tuning With Just a Pinch of Multilinguality
by: Shaham, Uri, et al.
Published: (2024)
by: Shaham, Uri, et al.
Published: (2024)
Grammatical Error Correction for Low-Resource Languages: The Case of Zarma
by: Keita, Mamadou K., et al.
Published: (2024)
by: Keita, Mamadou K., et al.
Published: (2024)
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities
by: Jiang, Shixin, et al.
Published: (2024)
by: Jiang, Shixin, et al.
Published: (2024)
IMPARA-GED: Grammatical Error Detection is Boosting Reference-free Grammatical Error Quality Estimator
by: Sakai, Yusuke, et al.
Published: (2025)
by: Sakai, Yusuke, et al.
Published: (2025)
OmniStruct: Universal Text-to-Structure Generation across Diverse Schemas
by: Huang, James Y., et al.
Published: (2025)
by: Huang, James Y., et al.
Published: (2025)
Leveraging What's Overfixed: Post-Correction via LLM Grammatical Error Overcorrection
by: Park, Taehee, et al.
Published: (2025)
by: Park, Taehee, et al.
Published: (2025)
PromptMind Team at MEDIQA-CORR 2024: Improving Clinical Text Correction with Error Categorization and LLM Ensembles
by: Gundabathula, Satya Kesav, et al.
Published: (2024)
by: Gundabathula, Satya Kesav, et al.
Published: (2024)
GECTurk WEB: An Explainable Online Platform for Turkish Grammatical Error Detection and Correction
by: Gebeşçe, Ali, et al.
Published: (2024)
by: Gebeşçe, Ali, et al.
Published: (2024)
Align Once, Benefit Multilingually: Enforcing Multilingual Consistency for LLM Safety Alignment
by: Bu, Yuyan, et al.
Published: (2026)
by: Bu, Yuyan, et al.
Published: (2026)
Enhancing Grammatical Error Detection using BERT with Cleaned Lang-8 Dataset
by: Nihalani, Rahul, et al.
Published: (2024)
by: Nihalani, Rahul, et al.
Published: (2024)
OmniPred: Language Models as Universal Regressors
by: Song, Xingyou, et al.
Published: (2024)
by: Song, Xingyou, et al.
Published: (2024)
HyperCLOVA X 8B Omni
by: NAVER Cloud HyperCLOVA X Team
Published: (2026)
by: NAVER Cloud HyperCLOVA X Team
Published: (2026)
Similar Items
-
RoLegalGEC: Legal Domain Grammatical Error Detection and Correction Dataset for Romanian
by: Timpuriu, Mircea, et al.
Published: (2026) -
KoGEC : Korean Grammatical Error Correction with Pre-trained Translation Models
by: Kim, Taeeun, et al.
Published: (2025) -
LLMCL-GEC: Advancing Grammatical Error Correction with LLM-Driven Curriculum Learning
by: Fang, Tao, et al.
Published: (2024) -
CL$^2$GEC: A Multi-Discipline Benchmark for Continual Learning in Chinese Literature Grammatical Error Correction
by: Qin, Shang, et al.
Published: (2025) -
GPT-3.5 for Grammatical Error Correction
by: Katinskaia, Anisia, et al.
Published: (2024)