RoLegalGEC: Legal Domain Grammatical Error Detection and Correction Dataset for Romanian
Fuente:
arXiv
Saved in:
| Main Authors: | Timpuriu, Mircea, Cercel, Mihaela-Claudia, Cercel, Dumitru-Clementin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Parameter Efficient Multimodal Instruction Tuning for Romanian Vision Language Models
by: Dima, George-Andrei, et al.
Published: (2025)
by: Dima, George-Andrei, et al.
Published: (2025)
SaRoHead: Detecting Satire in a Multi-Domain Romanian News Headline Dataset
by: Vîrlan, Mihnea-Alexandru, et al.
Published: (2025)
by: Vîrlan, Mihnea-Alexandru, et al.
Published: (2025)
GRAF: Graph Retrieval Augmented by Facts for Romanian Legal Multi-Choice Question Answering
by: Crăciun, Cristian-George, et al.
Published: (2024)
by: Crăciun, Cristian-George, et al.
Published: (2024)
RoLargeSum: A Large Dialect-Aware Romanian News Dataset for Summary, Headline, and Keyword Generation
by: Avram, Andrei-Marius, et al.
Published: (2024)
by: Avram, Andrei-Marius, et al.
Published: (2024)
MoRoVoc: A Large Dataset for Geographical Variation Identification of the Spoken Romanian Language
by: Avram, Andrei-Marius, et al.
Published: (2025)
by: Avram, Andrei-Marius, et al.
Published: (2025)
MuSaRoNews: A Multidomain, Multimodal Satire Dataset from Romanian News Articles
by: Smădu, Răzvan-Alexandru, et al.
Published: (2025)
by: Smădu, Răzvan-Alexandru, et al.
Published: (2025)
RoIt-XMASA: Multi-Domain Multilingual Sentiment Analysis Dataset for Romanian and Italian
by: Avram, Andrei-Marius, et al.
Published: (2026)
by: Avram, Andrei-Marius, et al.
Published: (2026)
RoD-TAL: A Benchmark for Answering Questions in Romanian Driving License Exams
by: Man, Andrei Vlad, et al.
Published: (2025)
by: Man, Andrei Vlad, et al.
Published: (2025)
SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
by: Smădu, Răzvan-Alexandru, et al.
Published: (2025)
by: Smădu, Răzvan-Alexandru, et al.
Published: (2025)
Introducing OmniGEC: A Silver Multilingual Dataset for Grammatical Error Correction
by: Kovalchuk, Roman, et al.
Published: (2025)
by: Kovalchuk, Roman, et al.
Published: (2025)
Air Pollution Forecasting in Bucharest
by: Şerban, Dragoş-Andrei, et al.
Published: (2025)
by: Şerban, Dragoş-Andrei, et al.
Published: (2025)
RoCoISLR: A Romanian Corpus for Isolated Sign Language Recognition
by: Rîpanu, Cătălin-Alexandru, et al.
Published: (2025)
by: Rîpanu, Cătălin-Alexandru, et al.
Published: (2025)
Investigating the Impact of Semi-Supervised Methods with Data Augmentation on Offensive Language Detection in Romanian Language
by: Nicola, Elena-Beatrice, et al.
Published: (2024)
by: Nicola, Elena-Beatrice, et al.
Published: (2024)
RoQLlama: A Lightweight Romanian Adapted Language Model
by: Dima, George-Andrei, et al.
Published: (2024)
by: Dima, George-Andrei, et al.
Published: (2024)
Multimodal Learning with Augmentation Techniques for Natural Disaster Assessment
by: Urse, Adrian-Dinu, et al.
Published: (2025)
by: Urse, Adrian-Dinu, et al.
Published: (2025)
A Cross-Lingual Meta-Learning Method Based on Domain Adaptation for Speech Emotion Recognition
by: Ion, David-Gabriel, et al.
Published: (2024)
by: Ion, David-Gabriel, et al.
Published: (2024)
Investigating Large Language Models for Complex Word Identification in Multilingual and Multidomain Setups
by: Smădu, Răzvan-Alexandru, et al.
Published: (2024)
by: Smădu, Răzvan-Alexandru, et al.
Published: (2024)
Enhancing Romanian Offensive Language Detection through Knowledge Distillation, Multi-Task Learning, and Data Augmentation
by: Matei, Vlad-Cristian, et al.
Published: (2024)
by: Matei, Vlad-Cristian, et al.
Published: (2024)
Neural Grammatical Error Correction for Romanian
by: Cotet, Teodor-Mihai, et al.
Published: (2026)
by: Cotet, Teodor-Mihai, et al.
Published: (2026)
LLMCL-GEC: Advancing Grammatical Error Correction with LLM-Driven Curriculum Learning
by: Fang, Tao, et al.
Published: (2024)
by: Fang, Tao, et al.
Published: (2024)
KoGEC : Korean Grammatical Error Correction with Pre-trained Translation Models
by: Kim, Taeeun, et al.
Published: (2025)
by: Kim, Taeeun, et al.
Published: (2025)
PoPreRo: A New Dataset for Popularity Prediction of Romanian Reddit Posts
by: Rogoz, Ana-Cristina, et al.
Published: (2024)
by: Rogoz, Ana-Cristina, et al.
Published: (2024)
EUROPA: A Legal Multilingual Keyphrase Generation Dataset
by: Salaün, Olivier, et al.
Published: (2024)
by: Salaün, Olivier, et al.
Published: (2024)
CL$^2$GEC: A Multi-Discipline Benchmark for Continual Learning in Chinese Literature Grammatical Error Correction
by: Qin, Shang, et al.
Published: (2025)
by: Qin, Shang, et al.
Published: (2025)
IGAff: Benchmarking Adversarial Iterative and Genetic Affine Algorithms on Deep Neural Networks
by: Echim, Sebastian-Vasile, et al.
Published: (2025)
by: Echim, Sebastian-Vasile, et al.
Published: (2025)
Explainability-Driven Leaf Disease Classification Using Adversarial Training and Knowledge Distillation
by: Echim, Sebastian-Vasile, et al.
Published: (2023)
by: Echim, Sebastian-Vasile, et al.
Published: (2023)
ViLegalNLI: Natural Language Inference for Vietnamese Legal Texts
by: Duong, Nhung Thi-Hong, et al.
Published: (2026)
by: Duong, Nhung Thi-Hong, et al.
Published: (2026)
LegalLens: Leveraging LLMs for Legal Violation Identification in Unstructured Text
by: Bernsohn, Dor, et al.
Published: (2024)
by: Bernsohn, Dor, et al.
Published: (2024)
The Factuality of Large Language Models in the Legal Domain
by: Hamdani, Rajaa El, et al.
Published: (2024)
by: Hamdani, Rajaa El, et al.
Published: (2024)
Morpheme Boundary Detection & Grammatical Feature Prediction for Gujarati : Dataset & Model
by: Baxi, Jatayu, et al.
Published: (2021)
by: Baxi, Jatayu, et al.
Published: (2021)
NyayaMind- A Framework for Transparent Legal Reasoning and Judgment Prediction in the Indian Legal System
by: Shukla, Parjanya Aditya, et al.
Published: (2026)
by: Shukla, Parjanya Aditya, et al.
Published: (2026)
RoBiologyDataChoiceQA: A Romanian Dataset for improving Biology understanding of Large Language Models
by: Ghinea, Dragos-Dumitru, et al.
Published: (2025)
by: Ghinea, Dragos-Dumitru, et al.
Published: (2025)
HLDC: Hindi Legal Documents Corpus
by: Kapoor, Arnav, et al.
Published: (2022)
by: Kapoor, Arnav, et al.
Published: (2022)
Lawma: The Power of Specialization for Legal Annotation
by: Dominguez-Olmedo, Ricardo, et al.
Published: (2024)
by: Dominguez-Olmedo, Ricardo, et al.
Published: (2024)
A Novel Cartography-Based Curriculum Learning Method Applied on RoNLI: The First Romanian Natural Language Inference Corpus
by: Poesina, Eduard, et al.
Published: (2024)
by: Poesina, Eduard, et al.
Published: (2024)
ArbESC+: Arabic Enhanced Edit Selection System Combination for Grammatical Error Correction Resolving conflict and improving system combination in Arabic GEC
by: Alrehili, Ahlam, et al.
Published: (2025)
by: Alrehili, Ahlam, et al.
Published: (2025)
Indian Legal NLP Benchmarks : A Survey
by: Kalamkar, Prathamesh, et al.
Published: (2021)
by: Kalamkar, Prathamesh, et al.
Published: (2021)
Legal2LogicICL: Improving Generalization in Transforming Legal Cases to Logical Formulas via Diverse Few-Shot Learning
by: Xue, Jieying, et al.
Published: (2026)
by: Xue, Jieying, et al.
Published: (2026)
HalluGraph: Auditable Hallucination Detection for Legal RAG Systems via Knowledge Graph Alignment
by: Noël, Valentin, et al.
Published: (2025)
by: Noël, Valentin, et al.
Published: (2025)
Aligning Language Models for Icelandic Legal Text Summarization
by: Harðarson, Þórir Hrafn, et al.
Published: (2025)
by: Harðarson, Þórir Hrafn, et al.
Published: (2025)
Similar Items
-
Parameter Efficient Multimodal Instruction Tuning for Romanian Vision Language Models
by: Dima, George-Andrei, et al.
Published: (2025) -
SaRoHead: Detecting Satire in a Multi-Domain Romanian News Headline Dataset
by: Vîrlan, Mihnea-Alexandru, et al.
Published: (2025) -
GRAF: Graph Retrieval Augmented by Facts for Romanian Legal Multi-Choice Question Answering
by: Crăciun, Cristian-George, et al.
Published: (2024) -
RoLargeSum: A Large Dialect-Aware Romanian News Dataset for Summary, Headline, and Keyword Generation
by: Avram, Andrei-Marius, et al.
Published: (2024) -
MoRoVoc: A Large Dataset for Geographical Variation Identification of the Spoken Romanian Language
by: Avram, Andrei-Marius, et al.
Published: (2025)