EUROPA: A Legal Multilingual Keyphrase Generation Dataset
Fuente:
arXiv
Saved in:
| Main Authors: | Salaün, Olivier, Piedboeuf, Frédéric, Berre, Guillaume Le, Hermelo, David Alfonso, Langlais, Philippe |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On Evaluation Protocols for Data Augmentation in a Limited Data Scenario
by: Piedboeuf, Frédéric, et al.
Published: (2024)
by: Piedboeuf, Frédéric, et al.
Published: (2024)
Tagengo: A Multilingual Chat Dataset
by: Devine, Peter
Published: (2024)
by: Devine, Peter
Published: (2024)
Evaluating the Limits of Large Language Models in Multilingual Legal Reasoning
by: Ioannou, Antreas, et al.
Published: (2025)
by: Ioannou, Antreas, et al.
Published: (2025)
LongKey: Keyphrase Extraction for Long Documents
by: Alves, Jeovane Honorio, et al.
Published: (2024)
by: Alves, Jeovane Honorio, et al.
Published: (2024)
RoLegalGEC: Legal Domain Grammatical Error Detection and Correction Dataset for Romanian
by: Timpuriu, Mircea, et al.
Published: (2026)
by: Timpuriu, Mircea, et al.
Published: (2026)
WorldSpeech: A Multilingual Speech Corpus from Around the World
by: Asonitis, Antonis, et al.
Published: (2026)
by: Asonitis, Antonis, et al.
Published: (2026)
Detecting Relevant Information in High-Volume Chat Logs: Keyphrase Extraction for Grooming and Drug Dealing Forensic Analysis
by: Alves, Jeovane Honório, et al.
Published: (2023)
by: Alves, Jeovane Honório, et al.
Published: (2023)
Introducing OmniGEC: A Silver Multilingual Dataset for Grammatical Error Correction
by: Kovalchuk, Roman, et al.
Published: (2025)
by: Kovalchuk, Roman, et al.
Published: (2025)
MultiLegalPile: A 689GB Multilingual Legal Corpus
by: Niklaus, Joel, et al.
Published: (2023)
by: Niklaus, Joel, et al.
Published: (2023)
Unlocking Legal Knowledge: A Multilingual Dataset for Judicial Summarization in Switzerland
by: Rolshoven, Luca, et al.
Published: (2024)
by: Rolshoven, Luca, et al.
Published: (2024)
Improving Multilingual Instruction Finetuning via Linguistically Natural and Diverse Datasets
by: Indurthi, Sathish Reddy, et al.
Published: (2024)
by: Indurthi, Sathish Reddy, et al.
Published: (2024)
MedAidDialog: A Multilingual Multi-Turn Medical Dialogue Dataset for Accessible Healthcare
by: Nigam, Shubham Kumar, et al.
Published: (2026)
by: Nigam, Shubham Kumar, et al.
Published: (2026)
mCSQA: Multilingual Commonsense Reasoning Dataset with Unified Creation Strategy by Language Models and Humans
by: Sakai, Yusuke, et al.
Published: (2024)
by: Sakai, Yusuke, et al.
Published: (2024)
MORPHOGEN: A Multilingual Benchmark for Evaluating Gender-Aware Morphological Generation
by: Agarwal, Mehul, et al.
Published: (2026)
by: Agarwal, Mehul, et al.
Published: (2026)
Legal2LogicICL: Improving Generalization in Transforming Legal Cases to Logical Formulas via Diverse Few-Shot Learning
by: Xue, Jieying, et al.
Published: (2026)
by: Xue, Jieying, et al.
Published: (2026)
Multilingual Instruction Tuning With Just a Pinch of Multilinguality
by: Shaham, Uri, et al.
Published: (2024)
by: Shaham, Uri, et al.
Published: (2024)
NyayaMind- A Framework for Transparent Legal Reasoning and Judgment Prediction in the Indian Legal System
by: Shukla, Parjanya Aditya, et al.
Published: (2026)
by: Shukla, Parjanya Aditya, et al.
Published: (2026)
LegalLens: Leveraging LLMs for Legal Violation Identification in Unstructured Text
by: Bernsohn, Dor, et al.
Published: (2024)
by: Bernsohn, Dor, et al.
Published: (2024)
ViLegalNLI: Natural Language Inference for Vietnamese Legal Texts
by: Duong, Nhung Thi-Hong, et al.
Published: (2026)
by: Duong, Nhung Thi-Hong, et al.
Published: (2026)
Investigating Bias: A Multilingual Pipeline for Generating, Solving, and Evaluating Math Problems with LLMs
by: Mahran, Mariam, et al.
Published: (2025)
by: Mahran, Mariam, et al.
Published: (2025)
WolBanking77: Wolof Banking Speech Intent Classification Dataset
by: Kandji, Abdou Karim, et al.
Published: (2025)
by: Kandji, Abdou Karim, et al.
Published: (2025)
Indian Legal NLP Benchmarks : A Survey
by: Kalamkar, Prathamesh, et al.
Published: (2021)
by: Kalamkar, Prathamesh, et al.
Published: (2021)
The Challenge of Achieving Attributability in Multilingual Table-to-Text Generation with Question-Answer Blueprints
by: Haussmann, Aden
Published: (2025)
by: Haussmann, Aden
Published: (2025)
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
by: Alam, Firoj, et al.
Published: (2026)
by: Alam, Firoj, et al.
Published: (2026)
Align Once, Benefit Multilingually: Enforcing Multilingual Consistency for LLM Safety Alignment
by: Bu, Yuyan, et al.
Published: (2026)
by: Bu, Yuyan, et al.
Published: (2026)
SlovKE: A Large-Scale Dataset and LLM Evaluation for Slovak Keyphrase Extraction
by: Števaňák, David, et al.
Published: (2026)
by: Števaňák, David, et al.
Published: (2026)
LEGAL-UQA: A Low-Resource Urdu-English Dataset for Legal Question Answering
by: Faisal, Faizan, et al.
Published: (2024)
by: Faisal, Faizan, et al.
Published: (2024)
Multilingual Routing in Mixture-of-Experts
by: Bandarkar, Lucas, et al.
Published: (2025)
by: Bandarkar, Lucas, et al.
Published: (2025)
HLDC: Hindi Legal Documents Corpus
by: Kapoor, Arnav, et al.
Published: (2022)
by: Kapoor, Arnav, et al.
Published: (2022)
Lawma: The Power of Specialization for Legal Annotation
by: Dominguez-Olmedo, Ricardo, et al.
Published: (2024)
by: Dominguez-Olmedo, Ricardo, et al.
Published: (2024)
Better as Generators Than Classifiers: Leveraging LLMs and Synthetic Data for Low-Resource Multilingual Classification
by: Pecher, Branislav, et al.
Published: (2026)
by: Pecher, Branislav, et al.
Published: (2026)
mR3: Multilingual Rubric-Agnostic Reward Reasoning Models
by: Anugraha, David, et al.
Published: (2025)
by: Anugraha, David, et al.
Published: (2025)
HumVI: A Multilingual Dataset for Detecting Violent Incidents Impacting Humanitarian Aid
by: Lamba, Hemank, et al.
Published: (2024)
by: Lamba, Hemank, et al.
Published: (2024)
Do Multilingual LLMs Think In English?
by: Schut, Lisa, et al.
Published: (2025)
by: Schut, Lisa, et al.
Published: (2025)
CRAFT Your Dataset: Task-Specific Synthetic Dataset Generation Through Corpus Retrieval and Augmentation
by: Ziegler, Ingo, et al.
Published: (2024)
by: Ziegler, Ingo, et al.
Published: (2024)
ELSA: A Style Aligned Dataset for Emotionally Intelligent Language Generation
by: Gandhi, Vishal, et al.
Published: (2025)
by: Gandhi, Vishal, et al.
Published: (2025)
One Law, Many Languages: Benchmarking Multilingual Legal Reasoning for Judicial Support
by: Stern, Ronja, et al.
Published: (2023)
by: Stern, Ronja, et al.
Published: (2023)
SynthesizRR: Generating Diverse Datasets with Retrieval Augmentation
by: Divekar, Abhishek, et al.
Published: (2024)
by: Divekar, Abhishek, et al.
Published: (2024)
Aligning Language Models for Icelandic Legal Text Summarization
by: Harðarson, Þórir Hrafn, et al.
Published: (2025)
by: Harðarson, Þórir Hrafn, et al.
Published: (2025)
Towards Multilingual LLM Evaluation for European Languages
by: Thellmann, Klaudia, et al.
Published: (2024)
by: Thellmann, Klaudia, et al.
Published: (2024)
Similar Items
-
On Evaluation Protocols for Data Augmentation in a Limited Data Scenario
by: Piedboeuf, Frédéric, et al.
Published: (2024) -
Tagengo: A Multilingual Chat Dataset
by: Devine, Peter
Published: (2024) -
Evaluating the Limits of Large Language Models in Multilingual Legal Reasoning
by: Ioannou, Antreas, et al.
Published: (2025) -
LongKey: Keyphrase Extraction for Long Documents
by: Alves, Jeovane Honorio, et al.
Published: (2024) -
RoLegalGEC: Legal Domain Grammatical Error Detection and Correction Dataset for Romanian
by: Timpuriu, Mircea, et al.
Published: (2026)