MultiLegalPile: A 689GB Multilingual Legal Corpus
Fuente:
arXiv
Saved in:
| Main Authors: | Niklaus, Joel, Matoshi, Veton, Stürmer, Matthias, Chalkidis, Ilias, Ho, Daniel E. |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LEXTREME: A Multi-Lingual and Multi-Task Benchmark for the Legal Domain
by: Niklaus, Joel, et al.
Published: (2023)
by: Niklaus, Joel, et al.
Published: (2023)
One Law, Many Languages: Benchmarking Multilingual Legal Reasoning for Judicial Support
by: Stern, Ronja, et al.
Published: (2023)
by: Stern, Ronja, et al.
Published: (2023)
From Citations to Criticality: Predicting Legal Decision Influence in the Multilingual Swiss Jurisprudence
by: Stern, Ronja, et al.
Published: (2024)
by: Stern, Ronja, et al.
Published: (2024)
Unlocking Legal Knowledge: A Multilingual Dataset for Judicial Summarization in Switzerland
by: Rolshoven, Luca, et al.
Published: (2024)
by: Rolshoven, Luca, et al.
Published: (2024)
Towards Explainability and Fairness in Swiss Judgement Prediction: Benchmarking on a Multilingual Dataset
by: S, Santosh T. Y. S., et al.
Published: (2024)
by: S, Santosh T. Y. S., et al.
Published: (2024)
Anonymity at Risk? Assessing Re-Identification Capabilities of Large Language Models
by: Nyffenegger, Alex, et al.
Published: (2023)
by: Nyffenegger, Alex, et al.
Published: (2023)
LawInstruct: A Resource for Studying Language Model Adaptation to the Legal Domain
by: Niklaus, Joel, et al.
Published: (2024)
by: Niklaus, Joel, et al.
Published: (2024)
LEXam: Benchmarking Legal Reasoning on 340 Law Exams
by: Fan, Yu, et al.
Published: (2025)
by: Fan, Yu, et al.
Published: (2025)
SwiLTra-Bench: The Swiss Legal Translation Benchmark
by: Niklaus, Joel, et al.
Published: (2025)
by: Niklaus, Joel, et al.
Published: (2025)
Mitigating Manipulation and Enhancing Persuasion: A Reflective Multi-Agent Approach for Legal Argument Generation
by: Zhang, Li, et al.
Published: (2025)
by: Zhang, Li, et al.
Published: (2025)
MemeLens: Multilingual Multitask VLMs for Memes
by: Shahroor, Ali Ezzat, et al.
Published: (2026)
by: Shahroor, Ali Ezzat, et al.
Published: (2026)
SpokenNativQA: Multilingual Everyday Spoken Queries for LLMs
by: Alam, Firoj, et al.
Published: (2025)
by: Alam, Firoj, et al.
Published: (2025)
Temporal Concept Drift in Legal Judgment Prediction: Neural Baselines Across Three Epochs of Ukrainian Court Decisions
by: Ovcharov, Volodymyr
Published: (2026)
by: Ovcharov, Volodymyr
Published: (2026)
Thinking Longer, Not Always Smarter: Evaluating LLM Capabilities in Hierarchical Legal Reasoning
by: Zhang, Li, et al.
Published: (2025)
by: Zhang, Li, et al.
Published: (2025)
LLMs for Legal Subsumption in German Employment Contracts
by: Wardas, Oliver, et al.
Published: (2025)
by: Wardas, Oliver, et al.
Published: (2025)
Japanese Tort-case Dataset for Rationale-supported Legal Judgment Prediction
by: Yamada, Hiroaki, et al.
Published: (2023)
by: Yamada, Hiroaki, et al.
Published: (2023)
KIT-TIP-NLP at MultiPride: Continual Learning with Multilingual Foundation Model
by: HB, Barathi Ganesh, et al.
Published: (2026)
by: HB, Barathi Ganesh, et al.
Published: (2026)
Multi-Model Synthetic Training for Mission-Critical Small Language Models
by: Platt, Nolan, et al.
Published: (2025)
by: Platt, Nolan, et al.
Published: (2025)
Aligning LLMs for Multilingual Consistency in Enterprise Applications
by: Agarwal, Amit, et al.
Published: (2025)
by: Agarwal, Amit, et al.
Published: (2025)
Retrieval-Based Multi-Label Legal Annotation: Extensible, Data-Efficient and Hallucination-Free
by: Zhang, Li, et al.
Published: (2026)
by: Zhang, Li, et al.
Published: (2026)
MetaCheckGPT -- A Multi-task Hallucination Detector Using LLM Uncertainty and Meta-models
by: Mehta, Rahul, et al.
Published: (2024)
by: Mehta, Rahul, et al.
Published: (2024)
Structured Sentiment Analysis as Transition-based Dependency Graph Parsing
by: Fernández-González, Daniel
Published: (2023)
by: Fernández-González, Daniel
Published: (2023)
LegalGuardian: A Privacy-Preserving Framework for Secure Integration of Large Language Models in Legal Practice
by: Demir, M. Mikail, et al.
Published: (2025)
by: Demir, M. Mikail, et al.
Published: (2025)
RV-HATE: Reinforced Multi-Module Voting for Implicit Hate Speech Detection
by: Lee, Yejin, et al.
Published: (2025)
by: Lee, Yejin, et al.
Published: (2025)
Multi-chain Graph Refinement and Selection for Reliable Reasoning in Large Language Models
by: Yang, Yujiao, et al.
Published: (2025)
by: Yang, Yujiao, et al.
Published: (2025)
ADE: Adaptive Dictionary Embeddings -- Scaling Multi-Anchor Representations to Large Language Models
by: Demirci, Orhan, et al.
Published: (2026)
by: Demirci, Orhan, et al.
Published: (2026)
A Graph-based Approach for Multi-Modal Question Answering from Flowcharts in Telecom Documents
by: Soman, Sumit, et al.
Published: (2025)
by: Soman, Sumit, et al.
Published: (2025)
HR-MultiWOZ: A Task Oriented Dialogue (TOD) Dataset for HR LLM Agent
by: Xu, Weijie, et al.
Published: (2024)
by: Xu, Weijie, et al.
Published: (2024)
jina-embeddings-v4: Universal Embeddings for Multimodal Multilingual Retrieval
by: Günther, Michael, et al.
Published: (2025)
by: Günther, Michael, et al.
Published: (2025)
jina-embeddings-v3: Multilingual Embeddings With Task LoRA
by: Sturua, Saba, et al.
Published: (2024)
by: Sturua, Saba, et al.
Published: (2024)
Learning When to Think: Shaping Adaptive Reasoning in R1-Style Models via Multi-Stage RL
by: Tu, Songjun, et al.
Published: (2025)
by: Tu, Songjun, et al.
Published: (2025)
Raw Text is All you Need: Knowledge-intensive Multi-turn Instruction Tuning for Large Language Model
by: Hou, Xia, et al.
Published: (2024)
by: Hou, Xia, et al.
Published: (2024)
One Agent to Serve All: a Lite-Adaptive Stylized AI Assistant for Millions of Multi-Style Official Accounts
by: Fan, Xingyu, et al.
Published: (2025)
by: Fan, Xingyu, et al.
Published: (2025)
Triad: A Framework Leveraging a Multi-Role LLM-based Agent to Solve Knowledge Base Question Answering
by: Zong, Chang, et al.
Published: (2024)
by: Zong, Chang, et al.
Published: (2024)
The Illusion of Role Separation: Hidden Shortcuts in LLM Role Learning (and How to Fix Them)
by: Wang, Zihao, et al.
Published: (2025)
by: Wang, Zihao, et al.
Published: (2025)
Pretraining and Updates of Domain-Specific LLM: A Case Study in the Japanese Business Domain
by: Takahashi, Kosuke, et al.
Published: (2024)
by: Takahashi, Kosuke, et al.
Published: (2024)
Transforming and Combining Rewards for Aligning Large Language Models
by: Wang, Zihao, et al.
Published: (2024)
by: Wang, Zihao, et al.
Published: (2024)
Jina-ColBERT-v2: A General-Purpose Multilingual Late Interaction Retriever
by: Jha, Rohan, et al.
Published: (2024)
by: Jha, Rohan, et al.
Published: (2024)
Can Out-of-Distribution Evaluations Uncover Reliance on Shortcuts? A Case Study in Question Answering
by: Štefánik, Michal, et al.
Published: (2025)
by: Štefánik, Michal, et al.
Published: (2025)
Prompt Tuned Embedding Classification for Multi-Label Industry Sector Allocation
by: Buchner, Valentin Leonhard, et al.
Published: (2023)
by: Buchner, Valentin Leonhard, et al.
Published: (2023)
Similar Items
-
LEXTREME: A Multi-Lingual and Multi-Task Benchmark for the Legal Domain
by: Niklaus, Joel, et al.
Published: (2023) -
One Law, Many Languages: Benchmarking Multilingual Legal Reasoning for Judicial Support
by: Stern, Ronja, et al.
Published: (2023) -
From Citations to Criticality: Predicting Legal Decision Influence in the Multilingual Swiss Jurisprudence
by: Stern, Ronja, et al.
Published: (2024) -
Unlocking Legal Knowledge: A Multilingual Dataset for Judicial Summarization in Switzerland
by: Rolshoven, Luca, et al.
Published: (2024) -
Towards Explainability and Fairness in Swiss Judgement Prediction: Benchmarking on a Multilingual Dataset
by: S, Santosh T. Y. S., et al.
Published: (2024)