LEXTREME: A Multi-Lingual and Multi-Task Benchmark for the Legal Domain
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Niklaus, Joel, Matoshi, Veton, Rani, Pooja, Galassi, Andrea, Stürmer, Matthias, Chalkidis, Ilias |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MultiLegalPile: A 689GB Multilingual Legal Corpus
von: Niklaus, Joel, et al.
Veröffentlicht: (2023)
von: Niklaus, Joel, et al.
Veröffentlicht: (2023)
One Law, Many Languages: Benchmarking Multilingual Legal Reasoning for Judicial Support
von: Stern, Ronja, et al.
Veröffentlicht: (2023)
von: Stern, Ronja, et al.
Veröffentlicht: (2023)
From Citations to Criticality: Predicting Legal Decision Influence in the Multilingual Swiss Jurisprudence
von: Stern, Ronja, et al.
Veröffentlicht: (2024)
von: Stern, Ronja, et al.
Veröffentlicht: (2024)
Towards Explainability and Fairness in Swiss Judgement Prediction: Benchmarking on a Multilingual Dataset
von: S, Santosh T. Y. S., et al.
Veröffentlicht: (2024)
von: S, Santosh T. Y. S., et al.
Veröffentlicht: (2024)
Anonymity at Risk? Assessing Re-Identification Capabilities of Large Language Models
von: Nyffenegger, Alex, et al.
Veröffentlicht: (2023)
von: Nyffenegger, Alex, et al.
Veröffentlicht: (2023)
Unlocking Legal Knowledge: A Multilingual Dataset for Judicial Summarization in Switzerland
von: Rolshoven, Luca, et al.
Veröffentlicht: (2024)
von: Rolshoven, Luca, et al.
Veröffentlicht: (2024)
LawInstruct: A Resource for Studying Language Model Adaptation to the Legal Domain
von: Niklaus, Joel, et al.
Veröffentlicht: (2024)
von: Niklaus, Joel, et al.
Veröffentlicht: (2024)
LEXam: Benchmarking Legal Reasoning on 340 Law Exams
von: Fan, Yu, et al.
Veröffentlicht: (2025)
von: Fan, Yu, et al.
Veröffentlicht: (2025)
SwiLTra-Bench: The Swiss Legal Translation Benchmark
von: Niklaus, Joel, et al.
Veröffentlicht: (2025)
von: Niklaus, Joel, et al.
Veröffentlicht: (2025)
HR-MultiWOZ: A Task Oriented Dialogue (TOD) Dataset for HR LLM Agent
von: Xu, Weijie, et al.
Veröffentlicht: (2024)
von: Xu, Weijie, et al.
Veröffentlicht: (2024)
Mitigating Manipulation and Enhancing Persuasion: A Reflective Multi-Agent Approach for Legal Argument Generation
von: Zhang, Li, et al.
Veröffentlicht: (2025)
von: Zhang, Li, et al.
Veröffentlicht: (2025)
Pretraining and Updates of Domain-Specific LLM: A Case Study in the Japanese Business Domain
von: Takahashi, Kosuke, et al.
Veröffentlicht: (2024)
von: Takahashi, Kosuke, et al.
Veröffentlicht: (2024)
Multi-Model Synthetic Training for Mission-Critical Small Language Models
von: Platt, Nolan, et al.
Veröffentlicht: (2025)
von: Platt, Nolan, et al.
Veröffentlicht: (2025)
MORABLES: A Benchmark for Assessing Abstract Moral Reasoning in LLMs with Fables
von: Marcuzzo, Matteo, et al.
Veröffentlicht: (2025)
von: Marcuzzo, Matteo, et al.
Veröffentlicht: (2025)
K-MetBench: A Multi-Dimensional Benchmark for Fine-Grained Evaluation of Expert Reasoning, Locality, and Multimodality in Meteorology
von: Kim, Soyeon, et al.
Veröffentlicht: (2026)
von: Kim, Soyeon, et al.
Veröffentlicht: (2026)
Multi-Task Contrastive Learning for 8192-Token Bilingual Text Embeddings
von: Mohr, Isabelle, et al.
Veröffentlicht: (2024)
von: Mohr, Isabelle, et al.
Veröffentlicht: (2024)
MetaCheckGPT -- A Multi-task Hallucination Detector Using LLM Uncertainty and Meta-models
von: Mehta, Rahul, et al.
Veröffentlicht: (2024)
von: Mehta, Rahul, et al.
Veröffentlicht: (2024)
RV-HATE: Reinforced Multi-Module Voting for Implicit Hate Speech Detection
von: Lee, Yejin, et al.
Veröffentlicht: (2025)
von: Lee, Yejin, et al.
Veröffentlicht: (2025)
Multi-chain Graph Refinement and Selection for Reliable Reasoning in Large Language Models
von: Yang, Yujiao, et al.
Veröffentlicht: (2025)
von: Yang, Yujiao, et al.
Veröffentlicht: (2025)
ADE: Adaptive Dictionary Embeddings -- Scaling Multi-Anchor Representations to Large Language Models
von: Demirci, Orhan, et al.
Veröffentlicht: (2026)
von: Demirci, Orhan, et al.
Veröffentlicht: (2026)
A Graph-based Approach for Multi-Modal Question Answering from Flowcharts in Telecom Documents
von: Soman, Sumit, et al.
Veröffentlicht: (2025)
von: Soman, Sumit, et al.
Veröffentlicht: (2025)
Prompting Encoder Models for Zero-Shot Classification: A Cross-Domain Study in Italian
von: Auriemma, Serena, et al.
Veröffentlicht: (2024)
von: Auriemma, Serena, et al.
Veröffentlicht: (2024)
Learning When to Think: Shaping Adaptive Reasoning in R1-Style Models via Multi-Stage RL
von: Tu, Songjun, et al.
Veröffentlicht: (2025)
von: Tu, Songjun, et al.
Veröffentlicht: (2025)
Raw Text is All you Need: Knowledge-intensive Multi-turn Instruction Tuning for Large Language Model
von: Hou, Xia, et al.
Veröffentlicht: (2024)
von: Hou, Xia, et al.
Veröffentlicht: (2024)
Task Complexity Matters: An Empirical Study of Reasoning in LLMs for Sentiment Analysis
von: Huang, Donghao, et al.
Veröffentlicht: (2026)
von: Huang, Donghao, et al.
Veröffentlicht: (2026)
One Agent to Serve All: a Lite-Adaptive Stylized AI Assistant for Millions of Multi-Style Official Accounts
von: Fan, Xingyu, et al.
Veröffentlicht: (2025)
von: Fan, Xingyu, et al.
Veröffentlicht: (2025)
Triad: A Framework Leveraging a Multi-Role LLM-based Agent to Solve Knowledge Base Question Answering
von: Zong, Chang, et al.
Veröffentlicht: (2024)
von: Zong, Chang, et al.
Veröffentlicht: (2024)
Prompt Tuned Embedding Classification for Multi-Label Industry Sector Allocation
von: Buchner, Valentin Leonhard, et al.
Veröffentlicht: (2023)
von: Buchner, Valentin Leonhard, et al.
Veröffentlicht: (2023)
lmfaoooo at SemEval-2026 Task 1: Humor Is an Audience. Preference Modeling for Constrained Humor Generation
von: Tikhonov, Alexey, et al.
Veröffentlicht: (2026)
von: Tikhonov, Alexey, et al.
Veröffentlicht: (2026)
LayerTracer: A Joint Task-Particle and Vulnerable-Layer Analysis framework for Arbitrary Large Language Model Architectures
von: Wu, Yuhang, et al.
Veröffentlicht: (2026)
von: Wu, Yuhang, et al.
Veröffentlicht: (2026)
ART: Adaptive Response Tuning Framework -- A Multi-Agent Tournament-Based Approach to LLM Response Optimization
von: Khan, Omer Jauhar
Veröffentlicht: (2025)
von: Khan, Omer Jauhar
Veröffentlicht: (2025)
Mitigating LLM Hallucinations through Domain-Grounded Tiered Retrieval
von: Haque, Md. Asraful, et al.
Veröffentlicht: (2026)
von: Haque, Md. Asraful, et al.
Veröffentlicht: (2026)
Unsolvability Ceiling in Multi-LLM Routing: An Empirical Study of Evaluation Artifacts
von: Garg, Saloni, et al.
Veröffentlicht: (2026)
von: Garg, Saloni, et al.
Veröffentlicht: (2026)
Pun Unintended: LLMs and the Illusion of Humor Understanding
von: Zangari, Alessandro, et al.
Veröffentlicht: (2025)
von: Zangari, Alessandro, et al.
Veröffentlicht: (2025)
The Illusion of Role Separation: Hidden Shortcuts in LLM Role Learning (and How to Fix Them)
von: Wang, Zihao, et al.
Veröffentlicht: (2025)
von: Wang, Zihao, et al.
Veröffentlicht: (2025)
Transforming and Combining Rewards for Aligning Large Language Models
von: Wang, Zihao, et al.
Veröffentlicht: (2024)
von: Wang, Zihao, et al.
Veröffentlicht: (2024)
Reasoning Promotes Robustness in Theory of Mind Tasks
von: de Haan, Ian B., et al.
Veröffentlicht: (2026)
von: de Haan, Ian B., et al.
Veröffentlicht: (2026)
TSDS: Data Selection for Task-Specific Model Finetuning
von: Liu, Zifan, et al.
Veröffentlicht: (2024)
von: Liu, Zifan, et al.
Veröffentlicht: (2024)
KIT-TIP-NLP at MultiPride: Continual Learning with Multilingual Foundation Model
von: HB, Barathi Ganesh, et al.
Veröffentlicht: (2026)
von: HB, Barathi Ganesh, et al.
Veröffentlicht: (2026)
Can Out-of-Distribution Evaluations Uncover Reliance on Shortcuts? A Case Study in Question Answering
von: Štefánik, Michal, et al.
Veröffentlicht: (2025)
von: Štefánik, Michal, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MultiLegalPile: A 689GB Multilingual Legal Corpus
von: Niklaus, Joel, et al.
Veröffentlicht: (2023) -
One Law, Many Languages: Benchmarking Multilingual Legal Reasoning for Judicial Support
von: Stern, Ronja, et al.
Veröffentlicht: (2023) -
From Citations to Criticality: Predicting Legal Decision Influence in the Multilingual Swiss Jurisprudence
von: Stern, Ronja, et al.
Veröffentlicht: (2024) -
Towards Explainability and Fairness in Swiss Judgement Prediction: Benchmarking on a Multilingual Dataset
von: S, Santosh T. Y. S., et al.
Veröffentlicht: (2024) -
Anonymity at Risk? Assessing Re-Identification Capabilities of Large Language Models
von: Nyffenegger, Alex, et al.
Veröffentlicht: (2023)