LEXam: Benchmarking Legal Reasoning on 340 Law Exams
Fuente:
arXiv
Guardado en:
| Autores principales: | Fan, Yu, Ni, Jingwei, Merane, Jakob, Tian, Yang, Hermstrüwer, Yoan, Huang, Yinya, Akhtar, Mubashara, Salimbeni, Etienne, Geering, Florian, Dreyer, Oliver, Brunner, Daniel, Leippold, Markus, Sachan, Mrinmaya, Stremitzer, Alexander, Engel, Christoph, Ash, Elliott, Niklaus, Joel |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
AFaCTA: Assisting the Annotation of Factual Claim Detection with Reliable LLM Annotators
por: Ni, Jingwei, et al.
Publicado: (2024)
por: Ni, Jingwei, et al.
Publicado: (2024)
DIRAS: Efficient LLM Annotation of Document Relevance in Retrieval Augmented Generation
por: Ni, Jingwei, et al.
Publicado: (2024)
por: Ni, Jingwei, et al.
Publicado: (2024)
ReProbe: Efficient Test-Time Scaling of Multi-Step Reasoning by Probing Internal States of Large Language Models
por: Ni, Jingwei, et al.
Publicado: (2025)
por: Ni, Jingwei, et al.
Publicado: (2025)
Balancing Truthfulness and Informativeness with Uncertainty-Aware Instruction Fine-Tuning
por: Wu, Tianyi, et al.
Publicado: (2025)
por: Wu, Tianyi, et al.
Publicado: (2025)
Compose and Fuse: Revisiting the Foundational Bottlenecks in Multimodal Reasoning
por: Wang, Yucheng, et al.
Publicado: (2025)
por: Wang, Yucheng, et al.
Publicado: (2025)
Can Reasoning Help Large Language Models Capture Human Annotator Disagreement?
por: Ni, Jingwei, et al.
Publicado: (2025)
por: Ni, Jingwei, et al.
Publicado: (2025)
Chimera: Diagnosing Shortcut Learning in Visual-Language Understanding
por: Chi, Ziheng, et al.
Publicado: (2025)
por: Chi, Ziheng, et al.
Publicado: (2025)
Uncovering Hidden Correctness in LLM Causal Reasoning via Symbolic Verification
por: He, Paul, et al.
Publicado: (2026)
por: He, Paul, et al.
Publicado: (2026)
Benchmarking and Enhancing Text-to-Image Models for Generating Visual Representations in Early Arithmetic Education
por: Wang, Junling, et al.
Publicado: (2026)
por: Wang, Junling, et al.
Publicado: (2026)
Towards Faithful and Robust LLM Specialists for Evidence-Based Question-Answering
por: Schimanski, Tobias, et al.
Publicado: (2024)
por: Schimanski, Tobias, et al.
Publicado: (2024)
Modeling Motivated Reasoning in Law: Evaluating Strategic Role Conditioning in LLM Summarization
por: Cho, Eunjung, et al.
Publicado: (2025)
por: Cho, Eunjung, et al.
Publicado: (2025)
The Medium Is Not the Message: Deconfounding Document Embeddings via Linear Concept Erasure
por: Fan, Yu, et al.
Publicado: (2025)
por: Fan, Yu, et al.
Publicado: (2025)
Co-DETECT: Collaborative Discovery of Edge Cases in Text Classification
por: Xiong, Chenfei, et al.
Publicado: (2025)
por: Xiong, Chenfei, et al.
Publicado: (2025)
pdfQA: Diverse, Challenging, and Realistic Question Answering over PDFs
por: Schimanski, Tobias, et al.
Publicado: (2026)
por: Schimanski, Tobias, et al.
Publicado: (2026)
Agency Theory: Methodology, Analysis
por: Stremitzer, Alexander
Publicado: (2019)
por: Stremitzer, Alexander
Publicado: (2019)
Investigating the Zone of Proximal Development of Language Models for In-Context Learning
por: Cui, Peng, et al.
Publicado: (2025)
por: Cui, Peng, et al.
Publicado: (2025)
AI-assisted Automated Short Answer Grading of Handwritten University Level Mathematics Exams
por: Liu, Tianyi, et al.
Publicado: (2024)
por: Liu, Tianyi, et al.
Publicado: (2024)
Ev2R: Evaluating Evidence Retrieval in Automated Fact-Checking
por: Akhtar, Mubashara, et al.
Publicado: (2024)
por: Akhtar, Mubashara, et al.
Publicado: (2024)
Tackling the Root of Misinformation by Teaching Laypeople about Logical Fallacies via Socratic Questioning and Critical Argumentation
por: Shi, Minjing, et al.
Publicado: (2026)
por: Shi, Minjing, et al.
Publicado: (2026)
Genere e spazio urbano
por: Salimbeni, Alice
Publicado: (2024)
por: Salimbeni, Alice
Publicado: (2024)
Augusto Del Noce e l’origine della valutazione critica del moderno. L’incontro con Niccolò Machiavelli
por: Filippo Salimbeni
Publicado: (2024)
por: Filippo Salimbeni
Publicado: (2024)
Simulating Students or Sycophantic Problem Solving? On Misconception Faithfulness of LLM Simulators
por: Do, Heejin, et al.
Publicado: (2026)
por: Do, Heejin, et al.
Publicado: (2026)
Probing for Arithmetic Errors in Language Models
por: Sun, Yucheng, et al.
Publicado: (2025)
por: Sun, Yucheng, et al.
Publicado: (2025)
Variational Classification
por: Dhuliawala, Shehzaad, et al.
Publicado: (2023)
por: Dhuliawala, Shehzaad, et al.
Publicado: (2023)
AI-Assisted Human Evaluation of Machine Translation
por: Zouhar, Vilém, et al.
Publicado: (2024)
por: Zouhar, Vilém, et al.
Publicado: (2024)
How to Select Datapoints for Efficient Human Evaluation of NLG Models?
por: Zouhar, Vilém, et al.
Publicado: (2025)
por: Zouhar, Vilém, et al.
Publicado: (2025)
Automated Knowledge Concept Annotation and Question Representation Learning for Knowledge Tracing
por: Ozyurt, Yilmazcan, et al.
Publicado: (2024)
por: Ozyurt, Yilmazcan, et al.
Publicado: (2024)
SwiLTra-Bench: The Swiss Legal Translation Benchmark
por: Niklaus, Joel, et al.
Publicado: (2025)
por: Niklaus, Joel, et al.
Publicado: (2025)
GPT-4 as a Homework Tutor can Improve Student Engagement and Learning Outcomes
por: Vanzo, Alessandro, et al.
Publicado: (2024)
por: Vanzo, Alessandro, et al.
Publicado: (2024)
AutoTutor meets Large Language Models: A Language Model Tutor with Rich Pedagogy and Guardrails
por: Chowdhury, Sankalan Pal, et al.
Publicado: (2024)
por: Chowdhury, Sankalan Pal, et al.
Publicado: (2024)
Lean Meets Theoretical Computer Science: Scalable Synthesis of Theorem Proving Challenges in Formal-Informal Pairs
por: Zhang, Terry Jingchen, et al.
Publicado: (2025)
por: Zhang, Terry Jingchen, et al.
Publicado: (2025)
Translating Legalese: Enhancing Public Understanding of Court Opinions with Legal Summarizers
por: Ash, Elliott, et al.
Publicado: (2023)
por: Ash, Elliott, et al.
Publicado: (2023)
Efficiently Computing Susceptibility to Context in Language Models
por: Liu, Tianyu, et al.
Publicado: (2024)
por: Liu, Tianyu, et al.
Publicado: (2024)
Improving Large Language Model Safety with Contrastive Representation Learning
por: Simko, Samuel, et al.
Publicado: (2025)
por: Simko, Samuel, et al.
Publicado: (2025)
Unveiling the Visual Counting Bottleneck in Vision-Language Models
por: Pang, Xingzhou, et al.
Publicado: (2026)
por: Pang, Xingzhou, et al.
Publicado: (2026)
Personalized Exercise Recommendation with Semantically-Grounded Knowledge Tracing
por: Ozyurt, Yilmazcan, et al.
Publicado: (2025)
por: Ozyurt, Yilmazcan, et al.
Publicado: (2025)
Sample Smart, Not Hard: Correctness-First Decoding for Better Reasoning in LLMs
por: Li, Xueyan, et al.
Publicado: (2025)
por: Li, Xueyan, et al.
Publicado: (2025)
Grammar Control in Dialogue Response Generation for Language Learning Chatbots
por: Glandorf, Dominik, et al.
Publicado: (2025)
por: Glandorf, Dominik, et al.
Publicado: (2025)
Do Vision-Language Models Really Understand Visual Language?
por: Hou, Yifan, et al.
Publicado: (2024)
por: Hou, Yifan, et al.
Publicado: (2024)
What Do Language Models Learn in Context? The Structured Task Hypothesis
por: Li, Jiaoda, et al.
Publicado: (2024)
por: Li, Jiaoda, et al.
Publicado: (2024)
Ejemplares similares
-
AFaCTA: Assisting the Annotation of Factual Claim Detection with Reliable LLM Annotators
por: Ni, Jingwei, et al.
Publicado: (2024) -
DIRAS: Efficient LLM Annotation of Document Relevance in Retrieval Augmented Generation
por: Ni, Jingwei, et al.
Publicado: (2024) -
ReProbe: Efficient Test-Time Scaling of Multi-Step Reasoning by Probing Internal States of Large Language Models
por: Ni, Jingwei, et al.
Publicado: (2025) -
Balancing Truthfulness and Informativeness with Uncertainty-Aware Instruction Fine-Tuning
por: Wu, Tianyi, et al.
Publicado: (2025) -
Compose and Fuse: Revisiting the Foundational Bottlenecks in Multimodal Reasoning
por: Wang, Yucheng, et al.
Publicado: (2025)