Better Call CLAUSE: A Discrepancy Benchmark for Auditing LLMs Legal Reasoning Capabilities
Fuente:
arXiv
Salvato in:
| Autori principali: | Choudhury, Manan Roy, Chandramouli, Adithya, Anand, Mannan, Gupta, Vivek |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
FinRule-Bench: A Benchmark for Joint Reasoning over Financial Tables and Principles
di: Malarkkan, Arun Vignesh, et al.
Pubblicazione: (2026)
di: Malarkkan, Arun Vignesh, et al.
Pubblicazione: (2026)
CLAUSE: Agentic Neuro-Symbolic Knowledge Graph Reasoning via Dynamic Learnable Context Engineering
di: Zhao, Yang, et al.
Pubblicazione: (2025)
di: Zhao, Yang, et al.
Pubblicazione: (2025)
Exploiting LLMs' Reasoning Capability to Infer Implicit Concepts in Legal Information Retrieval
di: Nguyen, Hai-Long, et al.
Pubblicazione: (2024)
di: Nguyen, Hai-Long, et al.
Pubblicazione: (2024)
CIRCUIT: A Benchmark for Circuit Interpretation and Reasoning Capabilities of LLMs
di: Skelic, Lejla, et al.
Pubblicazione: (2025)
di: Skelic, Lejla, et al.
Pubblicazione: (2025)
DRAGON: A Benchmark for Evidence-Grounded Visual Reasoning over Diagrams
di: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Pubblicazione: (2026)
di: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Pubblicazione: (2026)
Benchmarking Spatiotemporal Reasoning in LLMs and Reasoning Models: Capabilities and Challenges
di: Quan, Pengrui, et al.
Pubblicazione: (2025)
di: Quan, Pengrui, et al.
Pubblicazione: (2025)
HiBench: Benchmarking LLMs Capability on Hierarchical Structure Reasoning
di: Jiang, Zhuohang, et al.
Pubblicazione: (2025)
di: Jiang, Zhuohang, et al.
Pubblicazione: (2025)
Do Moral Judgment and Reasoning Capability of LLMs Change with Language? A Study using the Multilingual Defining Issues Test
di: Khandelwal, Aditi, et al.
Pubblicazione: (2024)
di: Khandelwal, Aditi, et al.
Pubblicazione: (2024)
Oolong: Evaluating Long Context Reasoning and Aggregation Capabilities
di: Bertsch, Amanda, et al.
Pubblicazione: (2025)
di: Bertsch, Amanda, et al.
Pubblicazione: (2025)
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases
di: Gan, Eric, et al.
Pubblicazione: (2026)
di: Gan, Eric, et al.
Pubblicazione: (2026)
Better Accuracies, Worse Reasoning: A Step-Level Audit of Medical Chain-of-Thought Distillation
di: Jiang, Zhaoyang, et al.
Pubblicazione: (2026)
di: Jiang, Zhaoyang, et al.
Pubblicazione: (2026)
Small Language Models for Agentic Systems: A Survey of Architectures, Capabilities, and Deployment Trade offs
di: Sharma, Raghav, et al.
Pubblicazione: (2025)
di: Sharma, Raghav, et al.
Pubblicazione: (2025)
Are LLMs Capable of Data-based Statistical and Causal Reasoning? Benchmarking Advanced Quantitative Reasoning with Data
di: Liu, Xiao, et al.
Pubblicazione: (2024)
di: Liu, Xiao, et al.
Pubblicazione: (2024)
CyberSOCEval: Benchmarking LLMs Capabilities for Malware Analysis and Threat Intelligence Reasoning
di: Deason, Lauren, et al.
Pubblicazione: (2025)
di: Deason, Lauren, et al.
Pubblicazione: (2025)
Benchmarking the Legal Reasoning of LLMs in Arabic Islamic Inheritance Cases
di: AlDahoul, Nouar, et al.
Pubblicazione: (2025)
di: AlDahoul, Nouar, et al.
Pubblicazione: (2025)
Auditing the Ethical Logic of Generative AI Models
di: Neuman, W. Russell, et al.
Pubblicazione: (2025)
di: Neuman, W. Russell, et al.
Pubblicazione: (2025)
Indian Legal NLP Benchmarks : A Survey
di: Kalamkar, Prathamesh, et al.
Pubblicazione: (2021)
di: Kalamkar, Prathamesh, et al.
Pubblicazione: (2021)
Ko-MuSR: A Multistep Soft Reasoning Benchmark for LLMs Capable of Understanding Korean
di: Park, Chanwoo, et al.
Pubblicazione: (2025)
di: Park, Chanwoo, et al.
Pubblicazione: (2025)
Retrieval-Infused Reasoning Sandbox: A Benchmark for Decoupling Retrieval and Reasoning Capabilities
di: Ying, Shuangshuang, et al.
Pubblicazione: (2026)
di: Ying, Shuangshuang, et al.
Pubblicazione: (2026)
CoRe: Benchmarking LLMs Code Reasoning Capabilities through Static Analysis Tasks
di: Xie, Danning, et al.
Pubblicazione: (2025)
di: Xie, Danning, et al.
Pubblicazione: (2025)
Are Your LLMs Capable of Stable Reasoning?
di: Liu, Junnan, et al.
Pubblicazione: (2024)
di: Liu, Junnan, et al.
Pubblicazione: (2024)
CHBench: A Cognitive Hierarchy Benchmark for Evaluating Strategic Reasoning Capability of LLMs
di: Liu, Hongtao, et al.
Pubblicazione: (2025)
di: Liu, Hongtao, et al.
Pubblicazione: (2025)
Exploring the psychology of LLMs' Moral and Legal Reasoning
di: Almeida, Guilherme F. C. F., et al.
Pubblicazione: (2023)
di: Almeida, Guilherme F. C. F., et al.
Pubblicazione: (2023)
Structured Definitions and Segmentations for Legal Reasoning in LLMs: A Study on Indian Legal Data
di: Khatri, Mann, et al.
Pubblicazione: (2025)
di: Khatri, Mann, et al.
Pubblicazione: (2025)
Understanding the Geospatial Reasoning Capabilities of LLMs: A Trajectory Recovery Perspective
di: Truong, Thinh Hung, et al.
Pubblicazione: (2025)
di: Truong, Thinh Hung, et al.
Pubblicazione: (2025)
NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls
di: Basu, Kinjal, et al.
Pubblicazione: (2024)
di: Basu, Kinjal, et al.
Pubblicazione: (2024)
Interpretable Emergent Language Using Inter-Agent Transformers
di: Bhardwaj, Mannan
Pubblicazione: (2025)
di: Bhardwaj, Mannan
Pubblicazione: (2025)
Explore the Reasoning Capability of LLMs in the Chess Testbed
di: Wang, Shu, et al.
Pubblicazione: (2024)
di: Wang, Shu, et al.
Pubblicazione: (2024)
EarthSpatialBench: Benchmarking Spatial Reasoning Capabilities of Multimodal LLMs on Earth Imagery
di: Xu, Zelin, et al.
Pubblicazione: (2026)
di: Xu, Zelin, et al.
Pubblicazione: (2026)
RESOLVE: Relational Reasoning with Symbolic and Object-Level Features Using Vector Symbolic Processing
di: Mejri, Mohamed, et al.
Pubblicazione: (2024)
di: Mejri, Mohamed, et al.
Pubblicazione: (2024)
Diagnosing LLM Judge Reliability: Conformal Prediction Sets and Transitivity Violations
di: Gupta, Manan, et al.
Pubblicazione: (2026)
di: Gupta, Manan, et al.
Pubblicazione: (2026)
Latent Phase-Shift Rollback: Inference-Time Error Correction via Residual Stream Monitoring and KV-Cache Steering
di: Gupta, Manan, et al.
Pubblicazione: (2026)
di: Gupta, Manan, et al.
Pubblicazione: (2026)
Auditing of AI: Legal, Ethical and Technical Approaches
di: Mokander, Jakob
Pubblicazione: (2024)
di: Mokander, Jakob
Pubblicazione: (2024)
NTSEBENCH: Cognitive Reasoning Benchmark for Vision Language Models
di: Pandya, Pranshu, et al.
Pubblicazione: (2024)
di: Pandya, Pranshu, et al.
Pubblicazione: (2024)
Unequal Voices: How LLMs Construct Constrained Queer Narratives
di: Ghosal, Atreya, et al.
Pubblicazione: (2025)
di: Ghosal, Atreya, et al.
Pubblicazione: (2025)
ALARB: An Arabic Legal Argument Reasoning Benchmark
di: Shairah, Harethah Abu, et al.
Pubblicazione: (2025)
di: Shairah, Harethah Abu, et al.
Pubblicazione: (2025)
ECG-Reasoning-Benchmark: A Benchmark for Evaluating Clinical Reasoning Capabilities in ECG Interpretation
di: Oh, Jungwoo, et al.
Pubblicazione: (2026)
di: Oh, Jungwoo, et al.
Pubblicazione: (2026)
Can LLMs Learn by Teaching for Better Reasoning? A Preliminary Study
di: Ning, Xuefei, et al.
Pubblicazione: (2024)
di: Ning, Xuefei, et al.
Pubblicazione: (2024)
LegalReasoner: Step-wised Verification-Correction for Legal Judgment Reasoning
di: Shi, Weijie, et al.
Pubblicazione: (2025)
di: Shi, Weijie, et al.
Pubblicazione: (2025)
Toward Better EHR Reasoning in LLMs: Reinforcement Learning with Expert Attention Guidance
di: Fang, Yue, et al.
Pubblicazione: (2025)
di: Fang, Yue, et al.
Pubblicazione: (2025)
Documenti analoghi
-
FinRule-Bench: A Benchmark for Joint Reasoning over Financial Tables and Principles
di: Malarkkan, Arun Vignesh, et al.
Pubblicazione: (2026) -
CLAUSE: Agentic Neuro-Symbolic Knowledge Graph Reasoning via Dynamic Learnable Context Engineering
di: Zhao, Yang, et al.
Pubblicazione: (2025) -
Exploiting LLMs' Reasoning Capability to Infer Implicit Concepts in Legal Information Retrieval
di: Nguyen, Hai-Long, et al.
Pubblicazione: (2024) -
CIRCUIT: A Benchmark for Circuit Interpretation and Reasoning Capabilities of LLMs
di: Skelic, Lejla, et al.
Pubblicazione: (2025) -
DRAGON: A Benchmark for Evidence-Grounded Visual Reasoning over Diagrams
di: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Pubblicazione: (2026)