Where Experts Disagree, Models Fail: Detecting Implicit Legal Citations in French Court Decisions
Fuente:
arXiv
Saved in:
| Main Authors: | Floro, Avrile, Dhorasoo, Tamara, Pellez, Soline, Holzenberger, Nils |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cultural Binding Heads in Language Models
by: Floro, Avrile, et al.
Published: (2026)
by: Floro, Avrile, et al.
Published: (2026)
BLT: Can Large Language Models Handle Basic Legal Text?
by: Blair-Stanek, Andrew, et al.
Published: (2023)
by: Blair-Stanek, Andrew, et al.
Published: (2023)
The Factuality of Large Language Models in the Legal Domain
by: Hamdani, Rajaa El, et al.
Published: (2024)
by: Hamdani, Rajaa El, et al.
Published: (2024)
Language Models and Logic Programs for Trustworthy Tax Reasoning
by: Jurayj, William, et al.
Published: (2025)
by: Jurayj, William, et al.
Published: (2025)
Can AI expose tax loopholes? Towards a new generation of legal policy assistants
by: Fratrič, Peter, et al.
Published: (2025)
by: Fratrič, Peter, et al.
Published: (2025)
AgentsCourt: Building Judicial Decision-Making Agents with Court Debate Simulation and Legal Knowledge Augmentation
by: He, Zhitao, et al.
Published: (2024)
by: He, Zhitao, et al.
Published: (2024)
LELA: An End-to-end LLM-based Entity Linking Framework with Zero-shot Domain Adaptation
by: Haffoudhi, Samy, et al.
Published: (2026)
by: Haffoudhi, Samy, et al.
Published: (2026)
Reasoning Fails Where Step Flow Breaks
by: Xu, Xiaoyu, et al.
Published: (2026)
by: Xu, Xiaoyu, et al.
Published: (2026)
LegalCiteBench: Evaluating Citation Reliability in Legal Language Models
by: Chen, Sijia, et al.
Published: (2026)
by: Chen, Sijia, et al.
Published: (2026)
GiusBERTo: A Legal Language Model for Personal Data De-identification in Italian Court of Auditors Decisions
by: Salierno, Giulio, et al.
Published: (2024)
by: Salierno, Giulio, et al.
Published: (2024)
Why do Experts Disagree on Existential Risk and P(doom)? A Survey of AI Experts
by: Field, Severin
Published: (2025)
by: Field, Severin
Published: (2025)
Where LLM Agents Fail and How They can Learn From Failures
by: Zhu, Kunlun, et al.
Published: (2025)
by: Zhu, Kunlun, et al.
Published: (2025)
Reframing Tax Law Entailment as Analogical Reasoning
by: Zou, Xinrui, et al.
Published: (2024)
by: Zou, Xinrui, et al.
Published: (2024)
OpenAI Cribbed Our Tax Example, But Can GPT-4 Really Do Tax?
by: Blair-Stanek, Andrew, et al.
Published: (2023)
by: Blair-Stanek, Andrew, et al.
Published: (2023)
Retrieval-Constrained Decoding Reveals Underestimated Parametric Knowledge in Language Models
by: Hamdani, Rajaa El, et al.
Published: (2025)
by: Hamdani, Rajaa El, et al.
Published: (2025)
Where Do AI Coding Agents Fail? An Empirical Study of Failed Agentic Pull Requests in GitHub
by: Ehsani, Ramtin, et al.
Published: (2026)
by: Ehsani, Ramtin, et al.
Published: (2026)
From Legal Text to Executable Decision Models: Evaluating Structured Representations for Legal Decision Model Generation
by: Graus, David
Published: (2026)
by: Graus, David
Published: (2026)
Can LLMs Identify Tax Abuse?
by: Blair-Stanek, Andrew, et al.
Published: (2025)
by: Blair-Stanek, Andrew, et al.
Published: (2025)
Robust Markov Decision Processes: A Place Where AI and Formal Methods Meet
by: Suilen, Marnix, et al.
Published: (2024)
by: Suilen, Marnix, et al.
Published: (2024)
From Citations to Criticality: Predicting Legal Decision Influence in the Multilingual Swiss Jurisprudence
by: Stern, Ronja, et al.
Published: (2024)
by: Stern, Ronja, et al.
Published: (2024)
Are LLMs Court-Ready? Evaluating Frontier Models on Indian Legal Reasoning
by: Juvekar, Kush, et al.
Published: (2025)
by: Juvekar, Kush, et al.
Published: (2025)
Logical Consistency Between Disagreeing Experts and Its Role in AI Safety
by: Corrada-Emmanuel, Andrés
Published: (2025)
by: Corrada-Emmanuel, Andrés
Published: (2025)
CourtPressGER: A German Court Decision to Press Release Summarization Dataset
by: Nagl, Sebastian, et al.
Published: (2025)
by: Nagl, Sebastian, et al.
Published: (2025)
Where and What: Reasoning Dynamic and Implicit Preferences in Situated Conversational Recommendation
by: Lin, Dongding, et al.
Published: (2026)
by: Lin, Dongding, et al.
Published: (2026)
Hypothesis Class Determines Explanation: Why Accurate Models Disagree on Feature Attribution
by: B, Thackshanaramana
Published: (2026)
by: B, Thackshanaramana
Published: (2026)
Why We Need World Models for AGI: Where LLMs Fail and How World Models May Outperform
by: Alaswad, Feisal, et al.
Published: (2026)
by: Alaswad, Feisal, et al.
Published: (2026)
Where Fake Citations Are Made: Tracing Field-Level Hallucination to Specific Neurons in LLMs
by: Chen, Yuefei, et al.
Published: (2026)
by: Chen, Yuefei, et al.
Published: (2026)
When Agents Disagree With Themselves: Measuring Behavioral Consistency in LLM-Based Agents
by: Mehta, Aman
Published: (2026)
by: Mehta, Aman
Published: (2026)
"Let's Agree to Disagree": Investigating the Disagreement Problem in Explainable AI for Text Summarization
by: Aswani, Seema, et al.
Published: (2024)
by: Aswani, Seema, et al.
Published: (2024)
Learning Who Disagrees: Demographic Importance Weighting for Modeling Annotator Distributions with DiADEM
by: Shetty, Samay U., et al.
Published: (2026)
by: Shetty, Samay U., et al.
Published: (2026)
When Audio and Text Disagree: Revealing Text Bias in Large Audio-Language Models
by: Wang, Cheng, et al.
Published: (2025)
by: Wang, Cheng, et al.
Published: (2025)
Magazine Supply Optimization: a Case-study
by: Nguyen, Duong, et al.
Published: (2024)
by: Nguyen, Duong, et al.
Published: (2024)
Persuadability and LLMs as Legal Decision Tools
by: Suttle, Oisin, et al.
Published: (2026)
by: Suttle, Oisin, et al.
Published: (2026)
Disagree and Commit: Degrees of Argumentation-based Agreements
by: Kampik, Timotheus, et al.
Published: (2024)
by: Kampik, Timotheus, et al.
Published: (2024)
When Models Disagree: Rethinking LLM Evaluation for Public Comment Analysis
by: Najera, Aisha, et al.
Published: (2026)
by: Najera, Aisha, et al.
Published: (2026)
When Documents Disagree: Measuring Institutional Variation in Transplant Guidance with Retrieval-Augmented Language Models
by: Li, Yubo, et al.
Published: (2026)
by: Li, Yubo, et al.
Published: (2026)
Understanding Adversarial Transfer: Why Representation-Space Attacks Fail Where Data-Space Attacks Succeed
by: Gupta, Isha, et al.
Published: (2025)
by: Gupta, Isha, et al.
Published: (2025)
XAI-LAW: A Logic Programming Tool for Modeling, Explaining, and Learning Legal Decisions
by: Dovier, Agostino, et al.
Published: (2026)
by: Dovier, Agostino, et al.
Published: (2026)
Why Chain of Thought Fails in Clinical Text Understanding
by: Wu, Jiageng, et al.
Published: (2025)
by: Wu, Jiageng, et al.
Published: (2025)
Assessing the Reliability of Large Language Models in the Bengali Legal Context: A Comparative Evaluation Using LLM-as-Judge and Legal Experts
by: Aftahee, Sabik, et al.
Published: (2025)
by: Aftahee, Sabik, et al.
Published: (2025)
Similar Items
-
Cultural Binding Heads in Language Models
by: Floro, Avrile, et al.
Published: (2026) -
BLT: Can Large Language Models Handle Basic Legal Text?
by: Blair-Stanek, Andrew, et al.
Published: (2023) -
The Factuality of Large Language Models in the Legal Domain
by: Hamdani, Rajaa El, et al.
Published: (2024) -
Language Models and Logic Programs for Trustworthy Tax Reasoning
by: Jurayj, William, et al.
Published: (2025) -
Can AI expose tax loopholes? Towards a new generation of legal policy assistants
by: Fratrič, Peter, et al.
Published: (2025)