ContractEval: Benchmarking LLMs for Clause-Level Legal Risk Identification in Commercial Contracts
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Shuang, Li, Zelong, Ma, Ruoyun, Zhao, Haiyan, Du, Mengnan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ContractEval: A Benchmark for Evaluating Contract-Satisfying Assertions in Code Generation
by: Lim, Soohan, et al.
Published: (2025)
by: Lim, Soohan, et al.
Published: (2025)
LLM Agents in Law: Taxonomy, Applications, and Challenges
by: Liu, Shuang, et al.
Published: (2026)
by: Liu, Shuang, et al.
Published: (2026)
A Survey of Classification Tasks and Approaches for Legal Contracts
by: Singh, Amrita, et al.
Published: (2025)
by: Singh, Amrita, et al.
Published: (2025)
Exploring Multilingual Probing in Large Language Models: A Cross-Language Analysis
by: Li, Daoyang, et al.
Published: (2024)
by: Li, Daoyang, et al.
Published: (2024)
Beyond Single Concept Vector: Modeling Concept Subspace in LLMs with Gaussian Distribution
by: Zhao, Haiyan, et al.
Published: (2024)
by: Zhao, Haiyan, et al.
Published: (2024)
SCALM: Detecting Bad Practices in Smart Contracts Through LLMs
by: Li, Zongwei, et al.
Published: (2025)
by: Li, Zongwei, et al.
Published: (2025)
Legal Compliance Evaluation of Smart Contracts Generated By Large Language Models
by: Wijayakoon, Chanuka, et al.
Published: (2025)
by: Wijayakoon, Chanuka, et al.
Published: (2025)
LAW: Legal Agentic Workflows for Custody and Fund Services Contracts
by: Watson, William, et al.
Published: (2024)
by: Watson, William, et al.
Published: (2024)
Attracting Commercial Artificial Intelligence Firms to Support National Security through Collaborative Contracts
by: Bowne, Andrew
Published: (2025)
by: Bowne, Andrew
Published: (2025)
Contract-Coding: Towards Repo-Level Generation via Structured Symbolic Paradigm
by: Lin, Yi, et al.
Published: (2026)
by: Lin, Yi, et al.
Published: (2026)
Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders
by: Shu, Dong, et al.
Published: (2025)
by: Shu, Dong, et al.
Published: (2025)
Teaching Machines to Code: Smart Contract Translation with LLMs
by: Karanjai, Rabimba, et al.
Published: (2024)
by: Karanjai, Rabimba, et al.
Published: (2024)
Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering
by: Zhao, Haiyan, et al.
Published: (2025)
by: Zhao, Haiyan, et al.
Published: (2025)
Logic Meets Magic: LLMs Cracking Smart Contract Vulnerabilities
by: Xiao, ZeKe, et al.
Published: (2025)
by: Xiao, ZeKe, et al.
Published: (2025)
HLS-Eval: A Benchmark and Framework for Evaluating LLMs on High-Level Synthesis Design Tasks
by: Abi-Karam, Stefan, et al.
Published: (2025)
by: Abi-Karam, Stefan, et al.
Published: (2025)
CON-QA: Privacy-Preserving QA using cloud LLMs in Contract Domain
by: Singh, Ajeet Kumar, et al.
Published: (2025)
by: Singh, Ajeet Kumar, et al.
Published: (2025)
Benchmarking Zero-Shot Reasoning Approaches for Error Detection in Solidity Smart Contracts
by: Sardenberg, Eduardo, et al.
Published: (2026)
by: Sardenberg, Eduardo, et al.
Published: (2026)
ContractSkill: Repairable Contract-Based Skills for Multimodal Web Agents
by: Lu, Zijian, et al.
Published: (2026)
by: Lu, Zijian, et al.
Published: (2026)
MM-Eval: A Hierarchical Benchmark for Modern Mongolian Evaluation in LLMs
by: Zhang, Mengyuan, et al.
Published: (2024)
by: Zhang, Mengyuan, et al.
Published: (2024)
Towards Secure Program Partitioning for Smart Contracts with LLM's In-Context Learning
by: Liu, Ye, et al.
Published: (2025)
by: Liu, Ye, et al.
Published: (2025)
GAIus: Combining Genai with Legal Clauses Retrieval for Knowledge-based Assistant
by: Matak, Michał, et al.
Published: (2025)
by: Matak, Michał, et al.
Published: (2025)
SmartEval: A Benchmark for Evaluating LLM-Generated Smart Contracts from Natural Language Specifications
by: Goel, Abhinav, et al.
Published: (2026)
by: Goel, Abhinav, et al.
Published: (2026)
Accessible Smart Contracts Verification: Synthesizing Formal Models with Tamed LLMs
by: Corazza, Jan, et al.
Published: (2025)
by: Corazza, Jan, et al.
Published: (2025)
MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs
by: Zhao, Chenchen, et al.
Published: (2025)
by: Zhao, Chenchen, et al.
Published: (2025)
PARCER as an Operational Contract to Reduce Variance, Cost, and Risk in LLM Systems
by: Filho, Elzo Brito dos Santos
Published: (2026)
by: Filho, Elzo Brito dos Santos
Published: (2026)
SolContractEval: A Benchmark for Evaluating Contract-Level Solidity Code Generation
by: Ye, Zhifan, et al.
Published: (2025)
by: Ye, Zhifan, et al.
Published: (2025)
Towards Automated Smart Contract Generation: Evaluation, Benchmarking, and Retrieval-Augmented Repair
by: Chen, Zaoyu, et al.
Published: (2025)
by: Chen, Zaoyu, et al.
Published: (2025)
ToolGate: Contract-Grounded and Verified Tool Execution for LLMs
by: Liu, Yanming, et al.
Published: (2026)
by: Liu, Yanming, et al.
Published: (2026)
SCDBench: A Benchmark for LLM-Based Smart Contract Decompilers
by: Qin, Kaihua, et al.
Published: (2026)
by: Qin, Kaihua, et al.
Published: (2026)
ArabLegalEval: A Multitask Benchmark for Assessing Arabic Legal Knowledge in Large Language Models
by: Hijazi, Faris, et al.
Published: (2024)
by: Hijazi, Faris, et al.
Published: (2024)
Contraction Actor-Critic: Contraction Metric-Guided Reinforcement Learning for Robust Path Tracking
by: Cho, Minjae, et al.
Published: (2025)
by: Cho, Minjae, et al.
Published: (2025)
GeoEval: Benchmark for Evaluating LLMs and Multi-Modal Models on Geometry Problem-Solving
by: Zhang, Jiaxin, et al.
Published: (2024)
by: Zhang, Jiaxin, et al.
Published: (2024)
AmbiGraph-Eval: Can LLMs Effectively Handle Ambiguous Graph Queries?
by: Tian, Yuchen, et al.
Published: (2025)
by: Tian, Yuchen, et al.
Published: (2025)
StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs
by: Yang, Jialin, et al.
Published: (2025)
by: Yang, Jialin, et al.
Published: (2025)
SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models
by: He, Zirui, et al.
Published: (2025)
by: He, Zirui, et al.
Published: (2025)
LegalLens: Leveraging LLMs for Legal Violation Identification in Unstructured Text
by: Bernsohn, Dor, et al.
Published: (2024)
by: Bernsohn, Dor, et al.
Published: (2024)
Evaluation Ethics of LLMs in Legal Domain
by: Zhang, Ruizhe, et al.
Published: (2024)
by: Zhang, Ruizhe, et al.
Published: (2024)
Logical foundations of Smart Contracts
by: Kalala, Kalonji
Published: (2025)
by: Kalala, Kalonji
Published: (2025)
Neural Contractive Dynamical Systems
by: Beik-Mohammadi, Hadi, et al.
Published: (2024)
by: Beik-Mohammadi, Hadi, et al.
Published: (2024)
VoiceAgentEval: A Dual-Dimensional Benchmark for Expert-Level Intelligent Voice-Agent Evaluation of Xbench's Professional-Aligned Series
by: Xu, Pengyu, et al.
Published: (2025)
by: Xu, Pengyu, et al.
Published: (2025)
Similar Items
-
ContractEval: A Benchmark for Evaluating Contract-Satisfying Assertions in Code Generation
by: Lim, Soohan, et al.
Published: (2025) -
LLM Agents in Law: Taxonomy, Applications, and Challenges
by: Liu, Shuang, et al.
Published: (2026) -
A Survey of Classification Tasks and Approaches for Legal Contracts
by: Singh, Amrita, et al.
Published: (2025) -
Exploring Multilingual Probing in Large Language Models: A Cross-Language Analysis
by: Li, Daoyang, et al.
Published: (2024) -
Beyond Single Concept Vector: Modeling Concept Subspace in LLMs with Gaussian Distribution
by: Zhao, Haiyan, et al.
Published: (2024)