SmartEval: A Benchmark for Evaluating LLM-Generated Smart Contracts from Natural Language Specifications
Fuente:
arXiv
Saved in:
| Main Authors: | Goel, Abhinav, Capponi, Agostino, Gliozzo, Alfio, Shah, Chaitya |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An end-to-end agentic pipeline for smart contract translation and quality evaluation
by: Goel, Abhinav, et al.
Published: (2026)
by: Goel, Abhinav, et al.
Published: (2026)
3D Topological Modeling and Multi-Agent Movement Simulation for Viral Infection Risk Analysis
by: Jabi, Wassim, et al.
Published: (2024)
by: Jabi, Wassim, et al.
Published: (2024)
ToolRosella: Translating Code Repositories into Standardized Tools for Scientific Agents
by: Di, Shimin, et al.
Published: (2026)
by: Di, Shimin, et al.
Published: (2026)
Securing Smart Contract Languages with a Unified Agentic Framework for Vulnerability Repair in Solidity and Move
by: Karanjai, Rabimba, et al.
Published: (2025)
by: Karanjai, Rabimba, et al.
Published: (2025)
$λ_A$: A Typed Lambda Calculus for LLM Agent Composition
by: Liu, Qin
Published: (2026)
by: Liu, Qin
Published: (2026)
Agent Behavioral Contracts: Formal Specification and Runtime Enforcement for Reliable Autonomous AI Agents
by: Bhardwaj, Varun Pratap
Published: (2026)
by: Bhardwaj, Varun Pratap
Published: (2026)
Object-Spatial Programming
by: Mars, Jason
Published: (2025)
by: Mars, Jason
Published: (2025)
Enhancing Holonic Architecture with Natural Language Processing for System of Systems
by: Ashfaq, Muhammad, et al.
Published: (2024)
by: Ashfaq, Muhammad, et al.
Published: (2024)
Deterministic vs. LLM-Controlled Orchestration for COBOL-to-Python Modernization
by: Lwin, Naing Oo, et al.
Published: (2026)
by: Lwin, Naing Oo, et al.
Published: (2026)
LongCLI-Bench: A Preliminary Benchmark and Study for Long-horizon Agentic Programming in Command-Line Interfaces
by: Feng, Yukang, et al.
Published: (2026)
by: Feng, Yukang, et al.
Published: (2026)
Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering
by: Zhou, Chenyu, et al.
Published: (2026)
by: Zhou, Chenyu, et al.
Published: (2026)
Analyzing Code Injection Attacks on LLM-based Multi-Agent Systems in Software Development
by: Bowers, Brian, et al.
Published: (2025)
by: Bowers, Brian, et al.
Published: (2025)
Weaving the Cosmos: WASM-Powered Interchain Communication for AI Enabled Smart Contracts
by: Karanjai, Rabimba, et al.
Published: (2025)
by: Karanjai, Rabimba, et al.
Published: (2025)
SWE-WebDevBench: Evaluating Coding Agent Application Platforms as Virtual Software Agencies
by: Saxena, Siddhant, et al.
Published: (2026)
by: Saxena, Siddhant, et al.
Published: (2026)
Runtime Composition in Dynamic System of Systems: A Systematic Review of Challenges, Solutions, Tools, and Evaluation Methods
by: Ashfaq, Muhammad, et al.
Published: (2025)
by: Ashfaq, Muhammad, et al.
Published: (2025)
Cognitive Agents Powered by Large Language Models for Agile Software Project Management
by: Cinkusz, Konrad, et al.
Published: (2025)
by: Cinkusz, Konrad, et al.
Published: (2025)
A Generic Modelling Framework for Last-Mile Delivery Systems
by: Gürcan, Önder, et al.
Published: (2025)
by: Gürcan, Önder, et al.
Published: (2025)
AutoFSM: A Multi-agent Framework for FSM Code Generation with IR and SystemC-Based Testing
by: Luo, Qiuming, et al.
Published: (2025)
by: Luo, Qiuming, et al.
Published: (2025)
HEAS: Hierarchical Evolutionary Agent-Based Simulation Framework for Multi-Objective Policy Search
by: Zhang, Ruiyu, et al.
Published: (2025)
by: Zhang, Ruiyu, et al.
Published: (2025)
The Lifecycle Workbench -- A Configurable Framework for Digitized Product Maintenance Services
by: Briechle, Dominique, et al.
Published: (2025)
by: Briechle, Dominique, et al.
Published: (2025)
Bridging the Prototype-Production Gap: A Multi-Agent System for Notebooks Transformation
by: Elhashemy, Hanya, et al.
Published: (2025)
by: Elhashemy, Hanya, et al.
Published: (2025)
ABMax: A JAX-based Agent-based Modeling Framework
by: Chaturvedi, Siddharth, et al.
Published: (2025)
by: Chaturvedi, Siddharth, et al.
Published: (2025)
CodeCureAgent: Automatic Classification and Repair of Static Analysis Warnings
by: Joos, Pascal, et al.
Published: (2025)
by: Joos, Pascal, et al.
Published: (2025)
REFINE: Enhancing Program Repair Agents through Context-Aware Patch Refinement
by: Pabba, Anvith, et al.
Published: (2025)
by: Pabba, Anvith, et al.
Published: (2025)
Sherlock: Reliable and Efficient Agentic Workflow Execution
by: Ro, Yeonju, et al.
Published: (2025)
by: Ro, Yeonju, et al.
Published: (2025)
Real-Time BDI Agents: a model and its implementation
by: Traldi, Andrea, et al.
Published: (2022)
by: Traldi, Andrea, et al.
Published: (2022)
Fairness in Multi-Agent Systems for Software Engineering: An SDLC-Oriented Rapid Review
by: Yang-Smith, Corey, et al.
Published: (2026)
by: Yang-Smith, Corey, et al.
Published: (2026)
Gecko: A Simulation Environment with Stateful Feedback for Refining Agent Tool Calls
by: Zhang, Zeyu, et al.
Published: (2026)
by: Zhang, Zeyu, et al.
Published: (2026)
The Hidden Bloat in Machine Learning Systems
by: Zhang, Huaifeng, et al.
Published: (2025)
by: Zhang, Huaifeng, et al.
Published: (2025)
A Step Towards a Universal Method for Modeling and Implementing Cross-Organizational Business Processes
by: Zeisler, Gerhard, et al.
Published: (2024)
by: Zeisler, Gerhard, et al.
Published: (2024)
The Specification Gap: Coordination Failure Under Partial Knowledge in Code Agents
by: Sartori, Camilo Chacón
Published: (2026)
by: Sartori, Camilo Chacón
Published: (2026)
AI-Generated Code Is Not Reproducible (Yet): An Empirical Study of Dependency Gaps in LLM-Based Coding Agents
by: Vangala, Bhanu Prakash, et al.
Published: (2025)
by: Vangala, Bhanu Prakash, et al.
Published: (2025)
SPEAR: An Engineering Case Study of Multi-Agent Coordination for Smart Contract Auditing
by: Chebolu, Indraveni, et al.
Published: (2026)
by: Chebolu, Indraveni, et al.
Published: (2026)
An Executable Benchmarking Suite for Tool-Using Agents
by: Zhong, Zhiqing, et al.
Published: (2026)
by: Zhong, Zhiqing, et al.
Published: (2026)
TrustTrade: Human-Inspired Selective Consensus Reduces Decision Uncertainty in LLM Trading Agents
by: Li, Minghan, et al.
Published: (2026)
by: Li, Minghan, et al.
Published: (2026)
Multi-Agent LLM Committees for Autonomous Software Beta Testing
by: Karanam, Sumanth Bharadwaj Hachalli, et al.
Published: (2025)
by: Karanam, Sumanth Bharadwaj Hachalli, et al.
Published: (2025)
LDP: An Identity-Aware Protocol for Multi-Agent LLM Systems
by: Prakash, Sunil
Published: (2026)
by: Prakash, Sunil
Published: (2026)
RIVA: Leveraging LLM Agents for Reliable Configuration Drift Detection
by: Abuzakuk, Sami, et al.
Published: (2026)
by: Abuzakuk, Sami, et al.
Published: (2026)
Leveraging Large Language Models for Institutional Portfolio Management: Persona-Based Ensembles
by: Abe, Yoshia, et al.
Published: (2024)
by: Abe, Yoshia, et al.
Published: (2024)
StockSim: A Dual-Mode Order-Level Simulator for Evaluating Multi-Agent LLMs in Financial Markets
by: Papadakis, Charidimos, et al.
Published: (2025)
by: Papadakis, Charidimos, et al.
Published: (2025)
Similar Items
-
An end-to-end agentic pipeline for smart contract translation and quality evaluation
by: Goel, Abhinav, et al.
Published: (2026) -
3D Topological Modeling and Multi-Agent Movement Simulation for Viral Infection Risk Analysis
by: Jabi, Wassim, et al.
Published: (2024) -
ToolRosella: Translating Code Repositories into Standardized Tools for Scientific Agents
by: Di, Shimin, et al.
Published: (2026) -
Securing Smart Contract Languages with a Unified Agentic Framework for Vulnerability Repair in Solidity and Move
by: Karanjai, Rabimba, et al.
Published: (2025) -
$λ_A$: A Typed Lambda Calculus for LLM Agent Composition
by: Liu, Qin
Published: (2026)