The Fusion of Large Language Models and Formal Methods for Trustworthy AI Agents: A Roadmap
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Yedi, Cai, Yufan, Zuo, Xinyue, Luan, Xiaokun, Wang, Kailong, Hou, Zhe, Zhang, Yifan, Wei, Zhiyuan, Sun, Meng, Sun, Jun, Sun, Jing, Dong, Jin Song |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PAT-Agent: Autoformalization for Model Checking
by: Zuo, Xinyue, et al.
Published: (2025)
by: Zuo, Xinyue, et al.
Published: (2025)
Towards Large Language Model Aided Program Refinement
by: Cai, Yufan, et al.
Published: (2024)
by: Cai, Yufan, et al.
Published: (2024)
Protecting Deep Learning Model Copyrights with Adversarial Example-Free Reuse Detection
by: Luan, Xiaokun, et al.
Published: (2024)
by: Luan, Xiaokun, et al.
Published: (2024)
Event-B Agent: Towards LLM Agent for Formal Model Synthesis and Repair
by: Wang, Hongshu, et al.
Published: (2026)
by: Wang, Hongshu, et al.
Published: (2026)
Beyond Correctness: Exposing LLM-generated Logical Flaws in Reasoning via Multi-step Automated Theorem Proving
by: Zheng, Xinyi, et al.
Published: (2025)
by: Zheng, Xinyi, et al.
Published: (2025)
Automata-Based Steering of Large Language Models for Diverse Structured Generation
by: Luan, Xiaokun, et al.
Published: (2025)
by: Luan, Xiaokun, et al.
Published: (2025)
ClawWorm: Self-Propagating Attacks Across LLM Agent Ecosystems
by: Zhang, Yihao, et al.
Published: (2026)
by: Zhang, Yihao, et al.
Published: (2026)
Towards Trustworthy Legal AI through LLM Agents and Formal Reasoning
by: Chen, Linze, et al.
Published: (2025)
by: Chen, Linze, et al.
Published: (2025)
LLM-enabled Applications Require System-Level Threat Monitoring
by: Zhang, Yedi, et al.
Published: (2026)
by: Zhang, Yedi, et al.
Published: (2026)
RACC: Representation-Aware Coverage Criteria for LLM Safety Testing
by: Wei, Zeming, et al.
Published: (2026)
by: Wei, Zeming, et al.
Published: (2026)
MaCTG: Multi-Agent Collaborative Thought Graph for Automatic Programming
by: Zhao, Zixiao, et al.
Published: (2024)
by: Zhao, Zixiao, et al.
Published: (2024)
Self-Organizing Multi-Agent Systems for Continuous Software Development
by: Lyu, Wenhan, et al.
Published: (2026)
by: Lyu, Wenhan, et al.
Published: (2026)
Reinforcement Learning with Negative Tests as Completeness Signal for Formal Specification Synthesis
by: Huang, Zhechong, et al.
Published: (2026)
by: Huang, Zhechong, et al.
Published: (2026)
Requirements Development and Formalization for Reliable Code Generation: A Multi-Agent Vision
by: Lu, Xu, et al.
Published: (2025)
by: Lu, Xu, et al.
Published: (2025)
Large Language Models are overconfident and amplify human bias
by: Sun, Fengfei, et al.
Published: (2025)
by: Sun, Fengfei, et al.
Published: (2025)
A Test Suite for Efficient Robustness Evaluation of Face Recognition Systems
by: Zhang, Ruihan, et al.
Published: (2025)
by: Zhang, Ruihan, et al.
Published: (2025)
REDriver: Runtime Enforcement for Autonomous Vehicles
by: Sun, Yang, et al.
Published: (2024)
by: Sun, Yang, et al.
Published: (2024)
Trustworthy AI Software Engineers
by: Aleti, Aldeida, et al.
Published: (2026)
by: Aleti, Aldeida, et al.
Published: (2026)
LLM App Store Analysis: A Vision and Roadmap
by: Zhao, Yanjie, et al.
Published: (2024)
by: Zhao, Yanjie, et al.
Published: (2024)
A Roadmap for Software Testing in Open Collaborative Development Environments
by: Wang, Qing, et al.
Published: (2024)
by: Wang, Qing, et al.
Published: (2024)
SceneGenAgent: Precise Industrial Scene Generation with Coding Agent
by: Xia, Xiao, et al.
Published: (2024)
by: Xia, Xiao, et al.
Published: (2024)
A Vulnerability Code Intent Summary Dataset
by: Huang, Yifan, et al.
Published: (2025)
by: Huang, Yifan, et al.
Published: (2025)
Iterative Experience Refinement of Software-Developing Agents
by: Qian, Chen, et al.
Published: (2024)
by: Qian, Chen, et al.
Published: (2024)
MuMuTestUp: Mutation-based Multi-Agent Test Case Update
by: Tian, Dawei, et al.
Published: (2026)
by: Tian, Dawei, et al.
Published: (2026)
ACAV: A Framework for Automatic Causality Analysis in Autonomous Vehicle Accident Recordings
by: Sun, Huijia, et al.
Published: (2024)
by: Sun, Huijia, et al.
Published: (2024)
Formal Architecture Descriptors as Navigation Primitives for AI Coding Agents
by: Jin, Ruoqi
Published: (2026)
by: Jin, Ruoqi
Published: (2026)
Formalizing UML State Machines for Automated Verification -- A Survey
by: André, Étienne, et al.
Published: (2024)
by: André, Étienne, et al.
Published: (2024)
Accountability of Robust and Reliable AI-Enabled Systems: A Preliminary Study and Roadmap
by: Scaramuzza, Filippo, et al.
Published: (2025)
by: Scaramuzza, Filippo, et al.
Published: (2025)
Predicting Developer Acceptance of AI-Generated Code Suggestions
by: Jiang, Jing, et al.
Published: (2026)
by: Jiang, Jing, et al.
Published: (2026)
Knowledge-Based Multi-Agent Framework for Automated Software Architecture Design
by: Zhang, Yiran, et al.
Published: (2025)
by: Zhang, Yiran, et al.
Published: (2025)
Co-Evolution of Types and Dependencies: Towards Repository-Level Type Inference for Python Code
by: Sun, Shuo, et al.
Published: (2025)
by: Sun, Shuo, et al.
Published: (2025)
Co-Saving: Resource Aware Multi-Agent Collaboration for Software Development
by: Qiu, Rennai, et al.
Published: (2025)
by: Qiu, Rennai, et al.
Published: (2025)
Towards a Framework for Operationalizing the Specification of Trustworthy AI Requirements
by: Villamizar, Hugo, et al.
Published: (2025)
by: Villamizar, Hugo, et al.
Published: (2025)
POLARIS: A framework to guide the development of Trustworthy AI systems
by: Baldassarre, Maria Teresa, et al.
Published: (2024)
by: Baldassarre, Maria Teresa, et al.
Published: (2024)
The Cream Rises to the Top: Efficient Reranking Method for Verilog Code Generation
by: Yang, Guang, et al.
Published: (2025)
by: Yang, Guang, et al.
Published: (2025)
FixDrive: Automatically Repairing Autonomous Vehicle Driving Behaviour for $0.08 per Violation
by: Sun, Yang, et al.
Published: (2025)
by: Sun, Yang, et al.
Published: (2025)
LLM for Mobile: An Initial Roadmap
by: Chen, Daihang, et al.
Published: (2024)
by: Chen, Daihang, et al.
Published: (2024)
Evolaris: A Roadmap to Self-Evolving Software Intelligence Management
by: Liu, Chengwei, et al.
Published: (2025)
by: Liu, Chengwei, et al.
Published: (2025)
Shepherd: A Runtime Substrate Empowering Meta-Agents with a Formalized Execution Trace
by: Yu, Simon, et al.
Published: (2026)
by: Yu, Simon, et al.
Published: (2026)
Towards Reliable Vector Database Management Systems: A Software Testing Roadmap for 2030
by: Wang, Shenao, et al.
Published: (2025)
by: Wang, Shenao, et al.
Published: (2025)
Similar Items
-
PAT-Agent: Autoformalization for Model Checking
by: Zuo, Xinyue, et al.
Published: (2025) -
Towards Large Language Model Aided Program Refinement
by: Cai, Yufan, et al.
Published: (2024) -
Protecting Deep Learning Model Copyrights with Adversarial Example-Free Reuse Detection
by: Luan, Xiaokun, et al.
Published: (2024) -
Event-B Agent: Towards LLM Agent for Formal Model Synthesis and Repair
by: Wang, Hongshu, et al.
Published: (2026) -
Beyond Correctness: Exposing LLM-generated Logical Flaws in Reasoning via Multi-step Automated Theorem Proving
by: Zheng, Xinyi, et al.
Published: (2025)