MANTRA: Synthesizing SMT-Validated Compliance Benchmarks for Tool-Using LLM Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Anand, Ashwani, Chatzi, Ivi, Raha, Ritam, Schmuck, Anne-Kathrin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Follow the STARs: Dynamic $ω$-Regular Shielding of Learned Policies
von: Anand, Ashwani, et al.
Veröffentlicht: (2025)
von: Anand, Ashwani, et al.
Veröffentlicht: (2025)
Neuro-Symbolic Compliance: Integrating LLMs and SMT Solvers for Automated Financial Legal Analysis
von: Hsia, Yung-Shen, et al.
Veröffentlicht: (2026)
von: Hsia, Yung-Shen, et al.
Veröffentlicht: (2026)
Loop Invariant Generation: A Hybrid Framework of Reasoning optimised LLMs and SMT Solvers
von: Bharti, Varun, et al.
Veröffentlicht: (2025)
von: Bharti, Varun, et al.
Veröffentlicht: (2025)
Quantitative Strategy Templates
von: Anand, Ashwani, et al.
Veröffentlicht: (2025)
von: Anand, Ashwani, et al.
Veröffentlicht: (2025)
Fair Quantitative Games
von: Anand, Ashwani, et al.
Veröffentlicht: (2025)
von: Anand, Ashwani, et al.
Veröffentlicht: (2025)
Towards a Certified Proof Checker for Deep Neural Network Verification
von: Desmartin, Remi, et al.
Veröffentlicht: (2023)
von: Desmartin, Remi, et al.
Veröffentlicht: (2023)
Synthesizing Permissive Winning Strategy Templates for Parity Games
von: Anand, Ashwani, et al.
Veröffentlicht: (2023)
von: Anand, Ashwani, et al.
Veröffentlicht: (2023)
An Encoding for CLP Problems in SMT-LIB
von: Amrollahi, Daneshvar, et al.
Veröffentlicht: (2024)
von: Amrollahi, Daneshvar, et al.
Veröffentlicht: (2024)
Geometry of Reason: Spectral Signatures of Valid Mathematical Reasoning
von: Noël, Valentin
Veröffentlicht: (2026)
von: Noël, Valentin
Veröffentlicht: (2026)
Localized Attractor Computations for Infinite-State Games (Full Version)
von: Schmuck, Anne-Kathrin, et al.
Veröffentlicht: (2024)
von: Schmuck, Anne-Kathrin, et al.
Veröffentlicht: (2024)
Universal Safety Controllers with Learned Prophecies
von: Finkbeiner, Bernd, et al.
Veröffentlicht: (2025)
von: Finkbeiner, Bernd, et al.
Veröffentlicht: (2025)
Synthesis of Universal Safety Controllers
von: Finkbeiner, Bernd, et al.
Veröffentlicht: (2025)
von: Finkbeiner, Bernd, et al.
Veröffentlicht: (2025)
Bolzano: Case Studies in LLM-Assisted Mathematical Research
von: Balko, Martin, et al.
Veröffentlicht: (2026)
von: Balko, Martin, et al.
Veröffentlicht: (2026)
JustLogic: A Comprehensive Benchmark for Evaluating Deductive Reasoning in Large Language Models
von: Chen, Michael K., et al.
Veröffentlicht: (2025)
von: Chen, Michael K., et al.
Veröffentlicht: (2025)
SATBench: Benchmarking LLMs' Logical Reasoning via Automated Puzzle Generation from SAT Formulas
von: Wei, Anjiang, et al.
Veröffentlicht: (2025)
von: Wei, Anjiang, et al.
Veröffentlicht: (2025)
LLM2SMT: Building an SMT Solver with Zero Human-Written Code
von: Janota, Mikoláš, et al.
Veröffentlicht: (2026)
von: Janota, Mikoláš, et al.
Veröffentlicht: (2026)
Hybrid Models for Natural Language Reasoning: The Case of Syllogistic Logic
von: Guzmán, Manuel Vargas, et al.
Veröffentlicht: (2025)
von: Guzmán, Manuel Vargas, et al.
Veröffentlicht: (2025)
Process discovery on deviant traces and other stranger things
von: Chesani, Federico, et al.
Veröffentlicht: (2021)
von: Chesani, Federico, et al.
Veröffentlicht: (2021)
Process-Driven Autoformalization in Lean 4
von: Lu, Jianqiao, et al.
Veröffentlicht: (2024)
von: Lu, Jianqiao, et al.
Veröffentlicht: (2024)
Interchangeable Token Embeddings for Extendable Vocabulary and Alpha-Equivalence
von: Işık, İlker, et al.
Veröffentlicht: (2024)
von: Işık, İlker, et al.
Veröffentlicht: (2024)
Putnam 2025 Problems in Rocq using Opus 4.6 and Rocq-MCP
von: Baudart, Guillaume, et al.
Veröffentlicht: (2026)
von: Baudart, Guillaume, et al.
Veröffentlicht: (2026)
Learning to Estimate System Specifications in Linear Temporal Logic using Transformers and Mamba
von: Işık, İlker, et al.
Veröffentlicht: (2024)
von: Işık, İlker, et al.
Veröffentlicht: (2024)
About Time: Model-free Reinforcement Learning with Timed Reward Machines
von: Roy, Rajarshi, et al.
Veröffentlicht: (2025)
von: Roy, Rajarshi, et al.
Veröffentlicht: (2025)
Inference of Qualitative Models from Steady-State Data via Weighted MaxSMT
von: Huvar, Ondřej, et al.
Veröffentlicht: (2026)
von: Huvar, Ondřej, et al.
Veröffentlicht: (2026)
Tokenization Multiplicity Leads to Arbitrary Price Variation in LLM-as-a-service
von: Chatzi, Ivi, et al.
Veröffentlicht: (2025)
von: Chatzi, Ivi, et al.
Veröffentlicht: (2025)
'Put the Car on the Stand': SMT-based Oracles for Investigating Decisions
von: Judson, Samuel, et al.
Veröffentlicht: (2023)
von: Judson, Samuel, et al.
Veröffentlicht: (2023)
MiniF2F in Rocq: Automatic Translation Between Proof Assistants -- A Case Study
von: Viennot, Jules, et al.
Veröffentlicht: (2025)
von: Viennot, Jules, et al.
Veröffentlicht: (2025)
An In-Context Learning Agent for Formal Theorem-Proving
von: Thakur, Amitayush, et al.
Veröffentlicht: (2023)
von: Thakur, Amitayush, et al.
Veröffentlicht: (2023)
The Expressive Power of Transformers with Chain of Thought
von: Merrill, William, et al.
Veröffentlicht: (2023)
von: Merrill, William, et al.
Veröffentlicht: (2023)
Lean-SMT: An SMT tactic for discharging proof goals in Lean
von: Mohamed, Abdalrhman, et al.
Veröffentlicht: (2025)
von: Mohamed, Abdalrhman, et al.
Veröffentlicht: (2025)
Boosting Few-Pixel Robustness Verification via Covering Verification Designs
von: Shapira, Yuval, et al.
Veröffentlicht: (2024)
von: Shapira, Yuval, et al.
Veröffentlicht: (2024)
Smart Choices and the Selection Monad
von: Abadi, Martin, et al.
Veröffentlicht: (2020)
von: Abadi, Martin, et al.
Veröffentlicht: (2020)
Mini-Batch Robustness Verification of Deep Neural Networks
von: Tzour-Shaday, Saar, et al.
Veröffentlicht: (2025)
von: Tzour-Shaday, Saar, et al.
Veröffentlicht: (2025)
Programmatic Reinforcement Learning: Navigating Gridworlds
von: Shabadi, Guruprerana, et al.
Veröffentlicht: (2024)
von: Shabadi, Guruprerana, et al.
Veröffentlicht: (2024)
Floating-Point Neural Networks Are Provably Robust Universal Approximators
von: Hwang, Geonho, et al.
Veröffentlicht: (2025)
von: Hwang, Geonho, et al.
Veröffentlicht: (2025)
Neural Network Verification is a Programming Language Challenge
von: Cordeiro, Lucas C., et al.
Veröffentlicht: (2025)
von: Cordeiro, Lucas C., et al.
Veröffentlicht: (2025)
Probabilistic unifying relations for modelling epistemic and aleatoric uncertainty: semantics and automated reasoning with theorem proving
von: Ye, Kangfeng, et al.
Veröffentlicht: (2023)
von: Ye, Kangfeng, et al.
Veröffentlicht: (2023)
Evaluating Implicit Regulatory Compliance in LLM Tool Invocation via Logic-Guided Synthesis
von: Song, Da, et al.
Veröffentlicht: (2026)
von: Song, Da, et al.
Veröffentlicht: (2026)
Transformers Can Learn Connectivity in Some Graphs but Not Others
von: Roy, Amit, et al.
Veröffentlicht: (2025)
von: Roy, Amit, et al.
Veröffentlicht: (2025)
Conflict-Aware Fusion: Mitigating Logic Inertia in Large Language Models via Structured Cognitive Priors
von: Bao, Qiming, et al.
Veröffentlicht: (2025)
von: Bao, Qiming, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Follow the STARs: Dynamic $ω$-Regular Shielding of Learned Policies
von: Anand, Ashwani, et al.
Veröffentlicht: (2025) -
Neuro-Symbolic Compliance: Integrating LLMs and SMT Solvers for Automated Financial Legal Analysis
von: Hsia, Yung-Shen, et al.
Veröffentlicht: (2026) -
Loop Invariant Generation: A Hybrid Framework of Reasoning optimised LLMs and SMT Solvers
von: Bharti, Varun, et al.
Veröffentlicht: (2025) -
Quantitative Strategy Templates
von: Anand, Ashwani, et al.
Veröffentlicht: (2025) -
Fair Quantitative Games
von: Anand, Ashwani, et al.
Veröffentlicht: (2025)