BLAST: Benchmarking LLMs with ASP-based Structured Testing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Santana, Manuel Alejandro Borroto, Coppolillo, Erica, Calimeri, Francesco, Manco, Giuseppe, Perri, Simona, Ricca, Francesco |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLASP: Fine-tuning Large Language Models for Answer Set Programming
von: Coppolillo, Erica, et al.
Veröffentlicht: (2024)
von: Coppolillo, Erica, et al.
Veröffentlicht: (2024)
2-ASP(Q) programs with weak constraints: Complexity and efficient implementation
von: Cuteri, Andrea, et al.
Veröffentlicht: (2026)
von: Cuteri, Andrea, et al.
Veröffentlicht: (2026)
Question Answering with LLMs and Learning from Answer Sets
von: Borroto, Manuel, et al.
Veröffentlicht: (2025)
von: Borroto, Manuel, et al.
Veröffentlicht: (2025)
Direct Encoding of Declare Constraints in ASP
von: Chiariello, Francesco, et al.
Veröffentlicht: (2024)
von: Chiariello, Francesco, et al.
Veröffentlicht: (2024)
ASP-based Multi-shot Reasoning via DLV2 with Incremental Grounding
von: Calimeri, Francesco, et al.
Veröffentlicht: (2024)
von: Calimeri, Francesco, et al.
Veröffentlicht: (2024)
Towards Automatic Composition of ASP Programs from Natural Language Specifications
von: Borroto, Manuel, et al.
Veröffentlicht: (2024)
von: Borroto, Manuel, et al.
Veröffentlicht: (2024)
Unmasking Conversational Bias in AI Multiagent Systems
von: Coppolillo, Erica, et al.
Veröffentlicht: (2025)
von: Coppolillo, Erica, et al.
Veröffentlicht: (2025)
Enumerating Minimal Unsatisfiable Cores of LTLf formulas
von: Ielo, Antonio, et al.
Veröffentlicht: (2024)
von: Ielo, Antonio, et al.
Veröffentlicht: (2024)
A Reliable Common-Sense Reasoning Socialbot Built Using LLMs and Goal-Directed ASP
von: Zeng, Yankai, et al.
Veröffentlicht: (2024)
von: Zeng, Yankai, et al.
Veröffentlicht: (2024)
ASP-Bench: From Natural Language to Logic Programs
von: Szeider, Stefan
Veröffentlicht: (2026)
von: Szeider, Stefan
Veröffentlicht: (2026)
Relevance meets Diversity: A User-Centric Framework for Knowledge Exploration through Recommendations
von: Coppolillo, Erica, et al.
Veröffentlicht: (2024)
von: Coppolillo, Erica, et al.
Veröffentlicht: (2024)
Unit Testing in ASP Revisited: Language and Test-Driven Development Environment
von: Amendola, Giovanni, et al.
Veröffentlicht: (2024)
von: Amendola, Giovanni, et al.
Veröffentlicht: (2024)
Improving ASP-based ORS Schedules through Machine Learning Predictions
von: Bruno, Pierangela, et al.
Veröffentlicht: (2025)
von: Bruno, Pierangela, et al.
Veröffentlicht: (2025)
Engagement-Driven Content Generation with Large Language Models
von: Coppolillo, Erica, et al.
Veröffentlicht: (2024)
von: Coppolillo, Erica, et al.
Veröffentlicht: (2024)
Looking Beyond Accuracy: A Holistic Benchmark of ECG Foundation Models
von: Filice, Francesca, et al.
Veröffentlicht: (2026)
von: Filice, Francesca, et al.
Veröffentlicht: (2026)
Inferring multiple helper Dafny assertions with LLMs
von: Silva, Álvaro, et al.
Veröffentlicht: (2025)
von: Silva, Álvaro, et al.
Veröffentlicht: (2025)
SpotIt+: Verification-based Text-to-SQL Evaluation with Database Constraints
von: Tremante, Andrew, et al.
Veröffentlicht: (2026)
von: Tremante, Andrew, et al.
Veröffentlicht: (2026)
s2n-bignum-bench: A practical benchmark for evaluating low-level code reasoning of LLMs
von: Rao, Balaji, et al.
Veröffentlicht: (2026)
von: Rao, Balaji, et al.
Veröffentlicht: (2026)
VeriThoughts: Enabling Automated Verilog Code Generation using Reasoning and Formal Verification
von: Yubeaton, Patrick, et al.
Veröffentlicht: (2025)
von: Yubeaton, Patrick, et al.
Veröffentlicht: (2025)
Integrating Reasoning Systems for Trustworthy AI, Proceedings of the 4th Workshop on Logic and Practice of Programming (LPOP)
von: Nerode, Anil, et al.
Veröffentlicht: (2024)
von: Nerode, Anil, et al.
Veröffentlicht: (2024)
Pearce's Characterisation in an Epistemic Domain
von: Su, Ezgi Iraz
Veröffentlicht: (2025)
von: Su, Ezgi Iraz
Veröffentlicht: (2025)
VEL: A Formally Verified Reasoner for OWL2 EL Profile
von: Ileri, Atalay Mert, et al.
Veröffentlicht: (2024)
von: Ileri, Atalay Mert, et al.
Veröffentlicht: (2024)
Relating Answer Set Programming and Many-sorted Logics for Formal Verification
von: Hansen, Zachary
Veröffentlicht: (2025)
von: Hansen, Zachary
Veröffentlicht: (2025)
Causal Past Logic for Runtime Verification of Distributed LLM Agent Workflows
von: Bollig, Benedikt
Veröffentlicht: (2026)
von: Bollig, Benedikt
Veröffentlicht: (2026)
Cobblestone: A Divide-and-Conquer Approach for Automating Formal Verification
von: Kasibatla, Saketh Ram, et al.
Veröffentlicht: (2024)
von: Kasibatla, Saketh Ram, et al.
Veröffentlicht: (2024)
A Neurosymbolic Approach to Loop Invariant Generation via Weakest Precondition Reasoning
von: King, Daragh, et al.
Veröffentlicht: (2025)
von: King, Daragh, et al.
Veröffentlicht: (2025)
A Reversible Semantics for Janus
von: Lanese, Ivan, et al.
Veröffentlicht: (2026)
von: Lanese, Ivan, et al.
Veröffentlicht: (2026)
A Certified Proof Checker for Deep Neural Network Verification in Imandra
von: Desmartin, Remi, et al.
Veröffentlicht: (2024)
von: Desmartin, Remi, et al.
Veröffentlicht: (2024)
Soda: An Object-Oriented Functional Language for Specifying Human-Centered Problems
von: Mendez, Julian Alfredo
Veröffentlicht: (2023)
von: Mendez, Julian Alfredo
Veröffentlicht: (2023)
Using ASP(Q) to Handle Inconsistent Prioritized Data
von: Bienvenu, Meghyn, et al.
Veröffentlicht: (2026)
von: Bienvenu, Meghyn, et al.
Veröffentlicht: (2026)
An ASP-based Solution to the Medical Appointment Scheduling Problem
von: Vozna, Alina, et al.
Veröffentlicht: (2026)
von: Vozna, Alina, et al.
Veröffentlicht: (2026)
An ASP-Based Framework for MUSes
von: Kabir, Mohimenul, et al.
Veröffentlicht: (2025)
von: Kabir, Mohimenul, et al.
Veröffentlicht: (2025)
VERINA: Benchmarking Verifiable Code Generation
von: Ye, Zhe, et al.
Veröffentlicht: (2025)
von: Ye, Zhe, et al.
Veröffentlicht: (2025)
Checkpoint-based rollback recovery in session programming
von: Mezzina, Claudio Antares, et al.
Veröffentlicht: (2023)
von: Mezzina, Claudio Antares, et al.
Veröffentlicht: (2023)
Transformer-Based Models Are Not Yet Perfect At Learning to Emulate Structural Recursion
von: Zhang, Dylan, et al.
Veröffentlicht: (2024)
von: Zhang, Dylan, et al.
Veröffentlicht: (2024)
CLEVER: A Curated Benchmark for Formally Verified Code Generation
von: Thakur, Amitayush, et al.
Veröffentlicht: (2025)
von: Thakur, Amitayush, et al.
Veröffentlicht: (2025)
Towards Mass Spectrum Analysis with ASP
von: Küchenmeister, Nils, et al.
Veröffentlicht: (2025)
von: Küchenmeister, Nils, et al.
Veröffentlicht: (2025)
From Blind Solvers to Logical Thinkers: Benchmarking LLMs' Logical Integrity on Faulty Mathematical Problems
von: Rahman, A M Muntasir, et al.
Veröffentlicht: (2024)
von: Rahman, A M Muntasir, et al.
Veröffentlicht: (2024)
SM-based Semantics for Answer Set Programs Containing Conditional Literals and Arithmetic
von: Hansen, Zachary, et al.
Veröffentlicht: (2025)
von: Hansen, Zachary, et al.
Veröffentlicht: (2025)
Diminution: On Reducing the Size of Grounding ASP Programs
von: Yang, HuanYu, et al.
Veröffentlicht: (2025)
von: Yang, HuanYu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LLASP: Fine-tuning Large Language Models for Answer Set Programming
von: Coppolillo, Erica, et al.
Veröffentlicht: (2024) -
2-ASP(Q) programs with weak constraints: Complexity and efficient implementation
von: Cuteri, Andrea, et al.
Veröffentlicht: (2026) -
Question Answering with LLMs and Learning from Answer Sets
von: Borroto, Manuel, et al.
Veröffentlicht: (2025) -
Direct Encoding of Declare Constraints in ASP
von: Chiariello, Francesco, et al.
Veröffentlicht: (2024) -
ASP-based Multi-shot Reasoning via DLV2 with Incremental Grounding
von: Calimeri, Francesco, et al.
Veröffentlicht: (2024)