$\forall$uto$\exists$val: Autonomous Assessment of LLMs in Formal Synthesis and Interpretation Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Karia, Rushang, Bramblett, Daniel, Dobhal, Daksh, Verma, Pulkit, Srivastava, Siddharth |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Autonomous Evaluation of LLMs for Truth Maintenance and Reasoning Tasks
by: Karia, Rushang, et al.
Published: (2024)
by: Karia, Rushang, et al.
Published: (2024)
Epistemic Exploration for Generalizable Planning and Learning in Non-Stationary Settings
by: Karia, Rushang, et al.
Published: (2024)
by: Karia, Rushang, et al.
Published: (2024)
Discovering and Learning Probabilistic Models of Black-Box AI Capabilities
by: Bramblett, Daniel, et al.
Published: (2025)
by: Bramblett, Daniel, et al.
Published: (2025)
Using Explainable AI and Hierarchical Planning for Outreach with Robots
by: Karia, Rushang, et al.
Published: (2024)
by: Karia, Rushang, et al.
Published: (2024)
Belief-State Query Policies for User-Aligned POMDPs
by: Bramblett, Daniel, et al.
Published: (2024)
by: Bramblett, Daniel, et al.
Published: (2024)
Teaching LLMs to Plan: Logical Chain-of-Thought Instruction Tuning for Symbolic Planning
by: Verma, Pulkit, et al.
Published: (2025)
by: Verma, Pulkit, et al.
Published: (2025)
Are LLMs good pragmatic speakers?
by: Jian, Mingyue, et al.
Published: (2024)
by: Jian, Mingyue, et al.
Published: (2024)
ITLC at SemEval-2026 Task 11: Normalization and Deterministic Parsing for Formal Reasoning in LLMs
by: Muhamad, Wicaksono Leksono, et al.
Published: (2026)
by: Muhamad, Wicaksono Leksono, et al.
Published: (2026)
On Conformant Planning and Model-Checking of $\exists^*\forall^*$ Hyperproperties
by: Beutner, Raven, et al.
Published: (2025)
by: Beutner, Raven, et al.
Published: (2025)
Large Language Models as Pokémon Battle Agents: Strategic Play and Content Generation
by: Jain, Daksh, et al.
Published: (2025)
by: Jain, Daksh, et al.
Published: (2025)
HALT-RAG: A Task-Adaptable Framework for Hallucination Detection with Calibrated NLI Ensembles and Abstention
by: Goswami, Saumya, et al.
Published: (2025)
by: Goswami, Saumya, et al.
Published: (2025)
Phonetic Perturbations Reveal Tokenizer-Rooted Safety Gaps in LLMs
by: Aswal, Darpan, et al.
Published: (2025)
by: Aswal, Darpan, et al.
Published: (2025)
Emergence and Effectiveness of Task Vectors in In-Context Learning: An Encoder Decoder Perspective
by: Han, Seungwook, et al.
Published: (2024)
by: Han, Seungwook, et al.
Published: (2024)
Leveraging Weighted Syntactic and Semantic Context Assessment Summary (wSSAS) Towards Text Categorization Using LLMs
by: Kathuria, Shreeya Verma, et al.
Published: (2026)
by: Kathuria, Shreeya Verma, et al.
Published: (2026)
Value Augmented Sampling for Language Model Alignment and Personalization
by: Han, Seungwook, et al.
Published: (2024)
by: Han, Seungwook, et al.
Published: (2024)
From Informal to Formal -- Incorporating and Evaluating LLMs on Natural Language Requirements to Verifiable Formal Proofs
by: Cao, Jialun, et al.
Published: (2025)
by: Cao, Jialun, et al.
Published: (2025)
Grammars of Formal Uncertainty: When to Trust LLMs in Automated Reasoning Tasks
by: Ganguly, Debargha, et al.
Published: (2025)
by: Ganguly, Debargha, et al.
Published: (2025)
LLMs in Interpreting Legal Documents
by: Corbo, Simone
Published: (2025)
by: Corbo, Simone
Published: (2025)
Towards Automated Fact-Checking of Real-World Claims: Exploring Task Formulation and Assessment with LLMs
by: Sahitaj, Premtim, et al.
Published: (2025)
by: Sahitaj, Premtim, et al.
Published: (2025)
Interpretability of Language Models via Task Spaces
by: Weber, Lucas, et al.
Published: (2024)
by: Weber, Lucas, et al.
Published: (2024)
PLANTS: A Novel Problem and Dataset for Summarization of Planning-Like (PL) Tasks
by: Pallagani, Vishal, et al.
Published: (2024)
by: Pallagani, Vishal, et al.
Published: (2024)
The Pitfalls of Publishing in the Age of LLMs: Strange and Surprising Adventures with a High-Impact NLP Journal
by: Verma, Rakesh M., et al.
Published: (2024)
by: Verma, Rakesh M., et al.
Published: (2024)
Steering LLMs for Formal Theorem Proving
by: Kirtania, Shashank, et al.
Published: (2025)
by: Kirtania, Shashank, et al.
Published: (2025)
Difficult Task Yes but Simple Task No: Unveiling the Laziness in Multimodal LLMs
by: Zhao, Sihang, et al.
Published: (2024)
by: Zhao, Sihang, et al.
Published: (2024)
Cognitive BASIC: An In-Model Interpreted Reasoning Language for LLMs
by: Kramer, Oliver
Published: (2025)
by: Kramer, Oliver
Published: (2025)
Automating Legal Interpretation with LLMs: Retrieval, Generation, and Evaluation
by: Luo, Kangcheng, et al.
Published: (2025)
by: Luo, Kangcheng, et al.
Published: (2025)
Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective
by: Chandna, Bhavik, et al.
Published: (2025)
by: Chandna, Bhavik, et al.
Published: (2025)
Interpretability Framework for LLMs in Undergraduate Calculus
by: Dakshit, Sagnik, et al.
Published: (2025)
by: Dakshit, Sagnik, et al.
Published: (2025)
Lived Experience Not Found: LLMs Struggle to Align with Experts on Addressing Adverse Drug Reactions from Psychiatric Medication Use
by: Chandra, Mohit, et al.
Published: (2024)
by: Chandra, Mohit, et al.
Published: (2024)
From Policy to Logic for Efficient and Interpretable Coverage Assessment
by: Pokharel, Rhitabrat, et al.
Published: (2026)
by: Pokharel, Rhitabrat, et al.
Published: (2026)
BASIL: Bayesian Assessment of Sycophancy in LLMs
by: Atwell, Katherine, et al.
Published: (2025)
by: Atwell, Katherine, et al.
Published: (2025)
UtilityMax Prompting: A Formal Framework for Multi-Objective Large Language Model Tasks
by: Marom, Ofir
Published: (2026)
by: Marom, Ofir
Published: (2026)
Bridging Legal Interpretation and Formal Logic: Faithfulness, Assumption, and the Future of AI Legal Reasoning
by: Wang, Olivia Peiyu, et al.
Published: (2026)
by: Wang, Olivia Peiyu, et al.
Published: (2026)
Vision-Language Interpreter for Robot Task Planning
by: Shirai, Keisuke, et al.
Published: (2023)
by: Shirai, Keisuke, et al.
Published: (2023)
Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability
by: Raimondi, Bianca, et al.
Published: (2025)
by: Raimondi, Bianca, et al.
Published: (2025)
On Behalf of the Stakeholders: Trends in NLP Model Interpretability in the Era of LLMs
by: Calderon, Nitay, et al.
Published: (2024)
by: Calderon, Nitay, et al.
Published: (2024)
Hessian-Enhanced Token Attribution (HETA): Interpreting Autoregressive LLMs
by: Pramanik, Vishal, et al.
Published: (2026)
by: Pramanik, Vishal, et al.
Published: (2026)
AI Planning: A Primer and Survey (Preliminary Report)
by: Chen, Dillon Z., et al.
Published: (2024)
by: Chen, Dillon Z., et al.
Published: (2024)
Interpretable Question Answering with Knowledge Graphs
by: Aneja, Kartikeya, et al.
Published: (2025)
by: Aneja, Kartikeya, et al.
Published: (2025)
Position: Avoid Overstretching LLMs for every Enterprise Task
by: Singh, Kuldeep, et al.
Published: (2026)
by: Singh, Kuldeep, et al.
Published: (2026)
Similar Items
-
Autonomous Evaluation of LLMs for Truth Maintenance and Reasoning Tasks
by: Karia, Rushang, et al.
Published: (2024) -
Epistemic Exploration for Generalizable Planning and Learning in Non-Stationary Settings
by: Karia, Rushang, et al.
Published: (2024) -
Discovering and Learning Probabilistic Models of Black-Box AI Capabilities
by: Bramblett, Daniel, et al.
Published: (2025) -
Using Explainable AI and Hierarchical Planning for Outreach with Robots
by: Karia, Rushang, et al.
Published: (2024) -
Belief-State Query Policies for User-Aligned POMDPs
by: Bramblett, Daniel, et al.
Published: (2024)