Current Practices for Building LLM-Powered Reasoning Tools Are Ad Hoc -- and We Can Do Better
Fuente:
arXiv
Guardado en:
| Autor principal: | Bembenek, Aaron |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A Declarative Language for Building And Orchestrating LLM-Powered Agent Workflows
por: Daunis, Ivan
Publicado: (2025)
por: Daunis, Ivan
Publicado: (2025)
Bit-Vector CHC Solving for Binary Analysis and Binary Analysis for Bit-Vector CHC Solving
por: Bembenek, Aaron, et al.
Publicado: (2026)
por: Bembenek, Aaron, et al.
Publicado: (2026)
Oracular Programming: A Modular Foundation for Building LLM-Enabled Software
por: Laurent, Jonathan, et al.
Publicado: (2025)
por: Laurent, Jonathan, et al.
Publicado: (2025)
Making Formulog Fast: An Argument for Unconventional Datalog Evaluation (Extended Version)
por: Bembenek, Aaron, et al.
Publicado: (2024)
por: Bembenek, Aaron, et al.
Publicado: (2024)
BetterV: Controlled Verilog Generation with Discriminative Guidance
por: Pei, Zehua, et al.
Publicado: (2024)
por: Pei, Zehua, et al.
Publicado: (2024)
An LLM-Tool Compiler for Fused Parallel Function Calling
por: Singh, Simranjit, et al.
Publicado: (2024)
por: Singh, Simranjit, et al.
Publicado: (2024)
Towards Adaptive, Scalable, and Robust Coordination of LLM Agents: A Dynamic Ad-Hoc Networking Perspective
por: Li, Rui, et al.
Publicado: (2026)
por: Li, Rui, et al.
Publicado: (2026)
Can We Trust LLM Detectors?
por: Sandhan, Jivnesh, et al.
Publicado: (2026)
por: Sandhan, Jivnesh, et al.
Publicado: (2026)
Towards LLM-based optimization compilers. Can LLMs learn how to apply a single peephole optimization? Reasoning is all LLMs need!
por: Fang, Xiangxin, et al.
Publicado: (2024)
por: Fang, Xiangxin, et al.
Publicado: (2024)
Towards Formal Verification of LLM-Generated Code from Natural Language Prompts
por: Councilman, Aaron, et al.
Publicado: (2025)
por: Councilman, Aaron, et al.
Publicado: (2025)
Hey Pentti, We Did It!: A Fully Vector-Symbolic Lisp
por: Tomkins-Flanagan, Eilene, et al.
Publicado: (2025)
por: Tomkins-Flanagan, Eilene, et al.
Publicado: (2025)
OBsmith: LLM-Powered JavaScript Obfuscator Testing
por: Jiang, Shan, et al.
Publicado: (2025)
por: Jiang, Shan, et al.
Publicado: (2025)
LLM-REVal: Can We Trust LLM Reviewers Yet?
por: Li, Rui, et al.
Publicado: (2025)
por: Li, Rui, et al.
Publicado: (2025)
Automatically Improving LLM-based Verilog Generation using EDA Tool Feedback
por: Blocklove, Jason, et al.
Publicado: (2024)
por: Blocklove, Jason, et al.
Publicado: (2024)
How Do Humans Write Code? Large Models Do It the Same Way Too
por: Li, Long, et al.
Publicado: (2024)
por: Li, Long, et al.
Publicado: (2024)
BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software
por: Zhang, Zehua, et al.
Publicado: (2025)
por: Zhang, Zehua, et al.
Publicado: (2025)
Symbol Correctness in Deep Neural Networks Containing Symbolic Layers
por: Bembenek, Aaron, et al.
Publicado: (2024)
por: Bembenek, Aaron, et al.
Publicado: (2024)
Can Language Models Solve Olympiad Programming?
por: Shi, Quan, et al.
Publicado: (2024)
por: Shi, Quan, et al.
Publicado: (2024)
Can LLMs Reason About Program Semantics? A Comprehensive Evaluation of LLMs on Formal Specification Inference
por: Le-Cong, Thanh, et al.
Publicado: (2025)
por: Le-Cong, Thanh, et al.
Publicado: (2025)
CodeARC: Benchmarking Reasoning Capabilities of LLM Agents for Inductive Program Synthesis
por: Wei, Anjiang, et al.
Publicado: (2025)
por: Wei, Anjiang, et al.
Publicado: (2025)
Improving LLM Code Reasoning via Semantic Equivalence Self-Play with Formal Verification
por: Barone, Antonio Valerio Miceli, et al.
Publicado: (2026)
por: Barone, Antonio Valerio Miceli, et al.
Publicado: (2026)
Learning From Mistakes Makes LLM Better Reasoner
por: An, Shengnan, et al.
Publicado: (2023)
por: An, Shengnan, et al.
Publicado: (2023)
From Tool Calling to Symbolic Thinking: LLMs in a Persistent Lisp Metaprogramming Loop
por: de la Torre, Jordi
Publicado: (2025)
por: de la Torre, Jordi
Publicado: (2025)
Can LLMs Learn by Teaching for Better Reasoning? A Preliminary Study
por: Ning, Xuefei, et al.
Publicado: (2024)
por: Ning, Xuefei, et al.
Publicado: (2024)
REAMS: Reasoning Enhanced Algorithm for Maths Solving
por: Singh, Eishkaran, et al.
Publicado: (2025)
por: Singh, Eishkaran, et al.
Publicado: (2025)
LiteCoOp: Lightweight Multi-LLM Shared-Tree Reasoning for Model-Serving Compiler Optimizations
por: Tang, Annabelle Sujun, et al.
Publicado: (2026)
por: Tang, Annabelle Sujun, et al.
Publicado: (2026)
ToolCaching: Towards Efficient Caching for LLM Tool-calling
por: Zhai, Yi, et al.
Publicado: (2026)
por: Zhai, Yi, et al.
Publicado: (2026)
LLMs versus the Halting Problem: Characterizing Program Termination Reasoning
por: Sultan, Oren, et al.
Publicado: (2026)
por: Sultan, Oren, et al.
Publicado: (2026)
Provable Coordination for LLM Agents via Message Sequence Charts
por: Bollig, Benedikt, et al.
Publicado: (2026)
por: Bollig, Benedikt, et al.
Publicado: (2026)
HYSYNTH: Context-Free LLM Approximation for Guiding Program Synthesis
por: Barke, Shraddha, et al.
Publicado: (2024)
por: Barke, Shraddha, et al.
Publicado: (2024)
DriftScript: A Domain-Specific Language for Programming Non-Axiomatic Reasoning Agents
por: Brady, Seamus
Publicado: (2026)
por: Brady, Seamus
Publicado: (2026)
Improving LLM Classification of Logical Errors by Integrating Error Relationship into Prompts
por: Lee, Yanggyu, et al.
Publicado: (2024)
por: Lee, Yanggyu, et al.
Publicado: (2024)
Better LLM Reasoning via Dual-Play
por: Zhang, Zhengxin, et al.
Publicado: (2025)
por: Zhang, Zhengxin, et al.
Publicado: (2025)
Testing and Understanding Erroneous Planning in LLM Agents through Synthesized User Inputs
por: Ji, Zhenlan, et al.
Publicado: (2024)
por: Ji, Zhenlan, et al.
Publicado: (2024)
Integrating Reasoning Systems for Trustworthy AI, Proceedings of the 4th Workshop on Logic and Practice of Programming (LPOP)
por: Nerode, Anil, et al.
Publicado: (2024)
por: Nerode, Anil, et al.
Publicado: (2024)
HyGenar: An LLM-Driven Hybrid Genetic Algorithm for Few-Shot Grammar Generation
por: Tang, Weizhi, et al.
Publicado: (2025)
por: Tang, Weizhi, et al.
Publicado: (2025)
ToolLibGen: Scalable Automatic Tool Creation and Aggregation for LLM Reasoning
por: Yue, Murong, et al.
Publicado: (2025)
por: Yue, Murong, et al.
Publicado: (2025)
Hey Pentti, We Did (More of) It!: A Vector-Symbolic Lisp With Residue Arithmetic
por: Hanley, Connor, et al.
Publicado: (2025)
por: Hanley, Connor, et al.
Publicado: (2025)
Do Before You Judge: Self-Reference as a Pathway to Better LLM Evaluation
por: Lin, Wei-Hsiang, et al.
Publicado: (2025)
por: Lin, Wei-Hsiang, et al.
Publicado: (2025)
Adaptive Inference-Time Compute: LLMs Can Predict if They Can Do Better, Even Mid-Generation
por: Manvi, Rohin, et al.
Publicado: (2024)
por: Manvi, Rohin, et al.
Publicado: (2024)
Ejemplares similares
-
A Declarative Language for Building And Orchestrating LLM-Powered Agent Workflows
por: Daunis, Ivan
Publicado: (2025) -
Bit-Vector CHC Solving for Binary Analysis and Binary Analysis for Bit-Vector CHC Solving
por: Bembenek, Aaron, et al.
Publicado: (2026) -
Oracular Programming: A Modular Foundation for Building LLM-Enabled Software
por: Laurent, Jonathan, et al.
Publicado: (2025) -
Making Formulog Fast: An Argument for Unconventional Datalog Evaluation (Extended Version)
por: Bembenek, Aaron, et al.
Publicado: (2024) -
BetterV: Controlled Verilog Generation with Discriminative Guidance
por: Pei, Zehua, et al.
Publicado: (2024)