Program Semantic Inequivalence Game with Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Miceli-Barone, Antonio Valerio, Belle, Vaishak, Payani, Ali |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improving LLM Code Reasoning via Semantic Equivalence Self-Play with Formal Verification
by: Barone, Antonio Valerio Miceli, et al.
Published: (2026)
by: Barone, Antonio Valerio Miceli, et al.
Published: (2026)
Used Car Salesbots? Honesty and Credulity of LLMs as Bargaining Agents under Partial Information
by: Miceli-Barone, Antonio Valerio, et al.
Published: (2026)
by: Miceli-Barone, Antonio Valerio, et al.
Published: (2026)
Neuro-symbolic Weak Supervision: Theory and Semantics
by: Upreti, Nijesh, et al.
Published: (2025)
by: Upreti, Nijesh, et al.
Published: (2025)
Emergent Representations of Program Semantics in Language Models Trained on Programs
by: Jin, Charles, et al.
Published: (2023)
by: Jin, Charles, et al.
Published: (2023)
Zero, Finite, and Infinite Belief History of Theory of Mind Reasoning in Large Language Models
by: Tang, Weizhi, et al.
Published: (2024)
by: Tang, Weizhi, et al.
Published: (2024)
APPL: A Prompt Programming Language for Harmonious Integration of Programs and Large Language Model Prompts
by: Dong, Honghua, et al.
Published: (2024)
by: Dong, Honghua, et al.
Published: (2024)
ToM-LM: Delegating Theory of Mind Reasoning to External Symbolic Executors in Large Language Models
by: Tang, Weizhi, et al.
Published: (2024)
by: Tang, Weizhi, et al.
Published: (2024)
EquiBench: Benchmarking Large Language Models' Reasoning about Program Semantics via Equivalence Checking
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
LTLBench: Towards Benchmarks for Evaluating Temporal Reasoning in Large Language Models
by: Tang, Weizhi, et al.
Published: (2024)
by: Tang, Weizhi, et al.
Published: (2024)
HyGenar: An LLM-Driven Hybrid Genetic Algorithm for Few-Shot Grammar Generation
by: Tang, Weizhi, et al.
Published: (2025)
by: Tang, Weizhi, et al.
Published: (2025)
Code Simulation Challenges for Large Language Models
by: La Malfa, Emanuele, et al.
Published: (2024)
by: La Malfa, Emanuele, et al.
Published: (2024)
Chain of Execution Supervision Promotes General Reasoning in Large Language Models
by: Chen, Nuo, et al.
Published: (2025)
by: Chen, Nuo, et al.
Published: (2025)
Logic Tensor Network-Enhanced Generative Adversarial Network
by: Upreti, Nijesh, et al.
Published: (2026)
by: Upreti, Nijesh, et al.
Published: (2026)
AIOS Compiler: LLM as Interpreter for Natural Language Programming and Flow Programming of AI Agents
by: Xu, Shuyuan, et al.
Published: (2024)
by: Xu, Shuyuan, et al.
Published: (2024)
PassNet: Scaling Large Language Models for Graph Compiler Pass Generation
by: Liu, Yiqun, et al.
Published: (2026)
by: Liu, Yiqun, et al.
Published: (2026)
Synthesizing Programmatic Reinforcement Learning Policies with Large Language Model Guided Search
by: Liu, Max, et al.
Published: (2024)
by: Liu, Max, et al.
Published: (2024)
Tilus: A Tile-Level GPGPU Programming Language for Low-Precision Computation
by: Ding, Yaoyao, et al.
Published: (2025)
by: Ding, Yaoyao, et al.
Published: (2025)
Bridging the Knowledge Void: Inference-time Acquisition of Unfamiliar Programming Languages for Coding Tasks
by: Shen, Chen, et al.
Published: (2026)
by: Shen, Chen, et al.
Published: (2026)
Program Synthesis using Inductive Logic Programming for the Abstraction and Reasoning Corpus
by: Rocha, Filipe Marinho, et al.
Published: (2024)
by: Rocha, Filipe Marinho, et al.
Published: (2024)
Large Language Models for Code Summarization
by: Szalontai, Balázs, et al.
Published: (2024)
by: Szalontai, Balázs, et al.
Published: (2024)
SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models
by: He, Zirui, et al.
Published: (2025)
by: He, Zirui, et al.
Published: (2025)
The Elements of Differentiable Programming
by: Blondel, Mathieu, et al.
Published: (2024)
by: Blondel, Mathieu, et al.
Published: (2024)
EnCompass: Enhancing Agent Programming with Search Over Program Execution Paths
by: Li, Zhening, et al.
Published: (2025)
by: Li, Zhening, et al.
Published: (2025)
Searching for Programmatic Policies in Semantic Spaces
by: Moraes, Rubens O., et al.
Published: (2024)
by: Moraes, Rubens O., et al.
Published: (2024)
Do Large Code Models Understand Programming Concepts? Counterfactual Analysis for Code Predicates
by: Hooda, Ashish, et al.
Published: (2024)
by: Hooda, Ashish, et al.
Published: (2024)
Probabilistic Programming with Programmable Variational Inference
by: Becker, McCoy R., et al.
Published: (2024)
by: Becker, McCoy R., et al.
Published: (2024)
Prism: Symbolic Superoptimization of Tensor Programs
by: Wu, Mengdi, et al.
Published: (2026)
by: Wu, Mengdi, et al.
Published: (2026)
Curriculum Learning for Small Code Language Models
by: Naïr, Marwa, et al.
Published: (2024)
by: Naïr, Marwa, et al.
Published: (2024)
ChatDBG: Augmenting Debugging with Large Language Models
by: Levin, Kyla H., et al.
Published: (2024)
by: Levin, Kyla H., et al.
Published: (2024)
Large Language Models Synergize with Automated Machine Learning
by: Xu, Jinglue, et al.
Published: (2024)
by: Xu, Jinglue, et al.
Published: (2024)
Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages
by: Li, Zihao, et al.
Published: (2024)
by: Li, Zihao, et al.
Published: (2024)
Neural Networks Decoded: Targeted and Robust Analysis of Neural Network Decisions via Causal Explanations and Reasoning
by: Diallo, Alec F., et al.
Published: (2024)
by: Diallo, Alec F., et al.
Published: (2024)
Generating Pragmatic Examples to Train Neural Program Synthesizers
by: Vaduguru, Saujas, et al.
Published: (2023)
by: Vaduguru, Saujas, et al.
Published: (2023)
Mirage: A Multi-Level Superoptimizer for Tensor Programs
by: Wu, Mengdi, et al.
Published: (2024)
by: Wu, Mengdi, et al.
Published: (2024)
ProofOptimizer: Training Language Models to Simplify Proofs without Human Demonstrations
by: Gu, Alex, et al.
Published: (2025)
by: Gu, Alex, et al.
Published: (2025)
Hexcute: A Compiler Framework for Automating Layout Synthesis in GPU Programs
by: Zhang, Xiao, et al.
Published: (2025)
by: Zhang, Xiao, et al.
Published: (2025)
Beyond Single Concept Vector: Modeling Concept Subspace in LLMs with Gaussian Distribution
by: Zhao, Haiyan, et al.
Published: (2024)
by: Zhao, Haiyan, et al.
Published: (2024)
A Multi-Expert Large Language Model Architecture for Verilog Code Generation
by: Nadimi, Bardia, et al.
Published: (2024)
by: Nadimi, Bardia, et al.
Published: (2024)
EcoSearch: A Constant-Delay Best-First Search Algorithm for Program Synthesis
by: Matricon, Théo, et al.
Published: (2024)
by: Matricon, Théo, et al.
Published: (2024)
ScenicNL: Generating Probabilistic Scenario Programs from Natural Language
by: Elmaaroufi, Karim, et al.
Published: (2024)
by: Elmaaroufi, Karim, et al.
Published: (2024)
Similar Items
-
Improving LLM Code Reasoning via Semantic Equivalence Self-Play with Formal Verification
by: Barone, Antonio Valerio Miceli, et al.
Published: (2026) -
Used Car Salesbots? Honesty and Credulity of LLMs as Bargaining Agents under Partial Information
by: Miceli-Barone, Antonio Valerio, et al.
Published: (2026) -
Neuro-symbolic Weak Supervision: Theory and Semantics
by: Upreti, Nijesh, et al.
Published: (2025) -
Emergent Representations of Program Semantics in Language Models Trained on Programs
by: Jin, Charles, et al.
Published: (2023) -
Zero, Finite, and Infinite Belief History of Theory of Mind Reasoning in Large Language Models
by: Tang, Weizhi, et al.
Published: (2024)