Models That Prove Their Own Correctness
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Amit, Noga, Goldwasser, Shafi, Paradise, Orr, Rothblum, Guy |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Learning Randomized Reductions
von: Erata, Ferhat, et al.
Veröffentlicht: (2024)
von: Erata, Ferhat, et al.
Veröffentlicht: (2024)
Oblivious Defense in ML Models: Backdoor Removal without Detection
von: Goldwasser, Shafi, et al.
Veröffentlicht: (2024)
von: Goldwasser, Shafi, et al.
Veröffentlicht: (2024)
CodeComplex: Dataset for Worst-Case Time Complexity Prediction
von: Baik, Seung-Yeop, et al.
Veröffentlicht: (2024)
von: Baik, Seung-Yeop, et al.
Veröffentlicht: (2024)
Automatizing Software Cognitive Complexity Reduction through Integer Linear Programming
von: Saborido, Rubén, et al.
Veröffentlicht: (2024)
von: Saborido, Rubén, et al.
Veröffentlicht: (2024)
On Computationally Efficient Multi-Class Calibration
von: Gopalan, Parikshit, et al.
Veröffentlicht: (2024)
von: Gopalan, Parikshit, et al.
Veröffentlicht: (2024)
How to Verify Any (Reasonable) Distribution Property: Computationally Sound Argument Systems for Distributions
von: Herman, Tal, et al.
Veröffentlicht: (2024)
von: Herman, Tal, et al.
Veröffentlicht: (2024)
Proving the Coding Interview: A Benchmark for Formally Verified Code Generation
von: Dougherty, Quinn, et al.
Veröffentlicht: (2025)
von: Dougherty, Quinn, et al.
Veröffentlicht: (2025)
GLLM: Self-Corrective G-Code Generation using Large Language Models with User Feedback
von: Abdelaal, Mohamed, et al.
Veröffentlicht: (2025)
von: Abdelaal, Mohamed, et al.
Veröffentlicht: (2025)
Which Alert Removals are Beneficial?
von: Amit, Idan
Veröffentlicht: (2026)
von: Amit, Idan
Veröffentlicht: (2026)
Calibration and Correctness of Language Models for Code
von: Spiess, Claudio, et al.
Veröffentlicht: (2024)
von: Spiess, Claudio, et al.
Veröffentlicht: (2024)
Computational Complexity of Edge Coverage Problem for Constrained Control Flow Graphs
von: Ruszil, Jakub, et al.
Veröffentlicht: (2026)
von: Ruszil, Jakub, et al.
Veröffentlicht: (2026)
Efficient Heuristics and Exact Methods for Pairwise Interaction Sampling
von: Fekete, Sándor P., et al.
Veröffentlicht: (2025)
von: Fekete, Sándor P., et al.
Veröffentlicht: (2025)
Accuracy vs. Accuracy: Computational Tradeoffs Between Classification Rates and Utility
von: Amit, Noga, et al.
Veröffentlicht: (2025)
von: Amit, Noga, et al.
Veröffentlicht: (2025)
ReflexiCoder: Teaching Large Language Models to Self-Reflect on Generated Code and Self-Correct It via Reinforcement Learning
von: Jiang, Juyong, et al.
Veröffentlicht: (2026)
von: Jiang, Juyong, et al.
Veröffentlicht: (2026)
miniCodeProps: a Minimal Benchmark for Proving Code Properties
von: Lohn, Evan, et al.
Veröffentlicht: (2024)
von: Lohn, Evan, et al.
Veröffentlicht: (2024)
Ensuring Functional Correctness of Large Code Models with Selective Generation
von: Jeong, Jaewoo, et al.
Veröffentlicht: (2025)
von: Jeong, Jaewoo, et al.
Veröffentlicht: (2025)
CoCoST: Automatic Complex Code Generation with Online Searching and Correctness Testing
von: He, Xinyi, et al.
Veröffentlicht: (2024)
von: He, Xinyi, et al.
Veröffentlicht: (2024)
A Large Scale Survey of Motivation in Software Development and Analysis of its Validity
von: Amit, Idan, et al.
Veröffentlicht: (2024)
von: Amit, Idan, et al.
Veröffentlicht: (2024)
Correctness Assessment of Code Generated by Large Language Models Using Internal Representations
von: Bui, Tuan-Dung, et al.
Veröffentlicht: (2025)
von: Bui, Tuan-Dung, et al.
Veröffentlicht: (2025)
The Constraint Tax: Measuring Validity-Correctness Tradeoffs in Structured Outputs for Small Language Models
von: Ray, Jaideep
Veröffentlicht: (2026)
von: Ray, Jaideep
Veröffentlicht: (2026)
FunPRM: Function-as-Step Process Reward Model with Meta Reward Correction for Code Generation
von: Zhang, Ruiyi, et al.
Veröffentlicht: (2026)
von: Zhang, Ruiyi, et al.
Veröffentlicht: (2026)
DEM: A Method for Certifying Deep Neural Network Classifier Outputs in Aerospace
von: Katz, Guy, et al.
Veröffentlicht: (2024)
von: Katz, Guy, et al.
Veröffentlicht: (2024)
Follow Your Nose -- Which Code Smells are Worth Chasing?
von: Amit, Idan, et al.
Veröffentlicht: (2021)
von: Amit, Idan, et al.
Veröffentlicht: (2021)
HackerRank-ASTRA: Evaluating Correctness & Consistency of Large Language Models on cross-domain multi-file project problems
von: Xing, Jun, et al.
Veröffentlicht: (2025)
von: Xing, Jun, et al.
Veröffentlicht: (2025)
Unsupervised Evaluation of Code LLMs with Round-Trip Correctness
von: Allamanis, Miltiadis, et al.
Veröffentlicht: (2024)
von: Allamanis, Miltiadis, et al.
Veröffentlicht: (2024)
Mechanistic Interpretability of Code Correctness in LLMs via Sparse Autoencoders
von: Tahimic, Kriz, et al.
Veröffentlicht: (2025)
von: Tahimic, Kriz, et al.
Veröffentlicht: (2025)
A Fast and Scalable Pathwise-Solver for Group Lasso and Elastic Net Penalized Regression via Block-Coordinate Descent
von: Yang, James, et al.
Veröffentlicht: (2024)
von: Yang, James, et al.
Veröffentlicht: (2024)
Additive Models Explained: A Computational Complexity Approach
von: Bassan, Shahaf, et al.
Veröffentlicht: (2025)
von: Bassan, Shahaf, et al.
Veröffentlicht: (2025)
Evaluating the Process Modeling Abilities of Large Language Models -- Preliminary Foundations and Results
von: Fettke, Peter, et al.
Veröffentlicht: (2025)
von: Fettke, Peter, et al.
Veröffentlicht: (2025)
ExplainFuzz: Explainable and Constraint-Conditioned Test Generation with Probabilistic Circuits
von: Baiget, Annaëlle, et al.
Veröffentlicht: (2026)
von: Baiget, Annaëlle, et al.
Veröffentlicht: (2026)
Evaluating Language Models for Efficient Code Generation
von: Liu, Jiawei, et al.
Veröffentlicht: (2024)
von: Liu, Jiawei, et al.
Veröffentlicht: (2024)
BugGen: A Self-Correcting Multi-Agent LLM Pipeline for Realistic RTL Bug Synthesis
von: Jasper, Surya, et al.
Veröffentlicht: (2025)
von: Jasper, Surya, et al.
Veröffentlicht: (2025)
Evaluation and Improvement of Fault Detection for Large Language Models
von: Hu, Qiang, et al.
Veröffentlicht: (2024)
von: Hu, Qiang, et al.
Veröffentlicht: (2024)
Towards Understanding What Code Language Models Learned
von: Ahmed, Toufique, et al.
Veröffentlicht: (2023)
von: Ahmed, Toufique, et al.
Veröffentlicht: (2023)
Mass-Producing Failures of Multimodal Systems with Language Models
von: Tong, Shengbang, et al.
Veröffentlicht: (2023)
von: Tong, Shengbang, et al.
Veröffentlicht: (2023)
CodeJudge: Evaluating Code Generation with Large Language Models
von: Tong, Weixi, et al.
Veröffentlicht: (2024)
von: Tong, Weixi, et al.
Veröffentlicht: (2024)
Quantifying Contamination in Evaluating Code Generation Capabilities of Language Models
von: Riddell, Martin, et al.
Veröffentlicht: (2024)
von: Riddell, Martin, et al.
Veröffentlicht: (2024)
OptLLM: Optimal Assignment of Queries to Large Language Models
von: Liu, Yueyue, et al.
Veröffentlicht: (2024)
von: Liu, Yueyue, et al.
Veröffentlicht: (2024)
Evaluating the Generalization Capabilities of Large Language Models on Code Reasoning
von: Yang, Rem, et al.
Veröffentlicht: (2025)
von: Yang, Rem, et al.
Veröffentlicht: (2025)
Data vs. Model Machine Learning Fairness Testing: An Empirical Study
von: Shome, Arumoy, et al.
Veröffentlicht: (2024)
von: Shome, Arumoy, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Learning Randomized Reductions
von: Erata, Ferhat, et al.
Veröffentlicht: (2024) -
Oblivious Defense in ML Models: Backdoor Removal without Detection
von: Goldwasser, Shafi, et al.
Veröffentlicht: (2024) -
CodeComplex: Dataset for Worst-Case Time Complexity Prediction
von: Baik, Seung-Yeop, et al.
Veröffentlicht: (2024) -
Automatizing Software Cognitive Complexity Reduction through Integer Linear Programming
von: Saborido, Rubén, et al.
Veröffentlicht: (2024) -
On Computationally Efficient Multi-Class Calibration
von: Gopalan, Parikshit, et al.
Veröffentlicht: (2024)