Breaking the Myth: Can Small Models Infer Postconditions Too?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Gehao, Wang, Zhenting, Zhai, Juan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
POSTCONDBENCH: Benchmarking Correctness and Completeness in Formal Postcondition Inference
von: Zhang, Gehao, et al.
Veröffentlicht: (2026)
von: Zhang, Gehao, et al.
Veröffentlicht: (2026)
Can Large Language Models Transform Natural Language Intent into Formal Method Postconditions?
von: Endres, Madeline, et al.
Veröffentlicht: (2023)
von: Endres, Madeline, et al.
Veröffentlicht: (2023)
Beyond Postconditions: Can Large Language Models infer Formal Contracts for Automatic Software Verification?
von: Richter, Cedric, et al.
Veröffentlicht: (2025)
von: Richter, Cedric, et al.
Veröffentlicht: (2025)
REPOFUSE: Repository-Level Code Completion with Fused Dual Context
von: Liang, Ming, et al.
Veröffentlicht: (2024)
von: Liang, Ming, et al.
Veröffentlicht: (2024)
Beyond Code Generation: Assessing Code LLM Maturity with Postconditions
von: He, Fusen, et al.
Veröffentlicht: (2024)
von: He, Fusen, et al.
Veröffentlicht: (2024)
Inferring Pluggable Types with Machine Learning
von: Siddiqui, Kazi Amanul Islam, et al.
Veröffentlicht: (2024)
von: Siddiqui, Kazi Amanul Islam, et al.
Veröffentlicht: (2024)
Inferring Code Correctness from Specification
von: Florian, Tambon, et al.
Veröffentlicht: (2026)
von: Florian, Tambon, et al.
Veröffentlicht: (2026)
Efficient DNN-Powered Software with Fair Sparse Models
von: Gao, Xuanqi, et al.
Veröffentlicht: (2024)
von: Gao, Xuanqi, et al.
Veröffentlicht: (2024)
Deploy-Master: Automating the Deployment of 50,000+ Agent-Ready Scientific Tools in One Day
von: Wang, Yi, et al.
Veröffentlicht: (2026)
von: Wang, Yi, et al.
Veröffentlicht: (2026)
DREAM: Debugging and Repairing AutoML Pipelines
von: Zhang, Xiaoyu, et al.
Veröffentlicht: (2023)
von: Zhang, Xiaoyu, et al.
Veröffentlicht: (2023)
Small Models, Big Tasks: An Exploratory Empirical Study on Small Language Models for Function Calling
von: Kavathekar, Ishan, et al.
Veröffentlicht: (2025)
von: Kavathekar, Ishan, et al.
Veröffentlicht: (2025)
CodeCircuit: Toward Inferring LLM-Generated Code Correctness via Attribution Graphs
von: He, Yicheng, et al.
Veröffentlicht: (2026)
von: He, Yicheng, et al.
Veröffentlicht: (2026)
Breaking the Illusion of Identity in LLM Tooling
von: Miller, Marek
Veröffentlicht: (2026)
von: Miller, Marek
Veröffentlicht: (2026)
Breaking Android with AI: A Deep Dive into LLM-Powered Exploitation
von: Perera, Wanni Vidulige Ishan, et al.
Veröffentlicht: (2025)
von: Perera, Wanni Vidulige Ishan, et al.
Veröffentlicht: (2025)
This Is Taking Too Long -- Investigating Time as a Proxy for Energy Consumption of LLMs
von: Krupp, Lars, et al.
Veröffentlicht: (2026)
von: Krupp, Lars, et al.
Veröffentlicht: (2026)
Reasoning over Precedents Alongside Statutes: Case-Augmented Deliberative Alignment for LLM Safety
von: Jin, Can, et al.
Veröffentlicht: (2026)
von: Jin, Can, et al.
Veröffentlicht: (2026)
Accuracy Can Lie: On the Impact of Surrogate Model in Configuration Tuning
von: Chen, Pengzhou, et al.
Veröffentlicht: (2025)
von: Chen, Pengzhou, et al.
Veröffentlicht: (2025)
Breaking Barriers in Software Testing: The Power of AI-Driven Automation
von: Naqvi, Saba, et al.
Veröffentlicht: (2025)
von: Naqvi, Saba, et al.
Veröffentlicht: (2025)
Large Language Models for In-File Vulnerability Localization Can Be "Lost in the End"
von: Sovrano, Francesco, et al.
Veröffentlicht: (2025)
von: Sovrano, Francesco, et al.
Veröffentlicht: (2025)
ProgramBench: Can Language Models Rebuild Programs From Scratch?
von: Yang, John, et al.
Veröffentlicht: (2026)
von: Yang, John, et al.
Veröffentlicht: (2026)
GPIoT: Tailoring Small Language Models for IoT Program Synthesis and Development
von: Shen, Leming, et al.
Veröffentlicht: (2025)
von: Shen, Leming, et al.
Veröffentlicht: (2025)
ASSURE: Metamorphic Testing for AI-powered Browser Extensions
von: Gao, Xuanqi, et al.
Veröffentlicht: (2025)
von: Gao, Xuanqi, et al.
Veröffentlicht: (2025)
LlamaRestTest: Effective REST API Testing with Small Language Models
von: Kim, Myeongsoo, et al.
Veröffentlicht: (2025)
von: Kim, Myeongsoo, et al.
Veröffentlicht: (2025)
Spreadsheet Modeling Experiments Using GPTs on Small Problem Statements and the Wall Task
von: Grossman, Thomas A., et al.
Veröffentlicht: (2026)
von: Grossman, Thomas A., et al.
Veröffentlicht: (2026)
EduBot -- Can LLMs Solve Personalized Learning and Programming Assignments?
von: Wang, Yibin, et al.
Veröffentlicht: (2025)
von: Wang, Yibin, et al.
Veröffentlicht: (2025)
When the Code Autopilot Breaks: Why LLMs Falter in Embedded Machine Learning
von: Morabito, Roberto, et al.
Veröffentlicht: (2025)
von: Morabito, Roberto, et al.
Veröffentlicht: (2025)
I Can Find You in Seconds! Leveraging Large Language Models for Code Authorship Attribution
von: Choi, Soohyeon, et al.
Veröffentlicht: (2025)
von: Choi, Soohyeon, et al.
Veröffentlicht: (2025)
Can AI Models Direct Each Other? Organizational Structure as a Probe into Training Limitations
von: Liu, Rui
Veröffentlicht: (2026)
von: Liu, Rui
Veröffentlicht: (2026)
Can Agents Fix Agent Issues?
von: Rahardja, Alfin Wijaya, et al.
Veröffentlicht: (2025)
von: Rahardja, Alfin Wijaya, et al.
Veröffentlicht: (2025)
Terminus-4B: Can a Smaller Model Replace Frontier LLMs at Agentic Execution Tasks?
von: Garg, Spandan, et al.
Veröffentlicht: (2026)
von: Garg, Spandan, et al.
Veröffentlicht: (2026)
Neuro-Symbolic Generation and Validation of Memory-Aware Formal Function Specifications
von: Zhang, Liao, et al.
Veröffentlicht: (2026)
von: Zhang, Liao, et al.
Veröffentlicht: (2026)
Safer Builders, Risky Maintainers: A Comparative Study of Breaking Changes in Human vs Agentic PRs
von: Ferdous, K M, et al.
Veröffentlicht: (2026)
von: Ferdous, K M, et al.
Veröffentlicht: (2026)
Energy-Aware Code Generation with LLMs: Benchmarking Small vs. Large Language Models for Sustainable AI Programming
von: Ashraf, Humza, et al.
Veröffentlicht: (2025)
von: Ashraf, Humza, et al.
Veröffentlicht: (2025)
Can Github issues be solved with Tree Of Thoughts?
von: La Rosa, Ricardo, et al.
Veröffentlicht: (2024)
von: La Rosa, Ricardo, et al.
Veröffentlicht: (2024)
False Friends in the Shell: Unveiling the Emoticon Semantic Confusion in Large Language Models
von: Jiang, Weipeng, et al.
Veröffentlicht: (2026)
von: Jiang, Weipeng, et al.
Veröffentlicht: (2026)
Can GPT-O1 Kill All Bugs? An Evaluation of GPT-Family LLMs on QuixBugs
von: Hu, Haichuan, et al.
Veröffentlicht: (2024)
von: Hu, Haichuan, et al.
Veröffentlicht: (2024)
Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering
von: Wang, Ruiqi, et al.
Veröffentlicht: (2025)
von: Wang, Ruiqi, et al.
Veröffentlicht: (2025)
Let the Barbarians In: How AI Can Accelerate Systems Performance Research
von: Cheng, Audrey, et al.
Veröffentlicht: (2025)
von: Cheng, Audrey, et al.
Veröffentlicht: (2025)
Can LLMs Replace Humans During Code Chunking?
von: Glasz, Christopher, et al.
Veröffentlicht: (2025)
von: Glasz, Christopher, et al.
Veröffentlicht: (2025)
Can LLM Generate Regression Tests for Software Commits?
von: Liu, Jing, et al.
Veröffentlicht: (2025)
von: Liu, Jing, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
POSTCONDBENCH: Benchmarking Correctness and Completeness in Formal Postcondition Inference
von: Zhang, Gehao, et al.
Veröffentlicht: (2026) -
Can Large Language Models Transform Natural Language Intent into Formal Method Postconditions?
von: Endres, Madeline, et al.
Veröffentlicht: (2023) -
Beyond Postconditions: Can Large Language Models infer Formal Contracts for Automatic Software Verification?
von: Richter, Cedric, et al.
Veröffentlicht: (2025) -
REPOFUSE: Repository-Level Code Completion with Fused Dual Context
von: Liang, Ming, et al.
Veröffentlicht: (2024) -
Beyond Code Generation: Assessing Code LLM Maturity with Postconditions
von: He, Fusen, et al.
Veröffentlicht: (2024)