OSVBench: Benchmarking LLMs on Specification Generation Tasks for Operating System Verification
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Shangyu, Jiang, Juyong, Zhao, Tiancheng, Shen, Jiasi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Quine: Realizing LLM Agents as Native POSIX Processes
by: Ke, Hao
Published: (2026)
by: Ke, Hao
Published: (2026)
HyperGraphOS: A Modern Meta-Operating System for the Scientific and Engineering Domains
by: Ceravola, Antonello, et al.
Published: (2024)
by: Ceravola, Antonello, et al.
Published: (2024)
VeruSAGE: A Study of Agent-Based Verification for Rust Systems
by: Yang, Chenyuan, et al.
Published: (2025)
by: Yang, Chenyuan, et al.
Published: (2025)
Scaling Inter-procedural Dataflow Analysis on the Cloud
by: Sun, Zewen, et al.
Published: (2024)
by: Sun, Zewen, et al.
Published: (2024)
A Survey on Large Language Models for Code Generation
by: Jiang, Juyong, et al.
Published: (2024)
by: Jiang, Juyong, et al.
Published: (2024)
Compiling Away the Overhead of Race Detection
by: Paznikov, Alexey, et al.
Published: (2025)
by: Paznikov, Alexey, et al.
Published: (2025)
Scalable and Accurate Application-Level Crash-Consistency Testing via Representative Testing
by: Gu, Yile, et al.
Published: (2025)
by: Gu, Yile, et al.
Published: (2025)
Can LLMs Enable Verification in Mainstream Programming?
by: Shefer, Aleksandr, et al.
Published: (2025)
by: Shefer, Aleksandr, et al.
Published: (2025)
Smaller = Weaker? Benchmarking Robustness of Quantized LLMs in Code Generation
by: Fang, Sen, et al.
Published: (2025)
by: Fang, Sen, et al.
Published: (2025)
CorrectHDL: Agentic HDL Design with LLMs Leveraging High-Level Synthesis as Reference
by: Xu, Kangwei, et al.
Published: (2025)
by: Xu, Kangwei, et al.
Published: (2025)
Can LLMs Reason About Program Semantics? A Comprehensive Evaluation of LLMs on Formal Specification Inference
by: Le-Cong, Thanh, et al.
Published: (2025)
by: Le-Cong, Thanh, et al.
Published: (2025)
It's Not the Size: Harness Design Determines Operational Stability in Small Language Models
by: Cho, Yong-eun
Published: (2026)
by: Cho, Yong-eun
Published: (2026)
A Preliminary Study of Multilingual Code Language Models for Code Generation Task Using Translated Benchmarks
by: Dandamudi, Rohit, et al.
Published: (2024)
by: Dandamudi, Rohit, et al.
Published: (2024)
Ranking LLM-Generated Loop Invariants for Program Verification
by: Chakraborty, Saikat, et al.
Published: (2023)
by: Chakraborty, Saikat, et al.
Published: (2023)
BODHI: Precise OS Kernel Specification Inference
by: Chang, Zhiming, et al.
Published: (2026)
by: Chang, Zhiming, et al.
Published: (2026)
PPM: Automated Generation of Diverse Programming Problems for Benchmarking Code Generation Models
by: Chen, Simin, et al.
Published: (2024)
by: Chen, Simin, et al.
Published: (2024)
dafny-annotator: AI-Assisted Verification of Dafny Programs
by: Poesia, Gabriel, et al.
Published: (2024)
by: Poesia, Gabriel, et al.
Published: (2024)
Benchmarking LLM Code Generation for Audio Programming with Visual Dataflow Languages
by: Zhang, William, et al.
Published: (2024)
by: Zhang, William, et al.
Published: (2024)
AInsteinBench: Benchmarking Coding Agents on Scientific Repositories
by: Duston, Titouan, et al.
Published: (2025)
by: Duston, Titouan, et al.
Published: (2025)
DafnyBench: A Benchmark for Formal Software Verification
by: Loughridge, Chloe, et al.
Published: (2024)
by: Loughridge, Chloe, et al.
Published: (2024)
AutoCode: LLMs as Problem Setters for Competitive Programming
by: Zhou, Shang, et al.
Published: (2025)
by: Zhou, Shang, et al.
Published: (2025)
CodePivot: Bootstrapping Multilingual Transpilation in LLMs via Reinforcement Learning without Parallel Corpora
by: Li, Shangyu, et al.
Published: (2026)
by: Li, Shangyu, et al.
Published: (2026)
Towards Repository-Level Program Verification with Large Language Models
by: Zhong, Si Cheng, et al.
Published: (2025)
by: Zhong, Si Cheng, et al.
Published: (2025)
A Problem-Oriented Perspective and Anchor Verification for Code Optimization
by: Ye, Tong, et al.
Published: (2024)
by: Ye, Tong, et al.
Published: (2024)
Automating Modelica Module Generation Using Large Language Models: A Case Study on Building Control Description Language
by: Wan, Hanlong, et al.
Published: (2025)
by: Wan, Hanlong, et al.
Published: (2025)
Reverse Chain: A Generic-Rule for LLMs to Master Multi-API Planning
by: Zhang, Yinger, et al.
Published: (2023)
by: Zhang, Yinger, et al.
Published: (2023)
MigGPT: Harnessing Large Language Models for Automated Migration of Out-of-Tree Linux Kernel Patches Across Versions
by: Dang, Pucheng, et al.
Published: (2025)
by: Dang, Pucheng, et al.
Published: (2025)
LLMs Lean on Priors, Not Programming Language Semantics
by: Thimmaiah, Aditya, et al.
Published: (2025)
by: Thimmaiah, Aditya, et al.
Published: (2025)
The New Compiler Stack: A Survey on the Synergy of LLMs and Compilers
by: Zhang, Shuoming, et al.
Published: (2026)
by: Zhang, Shuoming, et al.
Published: (2026)
LangGPT: Rethinking Structured Reusable Prompt Design Framework for LLMs from the Programming Language
by: Wang, Ming, et al.
Published: (2024)
by: Wang, Ming, et al.
Published: (2024)
CodeIF-Bench: Evaluating Instruction-Following Capabilities of Large Language Models in Interactive Code Generation
by: Wang, Peiding, et al.
Published: (2025)
by: Wang, Peiding, et al.
Published: (2025)
SmartEval: A Benchmark for Evaluating LLM-Generated Smart Contracts from Natural Language Specifications
by: Goel, Abhinav, et al.
Published: (2026)
by: Goel, Abhinav, et al.
Published: (2026)
Extending Data Spatial Semantics for Scale Agnostic Programming
by: Mars, Jason
Published: (2025)
by: Mars, Jason
Published: (2025)
Verus-SpecGym: An Agentic Environment for Evaluating Specification Autoformalization
by: Agarwal, Anmol, et al.
Published: (2026)
by: Agarwal, Anmol, et al.
Published: (2026)
Benchmarking Large Language Models for ABAP Code Generation: An Empirical Study on Iterative Improvement by Compiler Feedback
by: Wallraven, Stephan, et al.
Published: (2026)
by: Wallraven, Stephan, et al.
Published: (2026)
Beyond Postconditions: Can Large Language Models infer Formal Contracts for Automatic Software Verification?
by: Richter, Cedric, et al.
Published: (2025)
by: Richter, Cedric, et al.
Published: (2025)
SACTOR: LLM-Driven Correct and Idiomatic C to Rust Translation with Static Analysis and FFI-Based Verification
by: Zhou, Tianyang, et al.
Published: (2025)
by: Zhou, Tianyang, et al.
Published: (2025)
Assessing Code Understanding in LLMs
by: Laneve, Cosimo, et al.
Published: (2025)
by: Laneve, Cosimo, et al.
Published: (2025)
Configuration Validation with Large Language Models
by: Lian, Xinyu, et al.
Published: (2023)
by: Lian, Xinyu, et al.
Published: (2023)
AIOps Solutions for Incident Management: Technical Guidelines and A Comprehensive Literature Review
by: Remil, Youcef, et al.
Published: (2024)
by: Remil, Youcef, et al.
Published: (2024)
Similar Items
-
Quine: Realizing LLM Agents as Native POSIX Processes
by: Ke, Hao
Published: (2026) -
HyperGraphOS: A Modern Meta-Operating System for the Scientific and Engineering Domains
by: Ceravola, Antonello, et al.
Published: (2024) -
VeruSAGE: A Study of Agent-Based Verification for Rust Systems
by: Yang, Chenyuan, et al.
Published: (2025) -
Scaling Inter-procedural Dataflow Analysis on the Cloud
by: Sun, Zewen, et al.
Published: (2024) -
A Survey on Large Language Models for Code Generation
by: Jiang, Juyong, et al.
Published: (2024)