How to benchmark: the Measure-Explain-Test-Improve loop
Fuente:
arXiv
Saved in:
| Main Author: | Scherer, Gabriel |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Omnidirectional type inference for ML: principality any way
by: O'Brien, Alistair, et al.
Published: (2025)
by: O'Brien, Alistair, et al.
Published: (2025)
Unboxed data constructors -- or, how cpp decides a halting problem
by: Chataing, Nicolas, et al.
Published: (2023)
by: Chataing, Nicolas, et al.
Published: (2023)
Tail Modulo Cons, OCaml, and Relational Separation Logic
by: Allain, Clément, et al.
Published: (2024)
by: Allain, Clément, et al.
Published: (2024)
Comparison of Three Programming Error Measures for Explaining Variability in CS1 Grades
by: Švábenský, Valdemar, et al.
Published: (2024)
by: Švábenský, Valdemar, et al.
Published: (2024)
Explaining Explanations in Probabilistic Logic Programming
by: Vidal, Germán
Published: (2024)
by: Vidal, Germán
Published: (2024)
The nature of loops in programming
by: Meyer, Bertrand
Published: (2025)
by: Meyer, Bertrand
Published: (2025)
Detecting and Explaining (In-)equivalence of Context-Free Grammars
by: Schmellenkamp, Marko, et al.
Published: (2024)
by: Schmellenkamp, Marko, et al.
Published: (2024)
MHRC-Bench: A Multilingual Hardware Repository-Level Code Completion benchmark
by: Zou, Qingyun, et al.
Published: (2026)
by: Zou, Qingyun, et al.
Published: (2026)
A benchmark for vericoding: formally verified program synthesis
by: Bursuc, Sergiu, et al.
Published: (2025)
by: Bursuc, Sergiu, et al.
Published: (2025)
Membership Testing for Semantic Regular Expressions
by: Huang, Yifei, et al.
Published: (2024)
by: Huang, Yifei, et al.
Published: (2024)
Trace-Guided Synthesis of Effectful Test Generators
by: Zhou, Zhe, et al.
Published: (2026)
by: Zhou, Zhe, et al.
Published: (2026)
Bounded Exhaustive Random Program Generation for Testing Solidity Compilers
by: Ma, Haoyang, et al.
Published: (2025)
by: Ma, Haoyang, et al.
Published: (2025)
Automatic Code and Test Generation of Smart Contracts from Coordination Models
by: Selabi, Elvis Konjoh, et al.
Published: (2026)
by: Selabi, Elvis Konjoh, et al.
Published: (2026)
jMT: Testing Correctness of Java Memory Models (Extended Version)
by: Panneke, Lukas, et al.
Published: (2026)
by: Panneke, Lukas, et al.
Published: (2026)
Causal-Consistent Reversible Debugging: Improving CauDEr
by: González-Abril, Juan José, et al.
Published: (2024)
by: González-Abril, Juan José, et al.
Published: (2024)
WatChat: Explaining perplexing programs by debugging mental models
by: Chandra, Kartik, et al.
Published: (2024)
by: Chandra, Kartik, et al.
Published: (2024)
Testing, Credible Compilation, and Verification in the Axon Verified Compiler in Lean and Claude Code
by: Rinard, Martin
Published: (2026)
by: Rinard, Martin
Published: (2026)
Which Part of the Heap is Useful? Improving Heap Liveness Analysis
by: Kanvar, Vini, et al.
Published: (2024)
by: Kanvar, Vini, et al.
Published: (2024)
Formalizing Linear Motion G-code for Invariant Checking and Differential Testing of Fabrication Tools
by: He, Yumeng, et al.
Published: (2025)
by: He, Yumeng, et al.
Published: (2025)
Improving compiler support for SIMD offload using Arm Streaming SVE
by: Mohamed, Mohamed Husain Noor, et al.
Published: (2025)
by: Mohamed, Mohamed Husain Noor, et al.
Published: (2025)
Generating Functions Meet Occupation Measures: Invariant Synthesis for Probabilistic Loops (Extended Version)
by: Haase, Darion, et al.
Published: (2026)
by: Haase, Darion, et al.
Published: (2026)
$\textbf{PLUM}$: Improving Code LMs with Execution-Guided On-Policy Preference Learning Driven By Synthetic Test Cases
by: Zhang, Dylan, et al.
Published: (2024)
by: Zhang, Dylan, et al.
Published: (2024)
Converting IEC 61131-3 LD into SFC Using Large Language Model: Dataset and Testing
by: Zhang, Yimin, et al.
Published: (2025)
by: Zhang, Yimin, et al.
Published: (2025)
The Master-Slave Encoder Model for Improving Patent Text Summarization: A New Approach to Combining Specifications and Claims
by: Zhou, Shu, et al.
Published: (2024)
by: Zhou, Shu, et al.
Published: (2024)
Dynamic Program Slices Change How Developers Diagnose Gradual Run-Time Type Errors
by: Schwerter, Felipe Bañados, et al.
Published: (2025)
by: Schwerter, Felipe Bañados, et al.
Published: (2025)
SkipFlow: Improving the Precision of Points-to Analysis using Primitive Values and Predicate Edges
by: Kozak, David, et al.
Published: (2025)
by: Kozak, David, et al.
Published: (2025)
Programmable Property-Based Testing
by: Keles, Alperen, et al.
Published: (2026)
by: Keles, Alperen, et al.
Published: (2026)
Pyrosome: Verified Compilation for Modular Metatheory
by: Jamner, Dustin, et al.
Published: (2025)
by: Jamner, Dustin, et al.
Published: (2025)
How Programming Concepts and Neurons Are Shared in Code Language Models
by: Kargaran, Amir Hossein, et al.
Published: (2025)
by: Kargaran, Amir Hossein, et al.
Published: (2025)
Type-level Property Based Testing
by: Hansen, Thomas Ekström, et al.
Published: (2024)
by: Hansen, Thomas Ekström, et al.
Published: (2024)
Understanding How CodeLLMs (Mis)Predict Types with Activation Steering
by: Lucchetti, Francesca, et al.
Published: (2024)
by: Lucchetti, Francesca, et al.
Published: (2024)
Lua API and benchmark design using 3n+1 sequences: Comparing API elegance and raw speed in Redis and YottaDB databases
by: Hoyt, Berwyn
Published: (2024)
by: Hoyt, Berwyn
Published: (2024)
Pragma driven shared memory parallelism in Zig by supporting OpenMP loop directives
by: Kacs, David, et al.
Published: (2024)
by: Kacs, David, et al.
Published: (2024)
Etna: An Evaluation Platform for Property-Based Testing
by: Keles, Alperen, et al.
Published: (2026)
by: Keles, Alperen, et al.
Published: (2026)
LitmusKt: Concurrency Stress Testing for Kotlin
by: Lochmelis, Denis, et al.
Published: (2025)
by: Lochmelis, Denis, et al.
Published: (2025)
Conceptual Mutation Testing for Student Programming Misconceptions
by: Prasad, Siddhartha, et al.
Published: (2023)
by: Prasad, Siddhartha, et al.
Published: (2023)
Mica: Automated Differential Testing for OCaml Modules
by: Ng, Ernest, et al.
Published: (2024)
by: Ng, Ernest, et al.
Published: (2024)
How Do Humans Write Code? Large Models Do It the Same Way Too
by: Li, Long, et al.
Published: (2024)
by: Li, Long, et al.
Published: (2024)
CSSG: Measuring Code Similarity with Semantic Graphs
by: Lu, Yiyang, et al.
Published: (2026)
by: Lu, Yiyang, et al.
Published: (2026)
Testing the Unknown: A Framework for OpenMP Testing via Random Program Generation
by: Laguna, Ignacio, et al.
Published: (2024)
by: Laguna, Ignacio, et al.
Published: (2024)
Similar Items
-
Omnidirectional type inference for ML: principality any way
by: O'Brien, Alistair, et al.
Published: (2025) -
Unboxed data constructors -- or, how cpp decides a halting problem
by: Chataing, Nicolas, et al.
Published: (2023) -
Tail Modulo Cons, OCaml, and Relational Separation Logic
by: Allain, Clément, et al.
Published: (2024) -
Comparison of Three Programming Error Measures for Explaining Variability in CS1 Grades
by: Švábenský, Valdemar, et al.
Published: (2024) -
Explaining Explanations in Probabilistic Logic Programming
by: Vidal, Germán
Published: (2024)