ACES: Who Tests the Tests? Leave-One-Out AUC Consistency for Code Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sun, Hui, Zhang, Yun-Ji, Xie, Zheng, Liu, Ren-Biao, Du, Yali, Li, Xin-Ye, Li, Ming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Exploring Pass-Rate Reward in Reinforcement Learning for Code Generation
von: Li, Xin-Ye, et al.
Veröffentlicht: (2026)
von: Li, Xin-Ye, et al.
Veröffentlicht: (2026)
Post-Incorporating Code Structural Knowledge into Pretrained Models via ICL for Code Translation
von: Du, Yali, et al.
Veröffentlicht: (2025)
von: Du, Yali, et al.
Veröffentlicht: (2025)
Weakly Supervised AUC Optimization: A Unified Partial AUC Approach
von: Xie, Zheng, et al.
Veröffentlicht: (2023)
von: Xie, Zheng, et al.
Veröffentlicht: (2023)
Design-Specification Tiling for ICL-based CAD Code Generation
von: Du, Yali, et al.
Veröffentlicht: (2026)
von: Du, Yali, et al.
Veröffentlicht: (2026)
A Joint Learning Model with Variational Interaction for Multilingual Program Translation
von: Du, Yali, et al.
Veröffentlicht: (2024)
von: Du, Yali, et al.
Veröffentlicht: (2024)
Top Pass: Improve Code Generation by Pass@k-Maximized Code Ranking
von: Lyu, Zhi-Cun, et al.
Veröffentlicht: (2024)
von: Lyu, Zhi-Cun, et al.
Veröffentlicht: (2024)
Leave-One-Out Prediction for General Hypothesis Classes
von: Qian, Jian, et al.
Veröffentlicht: (2026)
von: Qian, Jian, et al.
Veröffentlicht: (2026)
An Iterative Test-and-Repair Framework for Competitive Code Generation
von: Tang, Lingxiao, et al.
Veröffentlicht: (2026)
von: Tang, Lingxiao, et al.
Veröffentlicht: (2026)
ACES: Accent Subspaces for Coupling, Explanations, and Stress-Testing in Automatic Speech Recognition
von: Parekh, Swapnil
Veröffentlicht: (2026)
von: Parekh, Swapnil
Veröffentlicht: (2026)
Evaluating the Test Adequacy of Benchmarks for LLMs on Code Generation
von: Xiangyue Liu, et al.
Veröffentlicht: (2025)
von: Xiangyue Liu, et al.
Veröffentlicht: (2025)
Weighted Leave-One-Out Cross Validation
von: Pronzato, Luc, et al.
Veröffentlicht: (2025)
von: Pronzato, Luc, et al.
Veröffentlicht: (2025)
Leave-One-Out Stable Conformal Prediction
von: Lee, Kiljae, et al.
Veröffentlicht: (2025)
von: Lee, Kiljae, et al.
Veröffentlicht: (2025)
Leave-One-Out Learning with Log-Loss
von: Fogel, Yaniv, et al.
Veröffentlicht: (2025)
von: Fogel, Yaniv, et al.
Veröffentlicht: (2025)
Mutation-based Consistency Testing for Evaluating the Code Understanding Capability of LLMs
von: Li, Ziyu, et al.
Veröffentlicht: (2024)
von: Li, Ziyu, et al.
Veröffentlicht: (2024)
Asymptotically Optimal Tests for One- and Two-Sample Problems
von: Grootveld, Arick, et al.
Veröffentlicht: (2026)
von: Grootveld, Arick, et al.
Veröffentlicht: (2026)
Structural Evaluation Metrics for SVG Generation via Leave-One-Out Analysis
von: Zhu, Haonan, et al.
Veröffentlicht: (2026)
von: Zhu, Haonan, et al.
Veröffentlicht: (2026)
Enhancing LLMs in Long Code Translation through Instrumentation and Program State Alignment
von: Xin-Ye, Li, et al.
Veröffentlicht: (2025)
von: Xin-Ye, Li, et al.
Veröffentlicht: (2025)
OODD: Test-time Out-of-Distribution Detection with Dynamic Dictionary
von: Yang, Yifeng, et al.
Veröffentlicht: (2025)
von: Yang, Yifeng, et al.
Veröffentlicht: (2025)
CodeContests+: High-Quality Test Case Generation for Competitive Programming
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
von: Wang, Zihan, et al.
Veröffentlicht: (2025)
MUCOCO: Automated Consistency Testing of Code LLMs
von: Chou, Chua Jin, et al.
Veröffentlicht: (2026)
von: Chou, Chua Jin, et al.
Veröffentlicht: (2026)
When LRP Diverges from Leave-One-Out in Transformers
von: You, Weiqiu, et al.
Veröffentlicht: (2025)
von: You, Weiqiu, et al.
Veröffentlicht: (2025)
Leave-One-Out-, Bootstrap- and Cross-Conformal Anomaly Detectors
von: Hennhöfer, Oliver, et al.
Veröffentlicht: (2024)
von: Hennhöfer, Oliver, et al.
Veröffentlicht: (2024)
Leave-One-Out Analysis for Nonconvex Robust Matrix Completion with General Thresholding Functions
von: Wang, Tianming, et al.
Veröffentlicht: (2024)
von: Wang, Tianming, et al.
Veröffentlicht: (2024)
S*: Test Time Scaling for Code Generation
von: Li, Dacheng, et al.
Veröffentlicht: (2025)
von: Li, Dacheng, et al.
Veröffentlicht: (2025)
Interval-Based AUC (iAUC): Extending ROC Analysis to Uncertainty-Aware Classification
von: Li, Yuqi, et al.
Veröffentlicht: (2026)
von: Li, Yuqi, et al.
Veröffentlicht: (2026)
ScaleRTL: Scaling LLMs with Reasoning Data and Test-Time Compute for Accurate RTL Code Generation
von: Deng, Chenhui, et al.
Veröffentlicht: (2025)
von: Deng, Chenhui, et al.
Veröffentlicht: (2025)
Navigating Pharmacogenomic Testing in Practice: Who to Test and When to Test
von: James M. Stevenson, et al.
Veröffentlicht: (2025)
von: James M. Stevenson, et al.
Veröffentlicht: (2025)
ACES: Generating Diverse Programming Puzzles with with Autotelic Generative Models
von: Pourcel, Julien, et al.
Veröffentlicht: (2023)
von: Pourcel, Julien, et al.
Veröffentlicht: (2023)
Measuring the Influence of Incorrect Code on Test Generation
von: Huang, Dong, et al.
Veröffentlicht: (2024)
von: Huang, Dong, et al.
Veröffentlicht: (2024)
Who Wrote this Code? Watermarking for Code Generation
von: Lee, Taehyun, et al.
Veröffentlicht: (2023)
von: Lee, Taehyun, et al.
Veröffentlicht: (2023)
Generalizing Test Cases for Comprehensive Test Scenario Coverage
von: Qi, Binhang, et al.
Veröffentlicht: (2026)
von: Qi, Binhang, et al.
Veröffentlicht: (2026)
Klear-CodeTest: Scalable Test Case Generation for Code Reinforcement Learning
von: Fu, Jia, et al.
Veröffentlicht: (2025)
von: Fu, Jia, et al.
Veröffentlicht: (2025)
Lares: LLM-driven Code Slice Semantic Search for Patch Presence Testing
von: Li, Siyuan, et al.
Veröffentlicht: (2025)
von: Li, Siyuan, et al.
Veröffentlicht: (2025)
Cross-validating causal discovery via Leave-One-Variable-Out
von: Schkoda, Daniela, et al.
Veröffentlicht: (2024)
von: Schkoda, Daniela, et al.
Veröffentlicht: (2024)
MCTS-Judge: Test-Time Scaling in LLM-as-a-Judge for Code Correctness Evaluation
von: Wang, Yutong, et al.
Veröffentlicht: (2025)
von: Wang, Yutong, et al.
Veröffentlicht: (2025)
Leaving No One Behind, Leaving No One Unaccountable
von: Glušac, Luka
Veröffentlicht: (2023)
von: Glušac, Luka
Veröffentlicht: (2023)
VeriScale: Adversarial Test-Suite Scaling for Verifiable Code Generation
von: Bai, Yifan, et al.
Veröffentlicht: (2026)
von: Bai, Yifan, et al.
Veröffentlicht: (2026)
Preserving AUC Fairness in Learning with Noisy Protected Groups
von: Wu, Mingyang, et al.
Veröffentlicht: (2025)
von: Wu, Mingyang, et al.
Veröffentlicht: (2025)
Confidence Intervals for AUC and pAUC by Empirical Likelihood
von: Yumin Zhao, et al.
Veröffentlicht: (2025)
von: Yumin Zhao, et al.
Veröffentlicht: (2025)
Test-time GNN Model Evaluation on Dynamic Graphs
von: Li, Bo, et al.
Veröffentlicht: (2025)
von: Li, Bo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Exploring Pass-Rate Reward in Reinforcement Learning for Code Generation
von: Li, Xin-Ye, et al.
Veröffentlicht: (2026) -
Post-Incorporating Code Structural Knowledge into Pretrained Models via ICL for Code Translation
von: Du, Yali, et al.
Veröffentlicht: (2025) -
Weakly Supervised AUC Optimization: A Unified Partial AUC Approach
von: Xie, Zheng, et al.
Veröffentlicht: (2023) -
Design-Specification Tiling for ICL-based CAD Code Generation
von: Du, Yali, et al.
Veröffentlicht: (2026) -
A Joint Learning Model with Variational Interaction for Multilingual Program Translation
von: Du, Yali, et al.
Veröffentlicht: (2024)