Guarded Repair for Harm-Aware Post-hoc Replacement of LLM Mathematical Reasoning
Fuente:
arXiv
Saved in:
| Main Author: | Xia, Haizhou |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
by: Tu, Xinming, et al.
Published: (2026)
by: Tu, Xinming, et al.
Published: (2026)
IRepair: An Intent-Aware Approach to Repair Data-Driven Errors in Large Language Models
by: Imtiaz, Sayem Mohammad, et al.
Published: (2025)
by: Imtiaz, Sayem Mohammad, et al.
Published: (2025)
Dissecting the SWE-Bench Leaderboards: Profiling Submitters and Architectures of LLM- and Agent-Based Repair Systems
by: Martinez, Matias, et al.
Published: (2025)
by: Martinez, Matias, et al.
Published: (2025)
Benchmarking Open-Source Safety Guard Models: A Comprehensive Evaluation
by: Harsh, Reetu Raj, et al.
Published: (2026)
by: Harsh, Reetu Raj, et al.
Published: (2026)
Pragmatic Reasoning improves LLM Code Generation
by: Cao, Zhuchen, et al.
Published: (2025)
by: Cao, Zhuchen, et al.
Published: (2025)
MemRepair: Hierarchical Memory for Agentic Repository-Level Vulnerability Repair
by: Liu, Simiao, et al.
Published: (2026)
by: Liu, Simiao, et al.
Published: (2026)
Post-hoc LLM-Supported Debugging of Distributed Processes
by: Schiese, Dennis, et al.
Published: (2025)
by: Schiese, Dennis, et al.
Published: (2025)
SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
by: Wei, Yuxiang, et al.
Published: (2025)
by: Wei, Yuxiang, et al.
Published: (2025)
Reasoning Efficiently Through Adaptive Chain-of-Thought Compression: A Self-Optimizing Framework
by: Huang, Kerui, et al.
Published: (2025)
by: Huang, Kerui, et al.
Published: (2025)
Towards Automated Smart Contract Generation: Evaluation, Benchmarking, and Retrieval-Augmented Repair
by: Chen, Zaoyu, et al.
Published: (2025)
by: Chen, Zaoyu, et al.
Published: (2025)
An LLM-as-Judge Metric for Bridging the Gap with Human Evaluation in SE Tasks
by: Zhou, Xin, et al.
Published: (2025)
by: Zhou, Xin, et al.
Published: (2025)
Exploring Large Language Models in Resolving Environment-Related Crash Bugs: Localizing and Repairing
by: Du, Xueying, et al.
Published: (2023)
by: Du, Xueying, et al.
Published: (2023)
Learner-Tailored Program Repair: A Solution Generator with Iterative Edit-Driven Retrieval Enhancement
by: Dai, Zhenlong, et al.
Published: (2026)
by: Dai, Zhenlong, et al.
Published: (2026)
Code Repair with LLMs gives an Exploration-Exploitation Tradeoff
by: Tang, Hao, et al.
Published: (2024)
by: Tang, Hao, et al.
Published: (2024)
Is Self-Repair a Silver Bullet for Code Generation?
by: Olausson, Theo X., et al.
Published: (2023)
by: Olausson, Theo X., et al.
Published: (2023)
Benchmarking Educational Program Repair
by: Koutcheme, Charles, et al.
Published: (2024)
by: Koutcheme, Charles, et al.
Published: (2024)
EigenData: A Self-Evolving Multi-Agent Platform for Function-Calling Data Synthesis, Auditing, and Repair
by: Chen, Jiaao, et al.
Published: (2026)
by: Chen, Jiaao, et al.
Published: (2026)
PrimeGuard: Safe and Helpful LLMs through Tuning-Free Routing
by: Manczak, Blazej, et al.
Published: (2024)
by: Manczak, Blazej, et al.
Published: (2024)
R2Vul: Learning to Reason about Software Vulnerabilities with Reinforcement Learning and Structured Reasoning Distillation
by: Weyssow, Martin, et al.
Published: (2025)
by: Weyssow, Martin, et al.
Published: (2025)
The Art of Repair: Optimizing Iterative Program Repair with Instruction-Tuned Models
by: Ruiz, Fernando Vallecillos, et al.
Published: (2025)
by: Ruiz, Fernando Vallecillos, et al.
Published: (2025)
Structure-Aware Fill-in-the-Middle Pretraining for Code
by: Gong, Linyuan, et al.
Published: (2025)
by: Gong, Linyuan, et al.
Published: (2025)
Verification Limits Code LLM Training
by: Gureja, Srishti, et al.
Published: (2025)
by: Gureja, Srishti, et al.
Published: (2025)
Crystal: Illuminating LLM Abilities on Language and Code
by: Tao, Tianhua, et al.
Published: (2024)
by: Tao, Tianhua, et al.
Published: (2024)
CRANE: Constrained Reasoning Injection for Code Agents via Nullspace Editing
by: Zhu, Mingzhi, et al.
Published: (2026)
by: Zhu, Mingzhi, et al.
Published: (2026)
SemCoder: Training Code Language Models with Comprehensive Semantics Reasoning
by: Ding, Yangruibo, et al.
Published: (2024)
by: Ding, Yangruibo, et al.
Published: (2024)
Planning to Explore: Curiosity-Driven Planning for LLM Test Generation
by: Amayuelas, Alfonso, et al.
Published: (2026)
by: Amayuelas, Alfonso, et al.
Published: (2026)
A Case Study of Web App Coding with OpenAI Reasoning Models
by: Cui, Yi
Published: (2024)
by: Cui, Yi
Published: (2024)
An evaluation of LLM code generation capabilities through graded exercises
by: Jiménez, Álvaro Barbero
Published: (2024)
by: Jiménez, Álvaro Barbero
Published: (2024)
ChainStream: An LLM-based Framework for Unified Synthetic Sensing
by: Liu, Jiacheng, et al.
Published: (2024)
by: Liu, Jiacheng, et al.
Published: (2024)
LocAgent: Graph-Guided LLM Agents for Code Localization
by: Chen, Zhaoling, et al.
Published: (2025)
by: Chen, Zhaoling, et al.
Published: (2025)
LLM Hallucinations in Practical Code Generation: Phenomena, Mechanism, and Mitigation
by: Zhang, Ziyao, et al.
Published: (2024)
by: Zhang, Ziyao, et al.
Published: (2024)
Learning to Ask: When LLM Agents Meet Unclear Instruction
by: Wang, Wenxuan, et al.
Published: (2024)
by: Wang, Wenxuan, et al.
Published: (2024)
Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data Contamination
by: Chen, Simin, et al.
Published: (2025)
by: Chen, Simin, et al.
Published: (2025)
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
by: Jiang, Hongchao, et al.
Published: (2025)
by: Jiang, Hongchao, et al.
Published: (2025)
WorkflowLLM: Enhancing Workflow Orchestration Capability of Large Language Models
by: Fan, Shengda, et al.
Published: (2024)
by: Fan, Shengda, et al.
Published: (2024)
Automated Business Process Analysis: An LLM-Based Approach to Value Assessment
by: De Michele, William, et al.
Published: (2025)
by: De Michele, William, et al.
Published: (2025)
Collaboration is all you need: LLM Assisted Safe Code Translation
by: Karanjai, Rabimba, et al.
Published: (2025)
by: Karanjai, Rabimba, et al.
Published: (2025)
Specifications: The missing link to making the development of LLM systems an engineering discipline
by: Stoica, Ion, et al.
Published: (2024)
by: Stoica, Ion, et al.
Published: (2024)
Issue Localization via LLM-Driven Iterative Code Graph Searching
by: Jiang, Zhonghao, et al.
Published: (2025)
by: Jiang, Zhonghao, et al.
Published: (2025)
Revision or Re-Solving? Decomposing Second-Pass Gains in Multi-LLM Pipelines
by: Ning, Jingjie, et al.
Published: (2026)
by: Ning, Jingjie, et al.
Published: (2026)
Similar Items
-
BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
by: Tu, Xinming, et al.
Published: (2026) -
IRepair: An Intent-Aware Approach to Repair Data-Driven Errors in Large Language Models
by: Imtiaz, Sayem Mohammad, et al.
Published: (2025) -
Dissecting the SWE-Bench Leaderboards: Profiling Submitters and Architectures of LLM- and Agent-Based Repair Systems
by: Martinez, Matias, et al.
Published: (2025) -
Benchmarking Open-Source Safety Guard Models: A Comprehensive Evaluation
by: Harsh, Reetu Raj, et al.
Published: (2026) -
Pragmatic Reasoning improves LLM Code Generation
by: Cao, Zhuchen, et al.
Published: (2025)