Demystifying the Silence of Correctness Bugs in PyTorch Compiler
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Meiziniu, Li, Dongze, Liu, Jianmeng, Cheung, Shing-Chi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enhancing Differential Testing With LLMs For Testing Deep Learning Libraries
by: Li, Meiziniu, et al.
Published: (2024)
by: Li, Meiziniu, et al.
Published: (2024)
COMET: Coverage-guided Model Generation For Deep Learning Library Testing
by: Li, Meiziniu, et al.
Published: (2022)
by: Li, Meiziniu, et al.
Published: (2022)
Bug In the Code Stack: Can LLMs Find Bugs in Large Python Code Stacks
by: Lee, Hokyung, et al.
Published: (2024)
by: Lee, Hokyung, et al.
Published: (2024)
LSPFuzz: Hunting Bugs in Language Servers
by: Zhu, Hengcheng, et al.
Published: (2025)
by: Zhu, Hengcheng, et al.
Published: (2025)
Abstain and Validate: A Dual-LLM Policy for Reducing Noise in Agentic Program Repair
by: Cambronero, José, et al.
Published: (2025)
by: Cambronero, José, et al.
Published: (2025)
AgentAssay: Token-Efficient Regression Testing for Non-Deterministic AI Agent Workflows
by: Bhardwaj, Varun Pratap
Published: (2026)
by: Bhardwaj, Varun Pratap
Published: (2026)
On the Mistaken Assumption of Interchangeable Deep Reinforcement Learning Implementations
by: Hundal, Rajdeep Singh, et al.
Published: (2025)
by: Hundal, Rajdeep Singh, et al.
Published: (2025)
From Untestable to Testable: Metamorphic Testing in the Age of LLMs
by: Terragni, Valerio
Published: (2026)
by: Terragni, Valerio
Published: (2026)
Assessing Data Augmentation-Induced Bias in Training and Testing of Machine Learning Models
by: More, Riddhi, et al.
Published: (2025)
by: More, Riddhi, et al.
Published: (2025)
An Analysis of LLM Fine-Tuning and Few-Shot Learning for Flaky Test Detection and Classification
by: More, Riddhi, et al.
Published: (2025)
by: More, Riddhi, et al.
Published: (2025)
Towards Explainable Test Case Prioritisation with Learning-to-Rank Models
by: Ramírez, Aurora, et al.
Published: (2024)
by: Ramírez, Aurora, et al.
Published: (2024)
A Systematic Approach for Assessing Large Language Models' Test Case Generation Capability
by: Chang, Hung-Fu, et al.
Published: (2025)
by: Chang, Hung-Fu, et al.
Published: (2025)
Test-Driven AI Agent Definition (TDAD): Compiling Tool-Using Agents from Behavioral Specifications
by: Rehan, Tzafrir
Published: (2026)
by: Rehan, Tzafrir
Published: (2026)
AuditRepairBench: A Paired-Execution Trace Corpus for Evaluator-Channel Ranking Instability in Agent Repair
by: Hu, Yuelin, et al.
Published: (2026)
by: Hu, Yuelin, et al.
Published: (2026)
CoTran: An LLM-based Code Translator using Reinforcement Learning with Feedback from Compiler and Symbolic Execution
by: Jana, Prithwish, et al.
Published: (2023)
by: Jana, Prithwish, et al.
Published: (2023)
AutoBridge: Automating Smart Device Integration with Centralized Platform
by: Liu, Siyuan, et al.
Published: (2025)
by: Liu, Siyuan, et al.
Published: (2025)
GALA: Multimodal Graph Alignment for Bug Localization in Automated Program Repair
by: Liu, Zhuoyao, et al.
Published: (2026)
by: Liu, Zhuoyao, et al.
Published: (2026)
RMCBench: Benchmarking Large Language Models' Resistance to Malicious Code
by: Chen, Jiachi, et al.
Published: (2024)
by: Chen, Jiachi, et al.
Published: (2024)
SpecOps: A Fully Automated AI Agent Testing Framework in Real-World GUI Environments
by: Ahmed, Syed Yusuf, et al.
Published: (2026)
by: Ahmed, Syed Yusuf, et al.
Published: (2026)
Addressing Data Leakage in HumanEval Using Combinatorial Test Design
by: Bradbury, Jeremy S., et al.
Published: (2024)
by: Bradbury, Jeremy S., et al.
Published: (2024)
Experience with GitHub Copilot for Developer Productivity at Zoominfo
by: Bakal, Gal, et al.
Published: (2025)
by: Bakal, Gal, et al.
Published: (2025)
CodeTracer: Towards Traceable Agent States
by: Li, Han, et al.
Published: (2026)
by: Li, Han, et al.
Published: (2026)
Mining Bug Repositories for Multi-Fault Programs
by: Callaghan, Dylan, et al.
Published: (2024)
by: Callaghan, Dylan, et al.
Published: (2024)
The Explabox: Model-Agnostic Machine Learning Transparency & Analysis
by: Robeer, Marcel, et al.
Published: (2024)
by: Robeer, Marcel, et al.
Published: (2024)
Runtime Execution Traces Guided Automated Program Repair with Multi-Agent Debate
by: Wu, Jiaqing, et al.
Published: (2026)
by: Wu, Jiaqing, et al.
Published: (2026)
Review Beats Planning: Dual-Model Interaction Patterns for Code Synthesis
by: Miller, Jan
Published: (2026)
by: Miller, Jan
Published: (2026)
CodeEvolve: LLM-Driven Evolutionary Optimization with Runtime-Enriched Target Selection for Multi-Language Code Enhancement
by: Borra, Ajay Krishna, et al.
Published: (2026)
by: Borra, Ajay Krishna, et al.
Published: (2026)
Empirical Analysis of Temporal and Spatial Fault Characteristics in Multi-Fault Bug Repositories
by: Callaghan, Dylan, et al.
Published: (2025)
by: Callaghan, Dylan, et al.
Published: (2025)
Monitoring Agentic Systems Before They're Reliable
by: Boston, Marisa Ferrara, et al.
Published: (2026)
by: Boston, Marisa Ferrara, et al.
Published: (2026)
AIRA: AI-Induced Risk Audit: A Structured Inspection Framework for AI-Generated Code
by: Parris, William M.
Published: (2026)
by: Parris, William M.
Published: (2026)
Towards a Probabilistic Framework for Analyzing and Improving LLM-Enabled Software
by: Baldonado, Juan Manuel, et al.
Published: (2025)
by: Baldonado, Juan Manuel, et al.
Published: (2025)
Multi-Agent Code Verification via Information Theory
by: Rajan, Shreshth
Published: (2025)
by: Rajan, Shreshth
Published: (2025)
MeDeT: Medical Device Digital Twins Creation with Few-shot Meta-learning
by: Sartaj, Hassan, et al.
Published: (2024)
by: Sartaj, Hassan, et al.
Published: (2024)
Navigating the growing field of research on AI for software testing -- the taxonomy for AI-augmented software testing and an ontology-driven literature survey
by: Schieferdecker, Ina K.
Published: (2025)
by: Schieferdecker, Ina K.
Published: (2025)
LSPRAG: LSP-Guided RAG for Language-Agnostic Real-Time Unit Test Generation
by: Go, Gwihwan, et al.
Published: (2025)
by: Go, Gwihwan, et al.
Published: (2025)
Understanding and Detecting Flaky Builds in GitHub Actions
by: Ge, Wenhao, et al.
Published: (2026)
by: Ge, Wenhao, et al.
Published: (2026)
AcTracer: Active Testing of Large Language Model via Multi-Stage Sampling
by: Huang, Yuheng, et al.
Published: (2024)
by: Huang, Yuheng, et al.
Published: (2024)
Fuzzing the brain: Automated stress testing for the safety of ML-driven neurostimulation
by: Downing, Mara, et al.
Published: (2025)
by: Downing, Mara, et al.
Published: (2025)
Automated Code Fix Suggestions for Accessibility Issues in Mobile Apps
by: Mehralian, Forough, et al.
Published: (2024)
by: Mehralian, Forough, et al.
Published: (2024)
DeepCodeProbe: Towards Understanding What Models Trained on Code Learn
by: Majdinasab, Vahid, et al.
Published: (2024)
by: Majdinasab, Vahid, et al.
Published: (2024)
Similar Items
-
Enhancing Differential Testing With LLMs For Testing Deep Learning Libraries
by: Li, Meiziniu, et al.
Published: (2024) -
COMET: Coverage-guided Model Generation For Deep Learning Library Testing
by: Li, Meiziniu, et al.
Published: (2022) -
Bug In the Code Stack: Can LLMs Find Bugs in Large Python Code Stacks
by: Lee, Hokyung, et al.
Published: (2024) -
LSPFuzz: Hunting Bugs in Language Servers
by: Zhu, Hengcheng, et al.
Published: (2025) -
Abstain and Validate: A Dual-LLM Policy for Reducing Noise in Agentic Program Repair
by: Cambronero, José, et al.
Published: (2025)