scicode-lint: Detecting Methodology Bugs in Scientific Python Code with LLM-Generated Patterns
Fuente:
arXiv
Enregistré dans:
| Auteur principal: | Samsonau, Sergey V. |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
An Empirical Evaluation of Locally Deployed LLMs for Bug Detection in Python Code
par: Vulićević, Jelena Ilić
Publié: (2026)
par: Vulićević, Jelena Ilić
Publié: (2026)
sciwrite-lint: Verification Infrastructure for the Age of Science Vibe-Writing
par: Samsonau, Sergey V
Publié: (2026)
par: Samsonau, Sergey V
Publié: (2026)
On the Impact of Code Comments for Automated Bug-Fixing: An Empirical Study
par: Vitale, Antonio, et autres
Publié: (2026)
par: Vitale, Antonio, et autres
Publié: (2026)
Are Sparse Autoencoders Useful for Java Function Bug Detection?
par: Melo, Rui, et autres
Publié: (2025)
par: Melo, Rui, et autres
Publié: (2025)
SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents
par: Mündler, Niels, et autres
Publié: (2024)
par: Mündler, Niels, et autres
Publié: (2024)
Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code
par: Galimzyanov, Timur, et autres
Publié: (2024)
par: Galimzyanov, Timur, et autres
Publié: (2024)
Enhancing LLM-Based Test Generation by Eliminating Covered Code
par: Xu, WeiZhe, et autres
Publié: (2026)
par: Xu, WeiZhe, et autres
Publié: (2026)
Semantic Voting: Execution-Grounded Consensus for LLM Code Generation
par: Jiang, Shan, et autres
Publié: (2026)
par: Jiang, Shan, et autres
Publié: (2026)
Automatic Detection of LLM-Generated Code: A Comparative Case Study of Contemporary Models Across Function and Class Granularities
par: Rahman, Musfiqur, et autres
Publié: (2024)
par: Rahman, Musfiqur, et autres
Publié: (2024)
Are Large Language Models Memorizing Bug Benchmarks?
par: Ramos, Daniel, et autres
Publié: (2024)
par: Ramos, Daniel, et autres
Publié: (2024)
How Efficient is LLM-Generated Code? A Rigorous & High-Standard Benchmark
par: Qiu, Ruizhong, et autres
Publié: (2024)
par: Qiu, Ruizhong, et autres
Publié: (2024)
OpenClassGen: A Large-Scale Corpus of Real-World Python Classes for LLM Research
par: Rahman, Musfiqur, et autres
Publié: (2025)
par: Rahman, Musfiqur, et autres
Publié: (2025)
Beyond Synthetic Benchmarks: Evaluating LLM Performance on Real-World Class-Level Code Generation
par: Rahman, Musfiqur, et autres
Publié: (2025)
par: Rahman, Musfiqur, et autres
Publié: (2025)
I Know Which LLM Wrote Your Code Last Summer: LLM generated Code Stylometry for Authorship Attribution
par: Bisztray, Tamas, et autres
Publié: (2025)
par: Bisztray, Tamas, et autres
Publié: (2025)
Imitation Game: Reproducing Deep Learning Bugs Leveraging an Intelligent Agent
par: Shah, Mehil B, et autres
Publié: (2025)
par: Shah, Mehil B, et autres
Publié: (2025)
Towards a Neural Debugger for Python
par: Beck, Maximilian, et autres
Publié: (2026)
par: Beck, Maximilian, et autres
Publié: (2026)
GREPO: A Benchmark for Graph Neural Networks on Repository-Level Bug Localization
par: Wang, Juntong, et autres
Publié: (2026)
par: Wang, Juntong, et autres
Publié: (2026)
TriagerX: Dual Transformers for Bug Triaging Tasks with Content and Interaction Based Rankings
par: Mamun, Md Afif Al, et autres
Publié: (2025)
par: Mamun, Md Afif Al, et autres
Publié: (2025)
A Stochastic Differential Equation Framework for Multi-Objective LLM Interactions: Dynamical Systems Analysis with Code Generation Applications
par: Shukla, Shivani, et autres
Publié: (2025)
par: Shukla, Shivani, et autres
Publié: (2025)
CodeTaste: Can LLMs Generate Human-Level Code Refactorings?
par: Thillen, Alex, et autres
Publié: (2026)
par: Thillen, Alex, et autres
Publié: (2026)
Operational Robustness of LLMs on Code Generation
par: Paul, Debalina Ghosh, et autres
Publié: (2026)
par: Paul, Debalina Ghosh, et autres
Publié: (2026)
Can Coding Agents Be General Agents?
par: Ivanov, Maksim, et autres
Publié: (2026)
par: Ivanov, Maksim, et autres
Publié: (2026)
Optimizing AI-Assisted Code Generation
par: Torka, Simon, et autres
Publié: (2024)
par: Torka, Simon, et autres
Publié: (2024)
The Struggles of LLMs in Cross-lingual Code Clone Detection
par: Moumoula, Micheline Bénédicte, et autres
Publié: (2024)
par: Moumoula, Micheline Bénédicte, et autres
Publié: (2024)
Code Generation by Differential Test Time Scaling
par: He, Yifeng, et autres
Publié: (2026)
par: He, Yifeng, et autres
Publié: (2026)
Clover: Closed-Loop Verifiable Code Generation
par: Sun, Chuyue, et autres
Publié: (2023)
par: Sun, Chuyue, et autres
Publié: (2023)
Functional Overlap Reranking for Neural Code Generation
par: To, Hung Quoc, et autres
Publié: (2023)
par: To, Hung Quoc, et autres
Publié: (2023)
Functional Programming Paradigm of Python for Scientific Computation Pipeline Integration
par: Zhang, Chen, et autres
Publié: (2024)
par: Zhang, Chen, et autres
Publié: (2024)
A Survey on Code Generation with LLM-based Agents
par: Dong, Yihong, et autres
Publié: (2025)
par: Dong, Yihong, et autres
Publié: (2025)
Benchmarking Reward Hack Detection in Code Environments via Contrastive Analysis
par: Deshpande, Darshan, et autres
Publié: (2026)
par: Deshpande, Darshan, et autres
Publié: (2026)
A Theoretical Analysis of Test-Driven Code Generation
par: Menet, Nicolas, et autres
Publié: (2026)
par: Menet, Nicolas, et autres
Publié: (2026)
Protocode: Prototype-Driven Interpretability for Code Generation in LLMs
par: Bodla, Krishna Vamshi, et autres
Publié: (2025)
par: Bodla, Krishna Vamshi, et autres
Publié: (2025)
Empirical Analysis and Detection of Hallucinations in LLM-Generated Bug Report Summaries
par: Nirujan, Hinduja, et autres
Publié: (2026)
par: Nirujan, Hinduja, et autres
Publié: (2026)
CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X
par: Zheng, Qinkai, et autres
Publié: (2023)
par: Zheng, Qinkai, et autres
Publié: (2023)
Exploring Pass-Rate Reward in Reinforcement Learning for Code Generation
par: Li, Xin-Ye, et autres
Publié: (2026)
par: Li, Xin-Ye, et autres
Publié: (2026)
Deep-Bench: Deep Learning Benchmark Dataset for Code Generation
par: Daghighfarsoodeh, Alireza, et autres
Publié: (2025)
par: Daghighfarsoodeh, Alireza, et autres
Publié: (2025)
Supersonic: Learning to Generate Source Code Optimizations in C/C++
par: Chen, Zimin, et autres
Publié: (2023)
par: Chen, Zimin, et autres
Publié: (2023)
A Survey of Bugs in AI-Generated Code
par: Gao, Ruofan, et autres
Publié: (2025)
par: Gao, Ruofan, et autres
Publié: (2025)
Co-Located Tests, Better AI Code: How Test Syntax Structure Affects Foundation Model Code Generation
par: Jacopin, Éric
Publié: (2026)
par: Jacopin, Éric
Publié: (2026)
LLM Benchmarking with LLaMA2: Evaluating Code Development Performance Across Multiple Programming Languages
par: Diehl, Patrick, et autres
Publié: (2025)
par: Diehl, Patrick, et autres
Publié: (2025)
Documents similaires
-
An Empirical Evaluation of Locally Deployed LLMs for Bug Detection in Python Code
par: Vulićević, Jelena Ilić
Publié: (2026) -
sciwrite-lint: Verification Infrastructure for the Age of Science Vibe-Writing
par: Samsonau, Sergey V
Publié: (2026) -
On the Impact of Code Comments for Automated Bug-Fixing: An Empirical Study
par: Vitale, Antonio, et autres
Publié: (2026) -
Are Sparse Autoencoders Useful for Java Function Bug Detection?
par: Melo, Rui, et autres
Publié: (2025) -
SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents
par: Mündler, Niels, et autres
Publié: (2024)