CODECLEANER: Elevating Standards with A Robust Data Contamination Mitigation Toolkit
Fuente:
arXiv
Saved in:
| Main Authors: | Cao, Jialun, Chen, Songqiang, Zhang, Wuqi, Lo, Hau Ching, Cheung, Shing-Chi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Concerned with Data Contamination? Assessing Countermeasures in Code Language Model
by: Cao, Jialun, et al.
Published: (2024)
by: Cao, Jialun, et al.
Published: (2024)
Isolating Language-Coding from Problem-Solving: Benchmarking LLMs with PseudoEval
by: Wu, Jiarong, et al.
Published: (2025)
by: Wu, Jiarong, et al.
Published: (2025)
ModelWisdom: An Integrated Toolkit for TLA+ Model Visualization, Digest and Repair
by: Chen, Zhiyong, et al.
Published: (2026)
by: Chen, Zhiyong, et al.
Published: (2026)
What Builds Effective In-Context Examples for Code Generation?
by: Li, Dongze, et al.
Published: (2025)
by: Li, Dongze, et al.
Published: (2025)
Understanding and Bridging the Planner-Coder Gap: A Systematic Study on the Robustness of Multi-Agent Systems for Code Generation
by: Lyu, Zongyi, et al.
Published: (2025)
by: Lyu, Zongyi, et al.
Published: (2025)
MR-Coupler: Automated Metamorphic Test Generation via Functional Coupling Analysis
by: Xu, Congying, et al.
Published: (2026)
by: Xu, Congying, et al.
Published: (2026)
When LLMs Meet API Documentation: Can Retrieval Augmentation Aid Code Generation Just as It Helps Developers?
by: Chen, Jingyi, et al.
Published: (2025)
by: Chen, Jingyi, et al.
Published: (2025)
CAM: A Causality-based Analysis Framework for Multi-Agent Code Generation Systems
by: Lyu, Zongyi, et al.
Published: (2026)
by: Lyu, Zongyi, et al.
Published: (2026)
Can Emulating Semantic Translation Help LLMs with Code Translation? A Study Based on Pseudocode
by: Chen, Songqiang, et al.
Published: (2025)
by: Chen, Songqiang, et al.
Published: (2025)
DOMAINEVAL: An Auto-Constructed Benchmark for Multi-Domain Code Generation
by: Zhu, Qiming, et al.
Published: (2024)
by: Zhu, Qiming, et al.
Published: (2024)
EmbedAgent: Benchmarking Large Language Models in Embedded System Development
by: Xu, Ruiyang, et al.
Published: (2025)
by: Xu, Ruiyang, et al.
Published: (2025)
MR-Adopt: Automatic Deduction of Input Transformation Function for Metamorphic Testing
by: Xu, Congying, et al.
Published: (2024)
by: Xu, Congying, et al.
Published: (2024)
JavaBench: A Benchmark of Object-Oriented Code Generation for Evaluating Large Language Models
by: Cao, Jialun, et al.
Published: (2024)
by: Cao, Jialun, et al.
Published: (2024)
Word Closure-Based Metamorphic Testing for Machine Translation
by: Xie, Xiaoyuan, et al.
Published: (2023)
by: Xie, Xiaoyuan, et al.
Published: (2023)
RulER: Automated Rule-Based Semantic Error Localization and Repair for Code Translation
by: Jin, Shuo, et al.
Published: (2025)
by: Jin, Shuo, et al.
Published: (2025)
LiveFMBench: Unveiling the Power and Limits of Agentic Workflows in Specification Generation
by: Xu, Dong, et al.
Published: (2026)
by: Xu, Dong, et al.
Published: (2026)
From What to How: Bridging User Requirements with Software Development Using Large Language Models
by: He, Xiao, et al.
Published: (2026)
by: He, Xiao, et al.
Published: (2026)
Enhancing Differential Testing With LLMs For Testing Deep Learning Libraries
by: Li, Meiziniu, et al.
Published: (2024)
by: Li, Meiziniu, et al.
Published: (2024)
COMET: Coverage-guided Model Generation For Deep Learning Library Testing
by: Li, Meiziniu, et al.
Published: (2022)
by: Li, Meiziniu, et al.
Published: (2022)
Multi-Agent Systems for Dataset Adaptation in Software Engineering: Capabilities, Limitations, and Future Directions
by: Chen, Jingyi, et al.
Published: (2025)
by: Chen, Jingyi, et al.
Published: (2025)
Precision in Practice: Knowledge Guided Code Summarizing Grounded in Industrial Expectations
by: Li, Jintai, et al.
Published: (2026)
by: Li, Jintai, et al.
Published: (2026)
Can Large Language Models Model Programs Formally?
by: Chen, Zhiyong, et al.
Published: (2026)
by: Chen, Zhiyong, et al.
Published: (2026)
ReuseDroid: A VLM-empowered Android UI Test Migrator Boosted by Active Feedback
by: Li, Xiaolei, et al.
Published: (2025)
by: Li, Xiaolei, et al.
Published: (2025)
Towards Understanding the Bugs in Solidity Compiler
by: Ma, Haoyang, et al.
Published: (2024)
by: Ma, Haoyang, et al.
Published: (2024)
Automatic Build Repair for Test Cases using Incompatible Java Versions
by: Mak, Ching Hang, et al.
Published: (2024)
by: Mak, Ching Hang, et al.
Published: (2024)
Scaling Coding Agents via Atomic Skills
by: Ma, Yingwei, et al.
Published: (2026)
by: Ma, Yingwei, et al.
Published: (2026)
Deep Learning for Code Intelligence: Survey, Benchmark and Toolkit
by: Wan, Yao, et al.
Published: (2023)
by: Wan, Yao, et al.
Published: (2023)
Rigor, Reliability, and Reproducibility Matter: A Decade-Scale Survey of 572 Code Benchmarks
by: Cao, Jialun, et al.
Published: (2025)
by: Cao, Jialun, et al.
Published: (2025)
In-IDE Toolkit for Developers of AI-Based Features
by: Sokolov, Yaroslav, et al.
Published: (2026)
by: Sokolov, Yaroslav, et al.
Published: (2026)
From Expectation to Habit: Why Do Software Practitioners Adopt Fairness Toolkits?
by: Voria, Gianmario, et al.
Published: (2024)
by: Voria, Gianmario, et al.
Published: (2024)
Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data Contamination
by: Chen, Simin, et al.
Published: (2025)
by: Chen, Simin, et al.
Published: (2025)
Demystifying the Silence of Correctness Bugs in PyTorch Compiler
by: Li, Meiziniu, et al.
Published: (2026)
by: Li, Meiziniu, et al.
Published: (2026)
Bias Testing and Mitigation in LLM-based Code Generation
by: Huang, Dong, et al.
Published: (2023)
by: Huang, Dong, et al.
Published: (2023)
Across Programming Language Silos: A Study on Cross-Lingual Retrieval-augmented Code Generation
by: Zhu, Qiming, et al.
Published: (2025)
by: Zhu, Qiming, et al.
Published: (2025)
Fairness Is Not Just Ethical: Performance Trade-Off via Data Correlation Tuning to Mitigate Bias in ML Software
by: Xiao, Ying, et al.
Published: (2025)
by: Xiao, Ying, et al.
Published: (2025)
LLM-Powered Detection of Price Manipulation in DeFi
by: Liu, Lu, et al.
Published: (2025)
by: Liu, Lu, et al.
Published: (2025)
ReusStdFlow: A Standardized Reusability Framework for Dynamic Workflow Construction in Agentic AI
by: Zhang, Gaoyang, et al.
Published: (2026)
by: Zhang, Gaoyang, et al.
Published: (2026)
A Systematic Literature Review on Explainability for Machine/Deep Learning-based Software Engineering Research
by: Cao, Sicong, et al.
Published: (2024)
by: Cao, Sicong, et al.
Published: (2024)
Agentic Software Issue Resolution with Large Language Models: A Survey
by: Jiang, Zhonghao, et al.
Published: (2025)
by: Jiang, Zhonghao, et al.
Published: (2025)
Towards Fair Machine Learning Software: Understanding and Addressing Model Bias Through Counterfactual Thinking
by: Wang, Zichong, et al.
Published: (2023)
by: Wang, Zichong, et al.
Published: (2023)
Similar Items
-
Concerned with Data Contamination? Assessing Countermeasures in Code Language Model
by: Cao, Jialun, et al.
Published: (2024) -
Isolating Language-Coding from Problem-Solving: Benchmarking LLMs with PseudoEval
by: Wu, Jiarong, et al.
Published: (2025) -
ModelWisdom: An Integrated Toolkit for TLA+ Model Visualization, Digest and Repair
by: Chen, Zhiyong, et al.
Published: (2026) -
What Builds Effective In-Context Examples for Code Generation?
by: Li, Dongze, et al.
Published: (2025) -
Understanding and Bridging the Planner-Coder Gap: A Systematic Study on the Robustness of Multi-Agent Systems for Code Generation
by: Lyu, Zongyi, et al.
Published: (2025)