Investigating Training Data Detection in AI Coders
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Tianlin, Wei, Yunxiang, Li, Zhiming, Liu, Aishan, Guo, Qing, Liu, Xianglong, Sun, Dongning, Liu, Yang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
From Context to Intent: Reasoning-Guided Function-Level Code Completion
von: Li, Yanzhou, et al.
Veröffentlicht: (2025)
von: Li, Yanzhou, et al.
Veröffentlicht: (2025)
Software Development Life Cycle Perspective: A Survey of Benchmarks for Code Large Language Models and Agents
von: Wang, Kaixin, et al.
Veröffentlicht: (2025)
von: Wang, Kaixin, et al.
Veröffentlicht: (2025)
Latent Imitator: Generating Natural Individual Discriminatory Instances for Black-Box Fairness Testing
von: Xiao, Yisong, et al.
Veröffentlicht: (2023)
von: Xiao, Yisong, et al.
Veröffentlicht: (2023)
InCoder-32B: Code Foundation Model for Industrial Scenarios
von: Yang, Jian, et al.
Veröffentlicht: (2026)
von: Yang, Jian, et al.
Veröffentlicht: (2026)
Unveiling Project-Specific Bias in Neural Code Models
von: Li, Zhiming, et al.
Veröffentlicht: (2022)
von: Li, Zhiming, et al.
Veröffentlicht: (2022)
FullStack Bench: Evaluating LLMs as Full Stack Coders
von: Bytedance-Seed-Foundation-Code-Team, et al.
Veröffentlicht: (2024)
von: Bytedance-Seed-Foundation-Code-Team, et al.
Veröffentlicht: (2024)
IQuest-Coder-V1 Technical Report
von: Yang, Jian, et al.
Veröffentlicht: (2026)
von: Yang, Jian, et al.
Veröffentlicht: (2026)
WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation Metrics
von: Liu, Chenxu, et al.
Veröffentlicht: (2026)
von: Liu, Chenxu, et al.
Veröffentlicht: (2026)
JumpCoder: Go Beyond Autoregressive Coder via Online Modification
von: Chen, Mouxiang, et al.
Veröffentlicht: (2024)
von: Chen, Mouxiang, et al.
Veröffentlicht: (2024)
Code2Bench: Scaling Source and Rigor for Dynamic Benchmark Construction
von: Zhang, Zhe, et al.
Veröffentlicht: (2025)
von: Zhang, Zhe, et al.
Veröffentlicht: (2025)
AI Coders Are Among Us: Rethinking Programming Language Grammar Towards Efficient Code Generation
von: Sun, Zhensu, et al.
Veröffentlicht: (2024)
von: Sun, Zhensu, et al.
Veröffentlicht: (2024)
Unveiling Code Pre-Trained Models: Investigating Syntax and Semantics Capacities
von: Ma, Wei, et al.
Veröffentlicht: (2022)
von: Ma, Wei, et al.
Veröffentlicht: (2022)
FastCoder: Accelerating Repository-level Code Generation via Efficient Retrieval and Verification
von: Zhao, Qianhui, et al.
Veröffentlicht: (2025)
von: Zhao, Qianhui, et al.
Veröffentlicht: (2025)
PerfCoder: Large Language Models for Interpretable Code Performance Optimization
von: Yang, Jiuding, et al.
Veröffentlicht: (2025)
von: Yang, Jiuding, et al.
Veröffentlicht: (2025)
TransCoder: Towards Unified Transferable Code Representation Learning Inspired by Human Skills
von: Sun, Qiushi, et al.
Veröffentlicht: (2023)
von: Sun, Qiushi, et al.
Veröffentlicht: (2023)
StarCoder 2 and The Stack v2: The Next Generation
von: Lozhkov, Anton, et al.
Veröffentlicht: (2024)
von: Lozhkov, Anton, et al.
Veröffentlicht: (2024)
RedCoder: Automated Multi-Turn Red Teaming for Code LLMs
von: Mo, Wenjie Jacky, et al.
Veröffentlicht: (2025)
von: Mo, Wenjie Jacky, et al.
Veröffentlicht: (2025)
BDefects4NN: A Backdoor Defect Database for Controlled Localization Studies in Neural Networks
von: Xiao, Yisong, et al.
Veröffentlicht: (2024)
von: Xiao, Yisong, et al.
Veröffentlicht: (2024)
When Prompt Engineering Meets Software Engineering: CNL-P as Natural and Robust "APIs'' for Human-AI Interaction
von: Xing, Zhenchang, et al.
Veröffentlicht: (2025)
von: Xing, Zhenchang, et al.
Veröffentlicht: (2025)
AutoCoder: Enhancing Code Large Language Model with \textsc{AIEV-Instruct}
von: Lei, Bin, et al.
Veröffentlicht: (2024)
von: Lei, Bin, et al.
Veröffentlicht: (2024)
BabelCoder: Agentic Code Translation with Specification Alignment
von: Rabbi, Fazle, et al.
Veröffentlicht: (2025)
von: Rabbi, Fazle, et al.
Veröffentlicht: (2025)
o1-Coder: an o1 Replication for Coding
von: Zhang, Yuxiang, et al.
Veröffentlicht: (2024)
von: Zhang, Yuxiang, et al.
Veröffentlicht: (2024)
TraceCoder: A Trace-Driven Multi-Agent Framework for Automated Debugging of LLM-Generated Code
von: Huang, Jiangping, et al.
Veröffentlicht: (2026)
von: Huang, Jiangping, et al.
Veröffentlicht: (2026)
SemCoder: Training Code Language Models with Comprehensive Semantics Reasoning
von: Ding, Yangruibo, et al.
Veröffentlicht: (2024)
von: Ding, Yangruibo, et al.
Veröffentlicht: (2024)
Can AI Models Direct Each Other? Organizational Structure as a Probe into Training Limitations
von: Liu, Rui
Veröffentlicht: (2026)
von: Liu, Rui
Veröffentlicht: (2026)
HTMLCure: Turning Browser Experience into State Guided Repair for Interactive HTML
von: Wu, Jiajun, et al.
Veröffentlicht: (2026)
von: Wu, Jiajun, et al.
Veröffentlicht: (2026)
Ensemble-Based Uncertainty Estimation for Code Correctness Estimation
von: Wei, Yunxiang, et al.
Veröffentlicht: (2026)
von: Wei, Yunxiang, et al.
Veröffentlicht: (2026)
CodeChemist: Functional Knowledge Transfer for Low-Resource Code Generation via Test-Time Scaling
von: Wang, Kaixin, et al.
Veröffentlicht: (2025)
von: Wang, Kaixin, et al.
Veröffentlicht: (2025)
Rethinking Testing for LLM Applications: Characteristics, Challenges, and a Lightweight Interaction Protocol
von: Ma, Wei, et al.
Veröffentlicht: (2025)
von: Ma, Wei, et al.
Veröffentlicht: (2025)
WybeCoder: Verified Imperative Code Generation
von: Gloeckle, Fabian, et al.
Veröffentlicht: (2026)
von: Gloeckle, Fabian, et al.
Veröffentlicht: (2026)
AlignCoder: Aligning Retrieval with Target Intent for Repository-Level Code Completion
von: Jiang, Tianyue, et al.
Veröffentlicht: (2026)
von: Jiang, Tianyue, et al.
Veröffentlicht: (2026)
ObscuraCoder: Powering Efficient Code LM Pre-Training Via Obfuscation Grounding
von: Paul, Indraneil, et al.
Veröffentlicht: (2025)
von: Paul, Indraneil, et al.
Veröffentlicht: (2025)
SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks
von: Guo, Lianghong, et al.
Veröffentlicht: (2025)
von: Guo, Lianghong, et al.
Veröffentlicht: (2025)
RTLRepoCoder: Repository-Level RTL Code Completion through the Combination of Fine-Tuning and Retrieval Augmentation
von: Wu, Peiyang, et al.
Veröffentlicht: (2025)
von: Wu, Peiyang, et al.
Veröffentlicht: (2025)
State-of-the-art Small Language Coder Model: Mify-Coder
von: Parmar, Abhinav, et al.
Veröffentlicht: (2025)
von: Parmar, Abhinav, et al.
Veröffentlicht: (2025)
ShortCoder: Knowledge-Augmented Syntax Optimization for Token-Efficient Code Generation
von: Liu, Sicong, et al.
Veröffentlicht: (2026)
von: Liu, Sicong, et al.
Veröffentlicht: (2026)
VIBEPASS: Can Vibe Coders Really Pass the Vibe Check?
von: Bansal, Srijan, et al.
Veröffentlicht: (2026)
von: Bansal, Srijan, et al.
Veröffentlicht: (2026)
InverseCoder: Self-improving Instruction-Tuned Code LLMs with Inverse-Instruct
von: Wu, Yutong, et al.
Veröffentlicht: (2024)
von: Wu, Yutong, et al.
Veröffentlicht: (2024)
Automated Snippet-Alignment Data Augmentation for Code Translation
von: Zhang, Zhiming, et al.
Veröffentlicht: (2025)
von: Zhang, Zhiming, et al.
Veröffentlicht: (2025)
MemoCoder: Automated Function Synthesis using LLM-Supported Agents
von: Jia, Yiping, et al.
Veröffentlicht: (2025)
von: Jia, Yiping, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
From Context to Intent: Reasoning-Guided Function-Level Code Completion
von: Li, Yanzhou, et al.
Veröffentlicht: (2025) -
Software Development Life Cycle Perspective: A Survey of Benchmarks for Code Large Language Models and Agents
von: Wang, Kaixin, et al.
Veröffentlicht: (2025) -
Latent Imitator: Generating Natural Individual Discriminatory Instances for Black-Box Fairness Testing
von: Xiao, Yisong, et al.
Veröffentlicht: (2023) -
InCoder-32B: Code Foundation Model for Industrial Scenarios
von: Yang, Jian, et al.
Veröffentlicht: (2026) -
Unveiling Project-Specific Bias in Neural Code Models
von: Li, Zhiming, et al.
Veröffentlicht: (2022)