Large Language Models are Qualified Benchmark Builders: Rebuilding Pre-Training Datasets for Advancing Code Intelligence Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Kang, Mao, Xinjun, Wang, Shangwen, Wang, Yanlin, Zhang, Tanghaoran, Lin, Bo, Qin, Yihao, Zhang, Zhang, Lu, Yao, Al-Sabahi, Kamal |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CARLDA: An Approach for Stack Overflow API Mention Recognition Driven by Context and LLM‐Based Data Augmentation
by: Zhang Zhang, et al.
Published: (2025)
by: Zhang Zhang, et al.
Published: (2025)
AdaptEval: A Benchmark for Evaluating Large Language Models on Code Snippet Adaptation
by: Zhang, Tanghaoran, et al.
Published: (2026)
by: Zhang, Tanghaoran, et al.
Published: (2026)
Instruct or Interact? Exploring and Eliciting LLMs' Capability in Code Snippet Adaptation Through Prompt Engineering
by: Zhang, Tanghaoran, et al.
Published: (2024)
by: Zhang, Tanghaoran, et al.
Published: (2024)
Coding in a Bubble? Evaluating LLMs in Resolving Context Adaptation Bugs During Code Adaptation
by: Zhang, Tanghaoran, et al.
Published: (2026)
by: Zhang, Tanghaoran, et al.
Published: (2026)
Give LLMs a Security Course: Securing Retrieval-Augmented Code Generation via Knowledge Injection
by: Lin, Bo, et al.
Published: (2025)
by: Lin, Bo, et al.
Published: (2025)
Keep It Simple: Towards Accurate Vulnerability Detection for Large Code Graphs
by: Peng, Xin, et al.
Published: (2024)
by: Peng, Xin, et al.
Published: (2024)
Smoke and Mirrors: Jailbreaking LLM-based Code Generation via Implicit Malicious Prompts
by: Ouyang, Sheng, et al.
Published: (2025)
by: Ouyang, Sheng, et al.
Published: (2025)
Fault Localization from the Semantic Code Search Perspective
by: Qin, Yihao, et al.
Published: (2024)
by: Qin, Yihao, et al.
Published: (2024)
Large Language Models-Aided Program Debloating
by: Lin, Bo, et al.
Published: (2025)
by: Lin, Bo, et al.
Published: (2025)
A Decentralized Framework for Ethical Authorship Validation in Academic Publishing: Leveraging Self-Sovereign Identity and Blockchain Technology
by: Al-Sabahi, Kamal, et al.
Published: (2025)
by: Al-Sabahi, Kamal, et al.
Published: (2025)
BDiff: Block-aware and Accurate Text-based Code Differencing
by: Lu, Yao, et al.
Published: (2025)
by: Lu, Yao, et al.
Published: (2025)
Exploring the Security Threats of Knowledge Base Poisoning in Retrieval-Augmented Code Generation
by: Lin, Bo, et al.
Published: (2025)
by: Lin, Bo, et al.
Published: (2025)
Zigzag Codes Revisited: From Optimal Rebuilding to Small Skip Cost and Small Fields
by: Zhang, Wenqin, et al.
Published: (2025)
by: Zhang, Wenqin, et al.
Published: (2025)
AgentFL: Scaling LLM-based Fault Localization to Project-Level Context
by: Qin, Yihao, et al.
Published: (2024)
by: Qin, Yihao, et al.
Published: (2024)
Multi-head Sequence Tagging Model for Grammatical Error Correction
by: Al-Sabahi, Kamal, et al.
Published: (2024)
by: Al-Sabahi, Kamal, et al.
Published: (2024)
AutoBaxBuilder: Bootstrapping Code Security Benchmarking
by: von Arx, Tobias, et al.
Published: (2025)
by: von Arx, Tobias, et al.
Published: (2025)
Exploring the Security Threats of Retriever Backdoors in Retrieval-Augmented Code Generation
by: Li, Tian, et al.
Published: (2025)
by: Li, Tian, et al.
Published: (2025)
There are More Fish in the Sea: Automated Vulnerability Repair via Binary Templates
by: Lin, Bo, et al.
Published: (2024)
by: Lin, Bo, et al.
Published: (2024)
MEV in Binance Builder
by: Wang, Qin, et al.
Published: (2026)
by: Wang, Qin, et al.
Published: (2026)
A Discussion to Qualify Intelligence
by: Greer, Kieran
Published: (2014)
by: Greer, Kieran
Published: (2014)
A LLM Benchmark based on the Minecraft Builder Dialog Agent Task
by: Madge, Chris, et al.
Published: (2024)
by: Madge, Chris, et al.
Published: (2024)
Enhancing Mortality Prediction in Heart Failure Patients: Exploring Preprocessing Methods for Imbalanced Clinical Datasets
by: Kia, Hanif, et al.
Published: (2023)
by: Kia, Hanif, et al.
Published: (2023)
ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning
by: Zhang, Yiming, et al.
Published: (2026)
by: Zhang, Yiming, et al.
Published: (2026)
Mobile-Env: Building Qualified Evaluation Benchmarks for LLM-GUI Interaction
by: Zhang, Danyang, et al.
Published: (2023)
by: Zhang, Danyang, et al.
Published: (2023)
An extendable voltage boosting gain unit with a self‐balanced switched‐capacitors multilevel inverter using a single DC source
by: Weam Al‐Tameemi, et al.
Published: (2024)
by: Weam Al‐Tameemi, et al.
Published: (2024)
Towards an Understanding of Context Utilization in Code Intelligence
by: Wang, Yanlin, et al.
Published: (2025)
by: Wang, Yanlin, et al.
Published: (2025)
Decentralization of Ethereum's Builder Market
by: Yang, Sen, et al.
Published: (2024)
by: Yang, Sen, et al.
Published: (2024)
HumanEvo: An Evolution-aware Benchmark for More Realistic Evaluation of Repository-level Code Generation
by: Zheng, Dewu, et al.
Published: (2024)
by: Zheng, Dewu, et al.
Published: (2024)
Nahm sum identities for Cartan matrices of type $D_k$
by: Wang, Liuquan, et al.
Published: (2025)
by: Wang, Liuquan, et al.
Published: (2025)
CoderEval: A Benchmark of Pragmatic Code Generation with Generative Pre-trained Models
by: Yu, Hao, et al.
Published: (2023)
by: Yu, Hao, et al.
Published: (2023)
On the Variability of Source Code in Maven Package Rebuilds
by: Dietrich, Jens, et al.
Published: (2026)
by: Dietrich, Jens, et al.
Published: (2026)
DockSmith: Scaling Reliable Coding Environments via an Agentic Docker Builder
by: Zhang, Jiaran, et al.
Published: (2026)
by: Zhang, Jiaran, et al.
Published: (2026)
RepoTransBench: A Real-World Multilingual Benchmark for Repository-Level Code Translation
by: Wang, Yanli, et al.
Published: (2024)
by: Wang, Yanli, et al.
Published: (2024)
BuilderBench: The Building Blocks of Intelligent Agents
by: Ghugare, Raj, et al.
Published: (2025)
by: Ghugare, Raj, et al.
Published: (2025)
Top General Performance = Top Domain Performance? DomainCodeBench: A Multi-domain Code Generation Benchmark
by: Zheng, Dewu, et al.
Published: (2024)
by: Zheng, Dewu, et al.
Published: (2024)
FeedbackEval: A Benchmark for Evaluating Large Language Models in Feedback-Driven Code Repair Tasks
by: Dai, Dekun, et al.
Published: (2025)
by: Dai, Dekun, et al.
Published: (2025)
Flashback: Enhancing Proposer-Builder Design with Future-Block Auctions in Proof-of-Stake Ethereum
by: Mao, Yifan, et al.
Published: (2024)
by: Mao, Yifan, et al.
Published: (2024)
TaskWeaver: A Code-First Agent Framework
by: Qiao, Bo, et al.
Published: (2023)
by: Qiao, Bo, et al.
Published: (2023)
RealSec-bench: A Benchmark for Evaluating Secure Code Generation in Real-World Repositories
by: Wang, Yanlin, et al.
Published: (2026)
by: Wang, Yanlin, et al.
Published: (2026)
A Survey of Large Language Models for Code: Evolution, Benchmarking, and Future Trends
by: Zheng, Zibin, et al.
Published: (2023)
by: Zheng, Zibin, et al.
Published: (2023)
Similar Items
-
CARLDA: An Approach for Stack Overflow API Mention Recognition Driven by Context and LLM‐Based Data Augmentation
by: Zhang Zhang, et al.
Published: (2025) -
AdaptEval: A Benchmark for Evaluating Large Language Models on Code Snippet Adaptation
by: Zhang, Tanghaoran, et al.
Published: (2026) -
Instruct or Interact? Exploring and Eliciting LLMs' Capability in Code Snippet Adaptation Through Prompt Engineering
by: Zhang, Tanghaoran, et al.
Published: (2024) -
Coding in a Bubble? Evaluating LLMs in Resolving Context Adaptation Bugs During Code Adaptation
by: Zhang, Tanghaoran, et al.
Published: (2026) -
Give LLMs a Security Course: Securing Retrieval-Augmented Code Generation via Knowledge Injection
by: Lin, Bo, et al.
Published: (2025)