You Augment Me: Exploring ChatGPT-based Data Augmentation for Semantic Code Search
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yanlin, Guo, Lianghong, Shi, Ensheng, Chen, Wenqing, Chen, Jiachi, Zhong, Wanjun, Wang, Menghan, Li, Hui, Zhang, Hongyu, Lyu, Ziyu, Zheng, Zibin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When to Stop? Towards Efficient Code Generation in LLMs with Excess Token Prevention
by: Guo, Lianghong, et al.
Published: (2024)
by: Guo, Lianghong, et al.
Published: (2024)
Towards an Understanding of Large Language Models in Software Engineering Tasks
by: Zheng, Zibin, et al.
Published: (2023)
by: Zheng, Zibin, et al.
Published: (2023)
ShortCoder: Knowledge-Augmented Syntax Optimization for Token-Efficient Code Generation
by: Liu, Sicong, et al.
Published: (2026)
by: Liu, Sicong, et al.
Published: (2026)
DRAINCODE: Stealthy Energy Consumption Attacks on Retrieval-Augmented Code Generation via Context Poisoning
by: Wang, Yanlin, et al.
Published: (2026)
by: Wang, Yanlin, et al.
Published: (2026)
When ChatGPT Meets Smart Contract Vulnerability Detection: How Far Are We?
by: Chen, Chong, et al.
Published: (2023)
by: Chen, Chong, et al.
Published: (2023)
Agents in Software Engineering: Survey, Landscape, and Vision
by: Wang, Yanlin, et al.
Published: (2024)
by: Wang, Yanlin, et al.
Published: (2024)
SimpleDevQA: Benchmarking Large Language Models on Development Knowledge QA
by: Zhang, Jing, et al.
Published: (2025)
by: Zhang, Jing, et al.
Published: (2025)
AlignCoder: Aligning Retrieval with Target Intent for Repository-Level Code Completion
by: Jiang, Tianyue, et al.
Published: (2026)
by: Jiang, Tianyue, et al.
Published: (2026)
EffiReasonTrans: RL-Optimized Reasoning for Code Translation
by: Wang, Yanlin, et al.
Published: (2025)
by: Wang, Yanlin, et al.
Published: (2025)
LLM Hallucinations in Practical Code Generation: Phenomena, Mechanism, and Mitigation
by: Zhang, Ziyao, et al.
Published: (2024)
by: Zhang, Ziyao, et al.
Published: (2024)
Towards an Understanding of Context Utilization in Code Intelligence
by: Wang, Yanlin, et al.
Published: (2025)
by: Wang, Yanlin, et al.
Published: (2025)
Top General Performance = Top Domain Performance? DomainCodeBench: A Multi-domain Code Generation Benchmark
by: Zheng, Dewu, et al.
Published: (2024)
by: Zheng, Dewu, et al.
Published: (2024)
CoSQA+: Pioneering the Multi-Choice Code Search Benchmark with Test-Driven Agents
by: Gong, Jing, et al.
Published: (2024)
by: Gong, Jing, et al.
Published: (2024)
HumanEvo: An Evolution-aware Benchmark for More Realistic Evaluation of Repository-level Code Generation
by: Zheng, Dewu, et al.
Published: (2024)
by: Zheng, Dewu, et al.
Published: (2024)
Exploring ChatGPT-based Augmentation Strategies for Contrastive Aspect-based Sentiment Analysis
by: Xu, Lingling, et al.
Published: (2024)
by: Xu, Lingling, et al.
Published: (2024)
OmniGIRL: A Multilingual and Multimodal Benchmark for GitHub Issue Resolution
by: Guo, Lianghong, et al.
Published: (2025)
by: Guo, Lianghong, et al.
Published: (2025)
Enhancing SLM via ChatGPT and Dataset Augmentation
by: Pieper, Tom, et al.
Published: (2024)
by: Pieper, Tom, et al.
Published: (2024)
Identifying Smart Contract Security Issues in Code Snippets from Stack Overflow
by: Chen, Jiachi, et al.
Published: (2024)
by: Chen, Jiachi, et al.
Published: (2024)
RLCoder: Reinforcement Learning for Repository-Level Code Completion
by: Wang, Yanlin, et al.
Published: (2024)
by: Wang, Yanlin, et al.
Published: (2024)
Exploring ChatGPT's Capabilities on Vulnerability Management
by: Liu, Peiyu, et al.
Published: (2023)
by: Liu, Peiyu, et al.
Published: (2023)
Does Using ChatGPT Result in Human Cognitive Augmentation?
by: Fulbright, Ron, et al.
Published: (2024)
by: Fulbright, Ron, et al.
Published: (2024)
SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks
by: Guo, Lianghong, et al.
Published: (2025)
by: Guo, Lianghong, et al.
Published: (2025)
Primacy Effect of ChatGPT
by: Wang, Yiwei, et al.
Published: (2023)
by: Wang, Yiwei, et al.
Published: (2023)
RMCBench: Benchmarking Large Language Models' Resistance to Malicious Code
by: Chen, Jiachi, et al.
Published: (2024)
by: Chen, Jiachi, et al.
Published: (2024)
Can ChatGPT Read Who You Are?
by: Derner, Erik, et al.
Published: (2023)
by: Derner, Erik, et al.
Published: (2023)
A Survey of Large Language Models for Code: Evolution, Benchmarking, and Future Trends
by: Zheng, Zibin, et al.
Published: (2023)
by: Zheng, Zibin, et al.
Published: (2023)
An Empirical Study of ChatGPT-Related Projects and Their Issues on GitHub
by: Lin, Zheng, et al.
Published: (2024)
by: Lin, Zheng, et al.
Published: (2024)
Exploring ChatGPT's Empathic Abilities
by: Schaaff, Kristina, et al.
Published: (2023)
by: Schaaff, Kristina, et al.
Published: (2023)
RealSec-bench: A Benchmark for Evaluating Secure Code Generation in Real-World Repositories
by: Wang, Yanlin, et al.
Published: (2026)
by: Wang, Yanlin, et al.
Published: (2026)
How Prevalent is Gender Bias in ChatGPT? -- Exploring German and English ChatGPT Responses
by: Urchs, Stefanie, et al.
Published: (2023)
by: Urchs, Stefanie, et al.
Published: (2023)
Learning-by-teaching with ChatGPT: The effect of teachable ChatGPT agent on programming education
by: Chen, Angxuan, et al.
Published: (2024)
by: Chen, Angxuan, et al.
Published: (2024)
Learning by teaching with ChatGPT : The effect of teachable ChatGPT agent on programming education
by: Angxuan Chen, et al.
Published: (2025)
by: Angxuan Chen, et al.
Published: (2025)
Efficiently Detecting Reentrancy Vulnerabilities in Complex Smart Contracts
by: Wang, Zexu, et al.
Published: (2024)
by: Wang, Zexu, et al.
Published: (2024)
Advances and Frontiers of LLM-based Issue Resolution in Software Engineering: A Comprehensive Survey
by: Li, Caihua, et al.
Published: (2026)
by: Li, Caihua, et al.
Published: (2026)
ChatGPT Based Data Augmentation for Improved Parameter-Efficient Debiasing of LLMs
by: Han, Pengrui, et al.
Published: (2024)
by: Han, Pengrui, et al.
Published: (2024)
Wikipedia Contributions in the Wake of ChatGPT
by: Lyu, Liang, et al.
Published: (2025)
by: Lyu, Liang, et al.
Published: (2025)
How Secure is Code Generated by ChatGPT?
by: Khoury, Raphaël, et al.
Published: (2023)
by: Khoury, Raphaël, et al.
Published: (2023)
An Empirical Study of the Non-determinism of ChatGPT in Code Generation
by: Ouyang, Shuyin, et al.
Published: (2023)
by: Ouyang, Shuyin, et al.
Published: (2023)
RepoTransBench: A Real-World Multilingual Benchmark for Repository-Level Code Translation
by: Wang, Yanli, et al.
Published: (2024)
by: Wang, Yanli, et al.
Published: (2024)
"I Use ChatGPT to Humanize My Words": Affordances and Risks of ChatGPT to Autistic Users
by: Ma, Renkai, et al.
Published: (2026)
by: Ma, Renkai, et al.
Published: (2026)
Similar Items
-
When to Stop? Towards Efficient Code Generation in LLMs with Excess Token Prevention
by: Guo, Lianghong, et al.
Published: (2024) -
Towards an Understanding of Large Language Models in Software Engineering Tasks
by: Zheng, Zibin, et al.
Published: (2023) -
ShortCoder: Knowledge-Augmented Syntax Optimization for Token-Efficient Code Generation
by: Liu, Sicong, et al.
Published: (2026) -
DRAINCODE: Stealthy Energy Consumption Attacks on Retrieval-Augmented Code Generation via Context Poisoning
by: Wang, Yanlin, et al.
Published: (2026) -
When ChatGPT Meets Smart Contract Vulnerability Detection: How Far Are We?
by: Chen, Chong, et al.
Published: (2023)