Boosting Source Code Learning with Text-Oriented Data Augmentation: An Empirical Study
Fuente:
arXiv
Guardado en:
| Autores principales: | Dong, Zeming, Hu, Qiang, Guo, Yuejun, Zhang, Zhenya, Cordy, Maxime, Papadakis, Mike, Traon, Yves Le, Zhao, Jianjun |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
GenCode: A Generic Data Augmentation Framework for Boosting Deep Learning-Based Code Understanding
por: Dong, Zeming, et al.
Publicado: (2024)
por: Dong, Zeming, et al.
Publicado: (2024)
Defective Task Descriptions in LLM-Based Code Generation: Detection and Analysis
por: Akli, Amal, et al.
Publicado: (2026)
por: Akli, Amal, et al.
Publicado: (2026)
Software Fairness: An Analysis and Survey
por: Soremekun, Ezekiel, et al.
Publicado: (2022)
por: Soremekun, Ezekiel, et al.
Publicado: (2022)
One Model, Many Skills: Parameter-Efficient Fine-Tuning for Multitask Code Analysis
por: Akli, Amal, et al.
Publicado: (2026)
por: Akli, Amal, et al.
Publicado: (2026)
An Empirical Study of the Imbalance Issue in Software Vulnerability Detection
por: Guo, Yuejun, et al.
Publicado: (2026)
por: Guo, Yuejun, et al.
Publicado: (2026)
Learning Generalizable Multimodal Representations for Software Vulnerability Detection
por: Dong, Zeming, et al.
Publicado: (2026)
por: Dong, Zeming, et al.
Publicado: (2026)
When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions
por: Larbi, Maya, et al.
Publicado: (2025)
por: Larbi, Maya, et al.
Publicado: (2025)
On the Effectiveness of Hybrid Pooling in Mixup-Based Graph Learning for Language Processing
por: Dong, Zeming, et al.
Publicado: (2022)
por: Dong, Zeming, et al.
Publicado: (2022)
Foundation Models for Autonomous Driving System: An Initial Roadmap
por: Wu, Xiongfei, et al.
Publicado: (2025)
por: Wu, Xiongfei, et al.
Publicado: (2025)
When Prompt Under-Specification Improves Code Correctness: An Exploratory Study of Prompt Wording and Structure Effects on LLM-Based Code Generation
por: AKLI, Amal, et al.
Publicado: (2026)
por: AKLI, Amal, et al.
Publicado: (2026)
Inferring Code Correctness from Specification
por: Florian, Tambon, et al.
Publicado: (2026)
por: Florian, Tambon, et al.
Publicado: (2026)
You Can REST Now: Automated REST API Documentation and Testing via LLM-Assisted Request Mutations
por: Decrop, Alix, et al.
Publicado: (2024)
por: Decrop, Alix, et al.
Publicado: (2024)
Evaluating Pre-Trained Models for Multi-Language Vulnerability Patching
por: Khan, Zanis Ali, et al.
Publicado: (2025)
por: Khan, Zanis Ali, et al.
Publicado: (2025)
Boosting LLMs for Mutation Generation
por: Wang, Bo, et al.
Publicado: (2026)
por: Wang, Bo, et al.
Publicado: (2026)
An Empirical Study of Retrieval-Augmented Code Generation: Challenges and Opportunities
por: Yang, Zezhou, et al.
Publicado: (2025)
por: Yang, Zezhou, et al.
Publicado: (2025)
Do LLMs generate test oracles that capture the actual or the expected program behaviour?
por: Konstantinou, Michael, et al.
Publicado: (2024)
por: Konstantinou, Michael, et al.
Publicado: (2024)
How well LLM-based test generation techniques perform with newer LLM versions?
por: Konstantinou, Michael, et al.
Publicado: (2026)
por: Konstantinou, Michael, et al.
Publicado: (2026)
Unveiling Code Clones in Quantum Programming: An Empirical Study with Qiskit
por: Manoku, Kenta, et al.
Publicado: (2025)
por: Manoku, Kenta, et al.
Publicado: (2025)
On the Effectiveness of Training Data Optimization for LLM-based Code Generation: An Empirical Study
por: Kuang, Shiqi, et al.
Publicado: (2025)
por: Kuang, Shiqi, et al.
Publicado: (2025)
An Empirical Study of Policy-as-Code Adoption in Open-Source Software Projects
por: Foalem, Patrick Loic, et al.
Publicado: (2026)
por: Foalem, Patrick Loic, et al.
Publicado: (2026)
QuanBench: Benchmarking Quantum Code Generation with Large Language Models
por: Guo, Xiaoyu, et al.
Publicado: (2025)
por: Guo, Xiaoyu, et al.
Publicado: (2025)
Evaluation and Improvement of Fault Detection for Large Language Models
por: Hu, Qiang, et al.
Publicado: (2024)
por: Hu, Qiang, et al.
Publicado: (2024)
An Empirical Study on Method-Level Performance Evolution in Open-Source Java Projects
por: Shahedi, Kaveh, et al.
Publicado: (2025)
por: Shahedi, Kaveh, et al.
Publicado: (2025)
Toward Interactive Optimization of Source Code Differences: An Empirical Study of Its Performance
por: Yagi, Tsukasa, et al.
Publicado: (2024)
por: Yagi, Tsukasa, et al.
Publicado: (2024)
Unveiling Code Clone Patterns in Open Source VR Software: An Empirical Study
por: Chen, Huashan, et al.
Publicado: (2025)
por: Chen, Huashan, et al.
Publicado: (2025)
What to Retrieve for Effective Retrieval-Augmented Code Generation? An Empirical Study and Beyond
por: Gu, Wenchao, et al.
Publicado: (2025)
por: Gu, Wenchao, et al.
Publicado: (2025)
Code vs Serialized AST Inputs for LLM-Based Code Summarization: An Empirical Study
por: Dong, Shijia, et al.
Publicado: (2026)
por: Dong, Shijia, et al.
Publicado: (2026)
An Empirical Study on Bug Severity Estimation using Source Code Metrics and Static Analysis
por: Mashhadi, Ehsan, et al.
Publicado: (2022)
por: Mashhadi, Ehsan, et al.
Publicado: (2022)
An Empirical Study on Automatically Detecting AI-Generated Source Code: How Far Are We?
por: Suh, Hyunjae, et al.
Publicado: (2024)
por: Suh, Hyunjae, et al.
Publicado: (2024)
You Augment Me: Exploring ChatGPT-based Data Augmentation for Semantic Code Search
por: Wang, Yanlin, et al.
Publicado: (2024)
por: Wang, Yanlin, et al.
Publicado: (2024)
Distribution-aware Fairness Test Generation
por: Rajan, Sai Sathiesh, et al.
Publicado: (2023)
por: Rajan, Sai Sathiesh, et al.
Publicado: (2023)
YATE: The Role of Test Repair in LLM-Based Unit Test Generation
por: Konstantinou, Michael, et al.
Publicado: (2025)
por: Konstantinou, Michael, et al.
Publicado: (2025)
How Does Chunking Affect Retrieval-Augmented Code Completion? A Controlled Empirical Study
por: Wu, Xinjian, et al.
Publicado: (2026)
por: Wu, Xinjian, et al.
Publicado: (2026)
Detection of Technical Debt in Java Source Code
por: Hai, Nam Le, et al.
Publicado: (2024)
por: Hai, Nam Le, et al.
Publicado: (2024)
CodeDPO: Aligning Code Models with Self Generated and Verified Source Code
por: Zhang, Kechi, et al.
Publicado: (2024)
por: Zhang, Kechi, et al.
Publicado: (2024)
Source-Code Analysis of iFogSim for Simulating Distributed IoT Architectures: Coverage, Challenges, and Enhancements
por: Ndadji, Milliam Maxime Zekeng
Publicado: (2026)
por: Ndadji, Milliam Maxime Zekeng
Publicado: (2026)
On the Use of Agentic Coding Manifests: An Empirical Study of Claude Code
por: Chatlatanagulchai, Worawalan, et al.
Publicado: (2025)
por: Chatlatanagulchai, Worawalan, et al.
Publicado: (2025)
An Empirical Study on Influence-Based Pretraining Data Selection for Code Large Language Models
por: Xing, Chengli, et al.
Publicado: (2026)
por: Xing, Chengli, et al.
Publicado: (2026)
Large Language Model Unlearning for Source Code
por: Jiang, Xue, et al.
Publicado: (2025)
por: Jiang, Xue, et al.
Publicado: (2025)
ConAIR:Consistency-Augmented Iterative Interaction Framework to Enhance the Reliability of Code Generation
por: Dong, Jinhao, et al.
Publicado: (2024)
por: Dong, Jinhao, et al.
Publicado: (2024)
Ejemplares similares
-
GenCode: A Generic Data Augmentation Framework for Boosting Deep Learning-Based Code Understanding
por: Dong, Zeming, et al.
Publicado: (2024) -
Defective Task Descriptions in LLM-Based Code Generation: Detection and Analysis
por: Akli, Amal, et al.
Publicado: (2026) -
Software Fairness: An Analysis and Survey
por: Soremekun, Ezekiel, et al.
Publicado: (2022) -
One Model, Many Skills: Parameter-Efficient Fine-Tuning for Multitask Code Analysis
por: Akli, Amal, et al.
Publicado: (2026) -
An Empirical Study of the Imbalance Issue in Software Vulnerability Detection
por: Guo, Yuejun, et al.
Publicado: (2026)