cAST: Enhancing Code Retrieval-Augmented Generation with Structural Chunking via Abstract Syntax Tree
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Yilin, Zhao, Xinran, Wang, Zora Zhiruo, Yang, Chenyang, Wei, Jiayi, Wu, Tongshuang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CodeRAG-Bench: Can Retrieval Augment Code Generation?
by: Wang, Zora Zhiruo, et al.
Published: (2024)
by: Wang, Zora Zhiruo, et al.
Published: (2024)
Code-Craft: Hierarchical Graph-Based Code Summarization for Enhanced Context Retrieval
by: Sounthiraraj, David, et al.
Published: (2025)
by: Sounthiraraj, David, et al.
Published: (2025)
Optimizing Retrieval Augmented Generation for Object Constraint Language
by: Li, Kevin Chenhao, et al.
Published: (2025)
by: Li, Kevin Chenhao, et al.
Published: (2025)
Modular Layout Synthesis (MLS): Front-end Code via Structure Normalization and Constrained Generation
by: Liu, Chong, et al.
Published: (2025)
by: Liu, Chong, et al.
Published: (2025)
CodeRAG: Finding Relevant and Necessary Knowledge for Retrieval-Augmented Repository-Level Code Completion
by: Zhang, Sheng, et al.
Published: (2025)
by: Zhang, Sheng, et al.
Published: (2025)
GNN-Coder: Boosting Semantic Code Retrieval with Combined GNNs and Transformer
by: Ye, Yufan, et al.
Published: (2025)
by: Ye, Yufan, et al.
Published: (2025)
DeepCodeSeek: Real-Time API Retrieval for Context-Aware Code Generation
by: Esakkiraja, Esakkivel, et al.
Published: (2025)
by: Esakkiraja, Esakkivel, et al.
Published: (2025)
Context-Augmented Code Generation Using Programming Knowledge Graphs
by: Saberi, Iman, et al.
Published: (2024)
by: Saberi, Iman, et al.
Published: (2024)
Deep Code Search with Naming-Agnostic Contrastive Multi-View Learning
by: Feng, Jiadong, et al.
Published: (2024)
by: Feng, Jiadong, et al.
Published: (2024)
BLAZE: Cross-Language and Cross-Project Bug Localization via Dynamic Chunking and Hard Example Learning
by: Chakraborty, Partha, et al.
Published: (2024)
by: Chakraborty, Partha, et al.
Published: (2024)
Developing Retrieval Augmented Generation (RAG) based LLM Systems from PDFs: An Experience Report
by: Khan, Ayman Asad, et al.
Published: (2024)
by: Khan, Ayman Asad, et al.
Published: (2024)
In-Context Learning as an Effective Estimator of Functional Correctness of LLM-Generated Code
by: Das, Susmita, et al.
Published: (2025)
by: Das, Susmita, et al.
Published: (2025)
AI-assisted Coding with Cody: Lessons from Context Retrieval and Evaluation for Code Recommendations
by: Hartman, Jan, et al.
Published: (2024)
by: Hartman, Jan, et al.
Published: (2024)
Benchmarking Failures in Tool-Augmented Language Models
by: Treviño, Eduardo, et al.
Published: (2025)
by: Treviño, Eduardo, et al.
Published: (2025)
Iterative Self-Training for Code Generation via Reinforced Re-Ranking
by: Sorokin, Nikita, et al.
Published: (2025)
by: Sorokin, Nikita, et al.
Published: (2025)
RANGER -- Repository-Level Agent for Graph-Enhanced Retrieval
by: Shah, Pratik, et al.
Published: (2025)
by: Shah, Pratik, et al.
Published: (2025)
Which Programming Language and Model Work Best With LLM-as-a-Judge For Code Retrieval?
by: Roberts, Lucas, et al.
Published: (2025)
by: Roberts, Lucas, et al.
Published: (2025)
Multi-View Adaptive Contrastive Learning for Information Retrieval Based Fault Localization
by: Zhou, Chunying, et al.
Published: (2024)
by: Zhou, Chunying, et al.
Published: (2024)
Source Code Clone Detection Using Unsupervised Similarity Measures
by: Martinez-Gil, Jorge
Published: (2024)
by: Martinez-Gil, Jorge
Published: (2024)
How Does Chunking Affect Retrieval-Augmented Code Completion? A Controlled Empirical Study
by: Wu, Xinjian, et al.
Published: (2026)
by: Wu, Xinjian, et al.
Published: (2026)
MGS3: A Multi-Granularity Self-Supervised Code Search Framework
by: Li, Rui, et al.
Published: (2025)
by: Li, Rui, et al.
Published: (2025)
Domain-Specific Retrieval-Augmented Generation Using Vector Stores, Knowledge Graphs, and Tensor Factorization
by: Barron, Ryan C., et al.
Published: (2024)
by: Barron, Ryan C., et al.
Published: (2024)
SBAN: A Framework & Multi-Dimensional Dataset for Large Language Model Pre-Training and Software Code Mining
by: Jelodar, Hamed, et al.
Published: (2025)
by: Jelodar, Hamed, et al.
Published: (2025)
Fairness Testing in Retrieval-Augmented Generation: How Small Perturbations Reveal Bias in Small Language Models
by: de Oliveira, Matheus Vinicius da Silva, et al.
Published: (2025)
by: de Oliveira, Matheus Vinicius da Silva, et al.
Published: (2025)
Efficient and Scalable Provenance Tracking for LLM-Generated Code Snippets
by: Gurioli, Andrea, et al.
Published: (2026)
by: Gurioli, Andrea, et al.
Published: (2026)
Retrieval-Augmented Generation for Service Discovery: Chunking Strategies and Benchmarking
by: Pesl, Robin D., et al.
Published: (2025)
by: Pesl, Robin D., et al.
Published: (2025)
Rewriting the Code: A Simple Method for Large Language Model Augmented Code Search
by: Li, Haochen, et al.
Published: (2024)
by: Li, Haochen, et al.
Published: (2024)
AST-PAC: AST-guided Membership Inference for Code
by: Koohestani, Roham, et al.
Published: (2026)
by: Koohestani, Roham, et al.
Published: (2026)
AST-Enhanced or AST-Overloaded? The Surprising Impact of Hybrid Graph Representations on Code Clone Detection
by: Zhang, Zixian, et al.
Published: (2025)
by: Zhang, Zixian, et al.
Published: (2025)
Leveraging Graph-RAG and Prompt Engineering to Enhance LLM-Based Automated Requirement Traceability and Compliance Checks
by: Masoudifard, Arsalan, et al.
Published: (2024)
by: Masoudifard, Arsalan, et al.
Published: (2024)
Can Code Evaluation Metrics Detect Code Plagiarism?
by: Ebrahim, Fahad, et al.
Published: (2026)
by: Ebrahim, Fahad, et al.
Published: (2026)
Code vs Serialized AST Inputs for LLM-Based Code Summarization: An Empirical Study
by: Dong, Shijia, et al.
Published: (2026)
by: Dong, Shijia, et al.
Published: (2026)
Less is More: On the Importance of Data Quality for Unit Test Generation
by: Zhang, Junwei, et al.
Published: (2025)
by: Zhang, Junwei, et al.
Published: (2025)
MRWeb: An Exploration of Generating Multi-Page Resource-Aware Web Code from UI Designs
by: Wan, Yuxuan, et al.
Published: (2024)
by: Wan, Yuxuan, et al.
Published: (2024)
CodeMEM: AST-Guided Adaptive Memory for Repository-Level Iterative Code Generation
by: Wang, Peiding, et al.
Published: (2026)
by: Wang, Peiding, et al.
Published: (2026)
CSA-Trans: Code Structure Aware Transformer for AST
by: Oh, Saeyoon, et al.
Published: (2024)
by: Oh, Saeyoon, et al.
Published: (2024)
InfCode-C++: Intent-Guided Semantic Retrieval and AST-Structured Search for C++ Issue Resolution
by: Dong, Qingao, et al.
Published: (2025)
by: Dong, Qingao, et al.
Published: (2025)
Taxonomy of the Retrieval System Framework: Pitfalls and Paradigms
by: Shah, Deep, et al.
Published: (2026)
by: Shah, Deep, et al.
Published: (2026)
Improving Retrieval-Augmented Code Comment Generation by Retrieving for Generation
by: Lu, Hanzhen, et al.
Published: (2024)
by: Lu, Hanzhen, et al.
Published: (2024)
AST-T5: Structure-Aware Pretraining for Code Generation and Understanding
by: Gong, Linyuan, et al.
Published: (2024)
by: Gong, Linyuan, et al.
Published: (2024)
Similar Items
-
CodeRAG-Bench: Can Retrieval Augment Code Generation?
by: Wang, Zora Zhiruo, et al.
Published: (2024) -
Code-Craft: Hierarchical Graph-Based Code Summarization for Enhanced Context Retrieval
by: Sounthiraraj, David, et al.
Published: (2025) -
Optimizing Retrieval Augmented Generation for Object Constraint Language
by: Li, Kevin Chenhao, et al.
Published: (2025) -
Modular Layout Synthesis (MLS): Front-end Code via Structure Normalization and Constrained Generation
by: Liu, Chong, et al.
Published: (2025) -
CodeRAG: Finding Relevant and Necessary Knowledge for Retrieval-Augmented Repository-Level Code Completion
by: Zhang, Sheng, et al.
Published: (2025)