Less is More: DocString Compression in Code Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Guang, Zhou, Yu, Cheng, Wei, Zhang, Xiangyu, Chen, Xiang, Zhuo, Terry Yue, Liu, Ke, Zhou, Xin, Lo, David, Chen, Taolue |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Less is More: Towards Green Code Large Language Models via Unified Structural Pruning
by: Yang, Guang, et al.
Published: (2024)
by: Yang, Guang, et al.
Published: (2024)
Chain-of-Thought in Neural Code Generation: From and For Lightweight Language Models
by: Yang, Guang, et al.
Published: (2023)
by: Yang, Guang, et al.
Published: (2023)
Defending Code Language Models against Backdoor Attacks with Deceptive Cross-Entropy Loss
by: Yang, Guang, et al.
Published: (2024)
by: Yang, Guang, et al.
Published: (2024)
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation
by: Yang, Guang, et al.
Published: (2025)
by: Yang, Guang, et al.
Published: (2025)
Robustness, Security, Privacy, Explainability, Efficiency, and Usability of Large Language Models for Code
by: Yang, Zhou, et al.
Published: (2024)
by: Yang, Zhou, et al.
Published: (2024)
CodeScore-R: An Automated Robustness Metric for Assessing the FunctionalCorrectness of Code Synthesis
by: Yang, Guang, et al.
Published: (2024)
by: Yang, Guang, et al.
Published: (2024)
Anchor Attention, Small Cache: Code Generation with Large Language Models
by: Zhang, Xiangyu, et al.
Published: (2024)
by: Zhang, Xiangyu, et al.
Published: (2024)
The Cream Rises to the Top: Efficient Reranking Method for Verilog Code Generation
by: Yang, Guang, et al.
Published: (2025)
by: Yang, Guang, et al.
Published: (2025)
Less is More: On the Importance of Data Quality for Unit Test Generation
by: Zhang, Junwei, et al.
Published: (2025)
by: Zhang, Junwei, et al.
Published: (2025)
SimADFuzz: Simulation-Feedback Fuzz Testing for Autonomous Driving Systems
by: Yang, Huiwen, et al.
Published: (2024)
by: Yang, Huiwen, et al.
Published: (2024)
Hotfixing Large Language Models for Code
by: Yang, Zhou, et al.
Published: (2024)
by: Yang, Zhou, et al.
Published: (2024)
Task Abstention for Large Language Models in Code Generation
by: Zhou, Yanke, et al.
Published: (2026)
by: Zhou, Yanke, et al.
Published: (2026)
ICE-Score: Instructing Large Language Models to Evaluate Code
by: Zhuo, Terry Yue
Published: (2023)
by: Zhuo, Terry Yue
Published: (2023)
From Code to Courtroom: LLMs as the New Software Judges
by: He, Junda, et al.
Published: (2025)
by: He, Junda, et al.
Published: (2025)
VisDocSketcher: Towards Scalable Visual Documentation with Agentic Systems
by: Gomes, Luís F., et al.
Published: (2025)
by: Gomes, Luís F., et al.
Published: (2025)
Semantic Consensus Decoding: Backdoor Defense for Verilog Code Generation
by: Yang, Guang, et al.
Published: (2026)
by: Yang, Guang, et al.
Published: (2026)
From Mirage to Grounding: Towards Reliable Multimodal Circuit-to-Verilog Code Generation
by: Yang, Guang, et al.
Published: (2026)
by: Yang, Guang, et al.
Published: (2026)
Uncertainty Quantification for LLM-based Code Generation
by: Xu, Senrong, et al.
Published: (2026)
by: Xu, Senrong, et al.
Published: (2026)
Beyond C/C++: Probabilistic and LLM Methods for Next-Generation Software Reverse Engineering
by: Zhuo, Zhuo, et al.
Published: (2025)
by: Zhuo, Zhuo, et al.
Published: (2025)
CodeArena: A Collective Evaluation Platform for LLM Code Generation
by: Du, Mingzhe, et al.
Published: (2025)
by: Du, Mingzhe, et al.
Published: (2025)
Code Less, Align More: Efficient LLM Fine-tuning for Code Generation with Data Pruning
by: Tsai, Yun-Da, et al.
Published: (2024)
by: Tsai, Yun-Da, et al.
Published: (2024)
On the Usage of Continual Learning for Out-of-Distribution Generalization in Pre-trained Language Models of Code
by: Weyssow, Martin, et al.
Published: (2023)
by: Weyssow, Martin, et al.
Published: (2023)
Ecosystem of Large Language Models for Code
by: Yang, Zhou, et al.
Published: (2024)
by: Yang, Zhou, et al.
Published: (2024)
Debt Behind the AI Boom: A Large-Scale Empirical Study of AI-Generated Code in the Wild
by: Liu, Yue, et al.
Published: (2026)
by: Liu, Yue, et al.
Published: (2026)
Large Language Model for Vulnerability Detection: Emerging Results and Future Directions
by: Zhou, Xin, et al.
Published: (2024)
by: Zhou, Xin, et al.
Published: (2024)
Fixing Large Language Models' Specification Misunderstanding for Better Code Generation
by: Tian, Zhao, et al.
Published: (2023)
by: Tian, Zhao, et al.
Published: (2023)
Automatic Bi-modal Question Title Generation for Stack Overflow with Prompt Learning
by: Yang, Shaoyu, et al.
Published: (2024)
by: Yang, Shaoyu, et al.
Published: (2024)
Identifying and Mitigating API Misuse in Large Language Models
by: Zhuo, Terry Yue, et al.
Published: (2025)
by: Zhuo, Terry Yue, et al.
Published: (2025)
A Comprehensive Evaluation of Parameter-Efficient Fine-Tuning on Code Smell Detection
by: Zhang, Beiqi, et al.
Published: (2024)
by: Zhang, Beiqi, et al.
Published: (2024)
CODEPROMPTZIP: Code-specific Prompt Compression for Retrieval-Augmented Generation in Coding Tasks with LMs
by: He, Pengfei, et al.
Published: (2025)
by: He, Pengfei, et al.
Published: (2025)
Towards Efficient Verification of Constant-Time Cryptographic Implementations
by: Cai, Luwei, et al.
Published: (2024)
by: Cai, Luwei, et al.
Published: (2024)
Model-less Is the Best Model: Generating Pure Code Implementations to Replace On-Device DL Models
by: Zhou, Mingyi, et al.
Published: (2024)
by: Zhou, Mingyi, et al.
Published: (2024)
LLM-as-a-Judge for Software Engineering: Literature Review, Vision, and the Road Ahead
by: He, Junda, et al.
Published: (2025)
by: He, Junda, et al.
Published: (2025)
DocChecker: Bootstrapping Code Large Language Model for Detecting and Resolving Code-Comment Inconsistencies
by: Dau, Anh T. V., et al.
Published: (2023)
by: Dau, Anh T. V., et al.
Published: (2023)
GitChameleon: Unmasking the Version-Switching Capabilities of Code Generation Models
by: Islah, Nizar, et al.
Published: (2024)
by: Islah, Nizar, et al.
Published: (2024)
Bamboo: LLM-Driven Discovery of API-Permission Mappings in the Android Framework
by: Hu, Han, et al.
Published: (2025)
by: Hu, Han, et al.
Published: (2025)
Exploring Parameter-Efficient Fine-Tuning Techniques for Code Generation with Large Language Models
by: Weyssow, Martin, et al.
Published: (2023)
by: Weyssow, Martin, et al.
Published: (2023)
Synthesizing Inductive Invariants for Distributed Protocols via IC3 and Large Language Models
by: Cao, Weining, et al.
Published: (2026)
by: Cao, Weining, et al.
Published: (2026)
Large Language Model for Vulnerability Detection and Repair: Literature Review and the Road Ahead
by: Zhou, Xin, et al.
Published: (2024)
by: Zhou, Xin, et al.
Published: (2024)
Yet Even Less Is Even Better For Agentic, Reasoning, and Coding LLMs
by: CodeArts Model Team, et al.
Published: (2026)
by: CodeArts Model Team, et al.
Published: (2026)
Similar Items
-
Less is More: Towards Green Code Large Language Models via Unified Structural Pruning
by: Yang, Guang, et al.
Published: (2024) -
Chain-of-Thought in Neural Code Generation: From and For Lightweight Language Models
by: Yang, Guang, et al.
Published: (2023) -
Defending Code Language Models against Backdoor Attacks with Deceptive Cross-Entropy Loss
by: Yang, Guang, et al.
Published: (2024) -
CODE-DITING: A Reasoning-Based Metric for Functional Alignment in Code Evaluation
by: Yang, Guang, et al.
Published: (2025) -
Robustness, Security, Privacy, Explainability, Efficiency, and Usability of Large Language Models for Code
by: Yang, Zhou, et al.
Published: (2024)