Unveiling Code Pre-Trained Models: Investigating Syntax and Semantics Capacities
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Wei, Liu, Shangqing, Zhao, Mengjie, Xie, Xiaofei, Wang, Wenhan, Hu, Qiang, Zhang, Jie, Liu, Yang |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Do Code Semantics Help? A Comprehensive Study on Execution Trace-Based Information for Code Large Language Models
by: Wang, Jian, et al.
Published: (2025)
by: Wang, Jian, et al.
Published: (2025)
Intention is All You Need: Refining Your Code from Your Intention
by: Guo, Qi, et al.
Published: (2025)
by: Guo, Qi, et al.
Published: (2025)
SpecGen: Automated Generation of Formal Program Specifications via Large Language Models
by: Ma, Lezhi, et al.
Published: (2024)
by: Ma, Lezhi, et al.
Published: (2024)
Defects4C: Benchmarking Large Language Model Repair Capability with C/C++ Bugs
by: Wang, Jian, et al.
Published: (2025)
by: Wang, Jian, et al.
Published: (2025)
Improving Code Translation with Syntax-Guided and Semantic-aware Preference Optimization
by: Wu, Yuhan, et al.
Published: (2026)
by: Wu, Yuhan, et al.
Published: (2026)
SpecEval: Evaluating Code Comprehension in Large Language Models via Program Specifications
by: Ma, Lezhi, et al.
Published: (2024)
by: Ma, Lezhi, et al.
Published: (2024)
ContrastRepair: Enhancing Conversation-Based Automated Program Repair via Contrastive Test Case Pairs
by: Kong, Jiaolong, et al.
Published: (2024)
by: Kong, Jiaolong, et al.
Published: (2024)
FT2Ra: A Fine-Tuning-Inspired Approach to Retrieval-Augmented Code Completion
by: Guo, Qi, et al.
Published: (2024)
by: Guo, Qi, et al.
Published: (2024)
Unveiling Memorization in Code Models
by: Yang, Zhou, et al.
Published: (2023)
by: Yang, Zhou, et al.
Published: (2023)
Impact-driven Context Filtering For Cross-file Code Completion
by: Li, Yanzhou, et al.
Published: (2025)
by: Li, Yanzhou, et al.
Published: (2025)
Bias Unveiled: Investigating Social Bias in LLM-Generated Code
by: Ling, Lin, et al.
Published: (2024)
by: Ling, Lin, et al.
Published: (2024)
Same Signal, Different Semantics: A Cross-Framework Behavioral Analysis of Software Engineering Agents
by: Ma, Wei, et al.
Published: (2026)
by: Ma, Wei, et al.
Published: (2026)
CodeFort: Robust Training for Code Generation Models
by: Zhang, Yuhao, et al.
Published: (2024)
by: Zhang, Yuhao, et al.
Published: (2024)
GenCode: A Generic Data Augmentation Framework for Boosting Deep Learning-Based Code Understanding
by: Dong, Zeming, et al.
Published: (2024)
by: Dong, Zeming, et al.
Published: (2024)
Syntax Is Not Enough: An Empirical Study of Small Transformer Models for Neural Code Repair
by: Samant, Shaunak
Published: (2025)
by: Samant, Shaunak
Published: (2025)
Prompt Stability in Code LLMs: Measuring Sensitivity across Emotion- and Personality-Driven Variations
by: Ma, Wei, et al.
Published: (2025)
by: Ma, Wei, et al.
Published: (2025)
Large Language Model Supply Chain: Open Problems From the Security Perspective
by: Hu, Qiang, et al.
Published: (2024)
by: Hu, Qiang, et al.
Published: (2024)
From Context to Intent: Reasoning-Guided Function-Level Code Completion
by: Li, Yanzhou, et al.
Published: (2025)
by: Li, Yanzhou, et al.
Published: (2025)
A Comprehensive Study of Governance Issues in Decentralized Finance Applications
by: Ma, Wei, et al.
Published: (2023)
by: Ma, Wei, et al.
Published: (2023)
Coding-PTMs: How to Find Optimal Code Pre-trained Models for Code Embedding in Vulnerability Detection?
by: Zhao, Yu, et al.
Published: (2024)
by: Zhao, Yu, et al.
Published: (2024)
Evaluating Pre-Trained Models for Multi-Language Vulnerability Patching
by: Khan, Zanis Ali, et al.
Published: (2025)
by: Khan, Zanis Ali, et al.
Published: (2025)
Unveiling Code Clones in the Eclipse IIoT Software Ecosystem
by: Li, Zengyang, et al.
Published: (2026)
by: Li, Zengyang, et al.
Published: (2026)
Code Change Intention, Development Artifact and History Vulnerability: Putting Them Together for Vulnerability Fix Detection by LLM
by: Yang, Xu, et al.
Published: (2025)
by: Yang, Xu, et al.
Published: (2025)
Combining Fine-Tuning and LLM-based Agents for Intuitive Smart Contract Auditing with Justifications
by: Ma, Wei, et al.
Published: (2024)
by: Ma, Wei, et al.
Published: (2024)
PseudoBridge: Pseudo Code as the Bridge for Better Semantic and Logic Alignment in Code Retrieval
by: Li, Yixuan, et al.
Published: (2025)
by: Li, Yixuan, et al.
Published: (2025)
SynConfRoute: Syntax-Aware Routing for Efficient Code Completion with Small CodeLLMs
by: Thangarajah, Kishanthan, et al.
Published: (2026)
by: Thangarajah, Kishanthan, et al.
Published: (2026)
CodeImprove: Program Adaptation for Deep Code Models
by: Rathnasuriya, Ravishka, et al.
Published: (2025)
by: Rathnasuriya, Ravishka, et al.
Published: (2025)
AOCI: Symbolic-Semantic Indexing for Practical Repository-Scale Code Understanding with LLMs
by: Liu, Jinshi, et al.
Published: (2026)
by: Liu, Jinshi, et al.
Published: (2026)
Utilizing Source Code Syntax Patterns to Detect Bug Inducing Commits using Machine Learning Models
by: Nadim, Md, et al.
Published: (2022)
by: Nadim, Md, et al.
Published: (2022)
An Empirical Study on Noisy Label Learning for Program Understanding
by: Wang, Wenhan, et al.
Published: (2023)
by: Wang, Wenhan, et al.
Published: (2023)
Chimera: Harnessing Multi-Agent LLMs for Automatic Insider Threat Simulation
by: Yu, Jiongchi, et al.
Published: (2025)
by: Yu, Jiongchi, et al.
Published: (2025)
ShortCoder: Knowledge-Augmented Syntax Optimization for Token-Efficient Code Generation
by: Liu, Sicong, et al.
Published: (2026)
by: Liu, Sicong, et al.
Published: (2026)
Towards Summarizing Code Snippets Using Pre-Trained Transformers
by: Mastropaolo, Antonio, et al.
Published: (2024)
by: Mastropaolo, Antonio, et al.
Published: (2024)
CodeCoT: Tackling Code Syntax Errors in CoT Reasoning for Code Generation
by: Huang, Dong, et al.
Published: (2023)
by: Huang, Dong, et al.
Published: (2023)
CKGFuzzer: LLM-Based Fuzz Driver Generation Enhanced By Code Knowledge Graph
by: Xu, Hanxiang, et al.
Published: (2024)
by: Xu, Hanxiang, et al.
Published: (2024)
Evaluation and Improvement of Fault Detection for Large Language Models
by: Hu, Qiang, et al.
Published: (2024)
by: Hu, Qiang, et al.
Published: (2024)
How to Select Pre-Trained Code Models for Reuse? A Learning Perspective
by: Bi, Zhangqian, et al.
Published: (2025)
by: Bi, Zhangqian, et al.
Published: (2025)
CoderEval: A Benchmark of Pragmatic Code Generation with Generative Pre-trained Models
by: Yu, Hao, et al.
Published: (2023)
by: Yu, Hao, et al.
Published: (2023)
Unveiling Project-Specific Bias in Neural Code Models
by: Li, Zhiming, et al.
Published: (2022)
by: Li, Zhiming, et al.
Published: (2022)
Condor: A Code Discriminator Integrating General Semantics with Code Details
by: Liang, Qingyuan, et al.
Published: (2024)
by: Liang, Qingyuan, et al.
Published: (2024)
Similar Items
-
Do Code Semantics Help? A Comprehensive Study on Execution Trace-Based Information for Code Large Language Models
by: Wang, Jian, et al.
Published: (2025) -
Intention is All You Need: Refining Your Code from Your Intention
by: Guo, Qi, et al.
Published: (2025) -
SpecGen: Automated Generation of Formal Program Specifications via Large Language Models
by: Ma, Lezhi, et al.
Published: (2024) -
Defects4C: Benchmarking Large Language Model Repair Capability with C/C++ Bugs
by: Wang, Jian, et al.
Published: (2025) -
Improving Code Translation with Syntax-Guided and Semantic-aware Preference Optimization
by: Wu, Yuhan, et al.
Published: (2026)