Towards Trustworthy LLMs for Code: A Data-Centric Synergistic Auditing Framework
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Chong, Chen, Zhenpeng, Li, Tianlin, Zhao, Yilun, Liu, Yang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Teaching Code LLMs to Use Autocompletion Tools in Repository-Level Code Generation
by: Wang, Chong, et al.
Published: (2024)
by: Wang, Chong, et al.
Published: (2024)
A Study on Thinking Patterns of Large Reasoning Models in Code Generation
by: Halim, Kevin, et al.
Published: (2025)
by: Halim, Kevin, et al.
Published: (2025)
TREAT: A Code LLMs Trustworthiness / Reliability Evaluation and Testing Framework
by: Gao, Shuzheng, et al.
Published: (2025)
by: Gao, Shuzheng, et al.
Published: (2025)
A Dual-Loop Agent Framework for Automated Vulnerability Reproduction
by: Liu, Bin, et al.
Published: (2026)
by: Liu, Bin, et al.
Published: (2026)
Knowledge-Guided Multi-Agent Framework for Application-Level Software Code Generation
by: Xiong, Qian, et al.
Published: (2025)
by: Xiong, Qian, et al.
Published: (2025)
SGCR: A Specification-Grounded Framework for Trustworthy LLM Code Review
by: Wang, Kai, et al.
Published: (2025)
by: Wang, Kai, et al.
Published: (2025)
LLMs Are Not a Silver Bullet: A Case Study on Software Fairness
by: Li, Xinyue, et al.
Published: (2026)
by: Li, Xinyue, et al.
Published: (2026)
Promptware Engineering: Software Engineering for Prompt-Enabled Systems
by: Chen, Zhenpeng, et al.
Published: (2025)
by: Chen, Zhenpeng, et al.
Published: (2025)
RepoMark: A Data-Usage Auditing Framework for Code Large Language Models
by: Qu, Wenjie, et al.
Published: (2025)
by: Qu, Wenjie, et al.
Published: (2025)
Code Digital Twin: Empowering LLMs with Tacit Knowledge for Complex Software Development
by: Peng, Xin, et al.
Published: (2025)
by: Peng, Xin, et al.
Published: (2025)
Benchmarking and Revisiting Code Generation Assessment: A Mutation-Based Approach
by: Wang, Longtian, et al.
Published: (2025)
by: Wang, Longtian, et al.
Published: (2025)
Software Development Life Cycle Perspective: A Survey of Benchmarks for Code Large Language Models and Agents
by: Wang, Kaixin, et al.
Published: (2025)
by: Wang, Kaixin, et al.
Published: (2025)
Towards More Trustworthy Deep Code Models by Enabling Out-of-Distribution Detection
by: Yan, Yanfu, et al.
Published: (2025)
by: Yan, Yanfu, et al.
Published: (2025)
Towards a Framework for Operationalizing the Specification of Trustworthy AI Requirements
by: Villamizar, Hugo, et al.
Published: (2025)
by: Villamizar, Hugo, et al.
Published: (2025)
CodeChemist: Functional Knowledge Transfer for Low-Resource Code Generation via Test-Time Scaling
by: Wang, Kaixin, et al.
Published: (2025)
by: Wang, Kaixin, et al.
Published: (2025)
Engineering Trustworthy Software: A Mission for LLMs
by: Vieira, Marco
Published: (2024)
by: Vieira, Marco
Published: (2024)
Codev-Bench: How Do LLMs Understand Developer-Centric Code Completion?
by: Pan, Zhenyu, et al.
Published: (2024)
by: Pan, Zhenyu, et al.
Published: (2024)
Towards More Trustworthy and Interpretable LLMs for Code through Syntax-Grounded Explanations
by: Palacio, David N., et al.
Published: (2024)
by: Palacio, David N., et al.
Published: (2024)
LLMs Meet Library Evolution: Evaluating Deprecated API Usage in LLM-based Code Completion
by: Wang, Chong, et al.
Published: (2024)
by: Wang, Chong, et al.
Published: (2024)
Personality-Guided Code Generation Using Large Language Models
by: Guo, Yaoqi, et al.
Published: (2024)
by: Guo, Yaoqi, et al.
Published: (2024)
Synergistic Enhancement of Requirement-to-Code Traceability: A Framework Combining Large Language Model based Data Augmentation and an Advanced Encoder
by: Zhang, Jianzhang, et al.
Published: (2025)
by: Zhang, Jianzhang, et al.
Published: (2025)
From Context to Intent: Reasoning-Guided Function-Level Code Completion
by: Li, Yanzhou, et al.
Published: (2025)
by: Li, Yanzhou, et al.
Published: (2025)
Towards Mitigating API Hallucination in Code Generated by LLMs with Hierarchical Dependency Aware
by: Chen, Yujia, et al.
Published: (2025)
by: Chen, Yujia, et al.
Published: (2025)
Ensemble-Based Uncertainty Estimation for Code Correctness Estimation
by: Wei, Yunxiang, et al.
Published: (2026)
by: Wei, Yunxiang, et al.
Published: (2026)
Rethinking Technology Stack Selection with AI Coding Proficiency
by: Zhang, Xiaoyu, et al.
Published: (2025)
by: Zhang, Xiaoyu, et al.
Published: (2025)
Not All Tokens Matter: Data-Centric Optimization for Efficient Code Summarization
by: Afrin, Saima, et al.
Published: (2026)
by: Afrin, Saima, et al.
Published: (2026)
Agent-Based Software Artifact Evaluation
by: Wu, Zhaonan, et al.
Published: (2026)
by: Wu, Zhaonan, et al.
Published: (2026)
Can LLMs be Effective Code Contributors? A Study on Open-source Projects
by: Chong, Chun Jie, et al.
Published: (2026)
by: Chong, Chun Jie, et al.
Published: (2026)
Structured Safety Auditing for Balancing Code Correctness and Content Safety in LLM-Generated Code
by: Tan, Honghao, et al.
Published: (2026)
by: Tan, Honghao, et al.
Published: (2026)
Toward Automated and Trustworthy Scientific Analysis and Visualization with LLM-Generated Code
by: Chakroborti, Apu Kumar, et al.
Published: (2025)
by: Chakroborti, Apu Kumar, et al.
Published: (2025)
CodeFuse-Query: A Data-Centric Static Code Analysis System for Large-Scale Organizations
by: Xie, Xiaoheng, et al.
Published: (2024)
by: Xie, Xiaoheng, et al.
Published: (2024)
When to Stop? Towards Efficient Code Generation in LLMs with Excess Token Prevention
by: Guo, Lianghong, et al.
Published: (2024)
by: Guo, Lianghong, et al.
Published: (2024)
Towards Translating Real-World Code with LLMs: A Study of Translating to Rust
by: Eniser, Hasan Ferit, et al.
Published: (2024)
by: Eniser, Hasan Ferit, et al.
Published: (2024)
Unveiling Project-Specific Bias in Neural Code Models
by: Li, Zhiming, et al.
Published: (2022)
by: Li, Zhiming, et al.
Published: (2022)
Understanding the Effectiveness of Coverage Criteria for Large Language Models: A Special Angle from Jailbreak Attacks
by: Zhou, Shide, et al.
Published: (2024)
by: Zhou, Shide, et al.
Published: (2024)
Understanding LLM-Centric Challenges for Deep Learning Frameworks: An Empirical Analysis
by: Mu, Yanzhou, et al.
Published: (2025)
by: Mu, Yanzhou, et al.
Published: (2025)
A Systematic Literature Review of Code Hallucinations in LLMs: Characterization, Mitigation Methods, Challenges, and Future Directions for Reliable AI
by: Gao, Cuiyun, et al.
Published: (2025)
by: Gao, Cuiyun, et al.
Published: (2025)
NeuSemSlice: Towards Effective DNN Model Maintenance via Neuron-level Semantic Slicing
by: Zhou, Shide, et al.
Published: (2024)
by: Zhou, Shide, et al.
Published: (2024)
Unveiling Overlooked Performance Variance in Serverless Computing
by: Wen, Jinfeng, et al.
Published: (2023)
by: Wen, Jinfeng, et al.
Published: (2023)
An Empirical Study of Bugs in Modern LLM Agent Frameworks
by: Zhu, Xinxue, et al.
Published: (2026)
by: Zhu, Xinxue, et al.
Published: (2026)
Similar Items
-
Teaching Code LLMs to Use Autocompletion Tools in Repository-Level Code Generation
by: Wang, Chong, et al.
Published: (2024) -
A Study on Thinking Patterns of Large Reasoning Models in Code Generation
by: Halim, Kevin, et al.
Published: (2025) -
TREAT: A Code LLMs Trustworthiness / Reliability Evaluation and Testing Framework
by: Gao, Shuzheng, et al.
Published: (2025) -
A Dual-Loop Agent Framework for Automated Vulnerability Reproduction
by: Liu, Bin, et al.
Published: (2026) -
Knowledge-Guided Multi-Agent Framework for Application-Level Software Code Generation
by: Xiong, Qian, et al.
Published: (2025)