TREAT: A Code LLMs Trustworthiness / Reliability Evaluation and Testing Framework
Fuente:
arXiv
Saved in:
| Main Authors: | Gao, Shuzheng, Li, Eric John, Lam, Man Ho, Xiao, Jingyu, Wan, Yuxuan, Wang, Chaozheng, Tik, Ng Man, Lyu, Michael R. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CodeCrash: Exposing LLM Fragility to Misleading Natural Language in Code Reasoning
by: Lam, Man Ho, et al.
Published: (2025)
by: Lam, Man Ho, et al.
Published: (2025)
SEER: Enhancing Chain-of-Thought Code Generation through Self-Exploring Deep Reasoning
by: Gao, Shuzheng, et al.
Published: (2025)
by: Gao, Shuzheng, et al.
Published: (2025)
DesignBench: A Comprehensive Benchmark for MLLM-based Front-end Code Generation
by: Xiao, Jingyu, et al.
Published: (2025)
by: Xiao, Jingyu, et al.
Published: (2025)
ComUICoder: Component-based Reusable UI Code Generation for Complex Websites via Semantic Segmentation and Element-wise Feedback
by: Xiao, Jingyu, et al.
Published: (2026)
by: Xiao, Jingyu, et al.
Published: (2026)
A Systematic Evaluation of Large Code Models in API Suggestion: When, Which, and How
by: Wang, Chaozheng, et al.
Published: (2024)
by: Wang, Chaozheng, et al.
Published: (2024)
RAG or Fine-tuning? A Comparative Study on LCMs-based Code Completion in Industry
by: Wang, Chaozheng, et al.
Published: (2025)
by: Wang, Chaozheng, et al.
Published: (2025)
SWE-Chain: Benchmarking Coding Agents on Chained Release-Level Package Upgrades
by: Lam, Man Ho, et al.
Published: (2026)
by: Lam, Man Ho, et al.
Published: (2026)
SR-Eval: Evaluating LLMs on Code Generation under Stepwise Requirement Refinement
by: Zhan, Zexun, et al.
Published: (2025)
by: Zhan, Zexun, et al.
Published: (2025)
Search-Based LLMs for Code Optimization
by: Gao, Shuzheng, et al.
Published: (2024)
by: Gao, Shuzheng, et al.
Published: (2024)
Cascaded Code Editing: Large-Small Model Collaboration for Effective and Efficient Code Editing
by: Wang, Chaozheng, et al.
Published: (2026)
by: Wang, Chaozheng, et al.
Published: (2026)
JSProtect: A Scalable Obfuscation Framework for Mini-Games in WeChat
by: Li, Zhihao, et al.
Published: (2025)
by: Li, Zhihao, et al.
Published: (2025)
What Makes Good In-context Demonstrations for Code Intelligence Tasks with LLMs?
by: Gao, Shuzheng, et al.
Published: (2023)
by: Gao, Shuzheng, et al.
Published: (2023)
The Prompt Alchemist: Automated LLM-Tailored Prompt Optimization for Test Case Generation
by: Gao, Shuzheng, et al.
Published: (2025)
by: Gao, Shuzheng, et al.
Published: (2025)
From Runnable to Shippable: Multi-Agent Test-Driven Development for Generating Full-Stack Web Applications from Requirements
by: Wan, Yuxuan, et al.
Published: (2026)
by: Wan, Yuxuan, et al.
Published: (2026)
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs
by: Zhou, Zenghui, et al.
Published: (2026)
by: Zhou, Zenghui, et al.
Published: (2026)
ComplexCodeEval: A Benchmark for Evaluating Large Code Models on More Complex Code
by: Feng, Jia, et al.
Published: (2024)
by: Feng, Jia, et al.
Published: (2024)
Automatically Generating UI Code from Screenshot: A Divide-and-Conquer-Based Approach
by: Wan, Yuxuan, et al.
Published: (2024)
by: Wan, Yuxuan, et al.
Published: (2024)
EfficientUICoder: Efficient MLLM-based UI Code Generation via Input and Output Token Compression
by: Xiao, Jingyu, et al.
Published: (2025)
by: Xiao, Jingyu, et al.
Published: (2025)
Automatically Generating Web Applications from Requirements Via Multi-Agent Test-Driven Development
by: Wan, Yuxuan, et al.
Published: (2025)
by: Wan, Yuxuan, et al.
Published: (2025)
Towards Trustworthy LLMs for Code: A Data-Centric Synergistic Auditing Framework
by: Wang, Chong, et al.
Published: (2024)
by: Wang, Chong, et al.
Published: (2024)
Learning in the Wild: Towards Leveraging Unlabeled Data for Effectively Tuning Pre-trained Code Models
by: Gao, Shuzheng, et al.
Published: (2024)
by: Gao, Shuzheng, et al.
Published: (2024)
On Evaluating the Efficiency of Source Code Generated by LLMs
by: Niu, Changan, et al.
Published: (2024)
by: Niu, Changan, et al.
Published: (2024)
Exploring Multi-Lingual Bias of Large Code Models in Code Generation
by: Wang, Chaozheng, et al.
Published: (2024)
by: Wang, Chaozheng, et al.
Published: (2024)
MRWeb: An Exploration of Generating Multi-Page Resource-Aware Web Code from UI Designs
by: Wan, Yuxuan, et al.
Published: (2024)
by: Wan, Yuxuan, et al.
Published: (2024)
Position Paper: Programming Language Techniques for Bridging LLM Code Generation Semantic Gaps
by: Du, Yalong, et al.
Published: (2025)
by: Du, Yalong, et al.
Published: (2025)
Many-Objective Search-Based Coverage-Guided Automatic Test Generation for Deep Neural Networks
by: Li, Dongcheng, et al.
Published: (2024)
by: Li, Dongcheng, et al.
Published: (2024)
Ising-based Test Optimization and Benchmarking
by: Yang, Yige, et al.
Published: (2026)
by: Yang, Yige, et al.
Published: (2026)
SCALE: Constructing Structured Natural Language Comment Trees for Software Vulnerability Detection
by: Wen, Xin-Cheng, et al.
Published: (2024)
by: Wen, Xin-Cheng, et al.
Published: (2024)
90% Faster, 100% Code-Free: MLLM-Driven Zero-Code 3D Game Development
by: Yang, Runxin, et al.
Published: (2025)
by: Yang, Runxin, et al.
Published: (2025)
A Deep Dive into Retrieval-Augmented Generation for Code Completion: Experience on WeChat
by: Yang, Zezhou, et al.
Published: (2025)
by: Yang, Zezhou, et al.
Published: (2025)
BDiff: Block-aware and Accurate Text-based Code Differencing
by: Lu, Yao, et al.
Published: (2025)
by: Lu, Yao, et al.
Published: (2025)
Causes and Effects of Fitness Landscapes in System Test Generation: A Replication Study
by: Sahin, Omur, et al.
Published: (2025)
by: Sahin, Omur, et al.
Published: (2025)
SGCR: A Specification-Grounded Framework for Trustworthy LLM Code Review
by: Wang, Kai, et al.
Published: (2025)
by: Wang, Kai, et al.
Published: (2025)
UIBenchKit: A unified toolkit for design-to-code model evaluation
by: Le, Chinh T., et al.
Published: (2026)
by: Le, Chinh T., et al.
Published: (2026)
Automated Prompt Generation for Code Intelligence: An Empirical study and Experience in WeChat
by: Ji, Kexing, et al.
Published: (2025)
by: Ji, Kexing, et al.
Published: (2025)
MUCOCO: Automated Consistency Testing of Code LLMs
by: Chou, Chua Jin, et al.
Published: (2026)
by: Chou, Chua Jin, et al.
Published: (2026)
Interaction2Code: Benchmarking MLLM-based Interactive Webpage Code Generation from Interactive Prototyping
by: Xiao, Jingyu, et al.
Published: (2024)
by: Xiao, Jingyu, et al.
Published: (2024)
Algorithm-Based Pipeline for Reliable and Intent-Preserving Code Translation with LLMs
by: Dipto, Shahriar Rumi, et al.
Published: (2026)
by: Dipto, Shahriar Rumi, et al.
Published: (2026)
Toward Automated and Trustworthy Scientific Analysis and Visualization with LLM-Generated Code
by: Chakroborti, Apu Kumar, et al.
Published: (2025)
by: Chakroborti, Apu Kumar, et al.
Published: (2025)
SPVR: syntax-to-prompt vulnerability repair based on large language models
by: Wang, Ruoke, et al.
Published: (2024)
by: Wang, Ruoke, et al.
Published: (2024)
Similar Items
-
CodeCrash: Exposing LLM Fragility to Misleading Natural Language in Code Reasoning
by: Lam, Man Ho, et al.
Published: (2025) -
SEER: Enhancing Chain-of-Thought Code Generation through Self-Exploring Deep Reasoning
by: Gao, Shuzheng, et al.
Published: (2025) -
DesignBench: A Comprehensive Benchmark for MLLM-based Front-end Code Generation
by: Xiao, Jingyu, et al.
Published: (2025) -
ComUICoder: Component-based Reusable UI Code Generation for Complex Websites via Semantic Segmentation and Element-wise Feedback
by: Xiao, Jingyu, et al.
Published: (2026) -
A Systematic Evaluation of Large Code Models in API Suggestion: When, Which, and How
by: Wang, Chaozheng, et al.
Published: (2024)