SimCT: A Simple Consistency Test Protocol in LLMs Development Lifecycle
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Fufangchen, Jin, Guoqiang, Zhao, Rui, Huang, Jiangheng, Tan, Fei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Consistency Matters: Explore LLMs Consistency From a Black-Box Perspective
by: Zhao, Fufangchen, et al.
Published: (2024)
by: Zhao, Fufangchen, et al.
Published: (2024)
Prompting Large Language Models to Tackle the Full Software Development Lifecycle: A Case Study
by: Li, Bowen, et al.
Published: (2024)
by: Li, Bowen, et al.
Published: (2024)
Can LLMs Generate High-Quality Test Cases for Algorithm Problems? TestCase-Eval: A Systematic Evaluation of Fault Coverage and Exposure
by: Yang, Zheyuan, et al.
Published: (2025)
by: Yang, Zheyuan, et al.
Published: (2025)
TestExplora: Benchmarking LLMs for Proactive Bug Discovery via Repository-Level Test Generation
by: Liu, Steven, et al.
Published: (2026)
by: Liu, Steven, et al.
Published: (2026)
Securing Computer-Use Agents: A Unified Architecture-Lifecycle Framework for Deployment-Grounded Reliability
by: Chen, Zejian, et al.
Published: (2026)
by: Chen, Zejian, et al.
Published: (2026)
FairCoder: Evaluating Social Bias of LLMs in Code Generation
by: Du, Yongkang, et al.
Published: (2025)
by: Du, Yongkang, et al.
Published: (2025)
Benchmarking LLMs for Unit Test Generation from Real-World Functions
by: Huang, Dong, et al.
Published: (2025)
by: Huang, Dong, et al.
Published: (2025)
Bias Testing and Mitigation in Black Box LLMs using Metamorphic Relations
by: Salimian, Sina, et al.
Published: (2025)
by: Salimian, Sina, et al.
Published: (2025)
EvoCodeBench: An Evolving Code Generation Benchmark with Domain-Specific Evaluations
by: Li, Jia, et al.
Published: (2024)
by: Li, Jia, et al.
Published: (2024)
MultiFileTest: A Multi-File-Level LLM Unit Test Generation Benchmark and Impact of Error Fixing Mechanisms
by: Wang, Yibo, et al.
Published: (2025)
by: Wang, Yibo, et al.
Published: (2025)
A Comparative Study on the Impact of Test-Driven Development (TDD) and Behavior-Driven Development (BDD) on Enterprise Software Delivery Effectiveness
by: Cui, Jun
Published: (2024)
by: Cui, Jun
Published: (2024)
DiffuTester: Accelerating Unit Test Generation for Diffusion LLMs via Mining Structural Pattern
by: Yang, Lekang, et al.
Published: (2025)
by: Yang, Lekang, et al.
Published: (2025)
MathDuels: Evaluating LLMs as Problem Posers and Solvers
by: Xu, Zhiqiu, et al.
Published: (2026)
by: Xu, Zhiqiu, et al.
Published: (2026)
Can LLMs Generate Reliable Test Case Generators? A Study on Competition-Level Programming Problems
by: Cao, Yuhan, et al.
Published: (2025)
by: Cao, Yuhan, et al.
Published: (2025)
Showing LLM-Generated Code Selectively Based on Confidence of LLMs
by: Li, Jia, et al.
Published: (2024)
by: Li, Jia, et al.
Published: (2024)
CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based Verification
by: Tian, Yuchen, et al.
Published: (2024)
by: Tian, Yuchen, et al.
Published: (2024)
AutoTestForge: A Multidimensional Automated Testing Framework for Natural Language Processing Models
by: Xing, Hengrui, et al.
Published: (2025)
by: Xing, Hengrui, et al.
Published: (2025)
GUI Test Migration via Abstraction and Concretization
by: Zhang, Yakun, et al.
Published: (2024)
by: Zhang, Yakun, et al.
Published: (2024)
Evaluating LLMs on Sequential API Call Through Automated Test Generation
by: Huang, Yuheng, et al.
Published: (2025)
by: Huang, Yuheng, et al.
Published: (2025)
Multi-Programming Language Sandbox for LLMs
by: Dou, Shihan, et al.
Published: (2024)
by: Dou, Shihan, et al.
Published: (2024)
Measuring the Influence of Incorrect Code on Test Generation
by: Huang, Dong, et al.
Published: (2024)
by: Huang, Dong, et al.
Published: (2024)
Unlocking a New Rust Programming Experience: Fast and Slow Thinking with LLMs to Conquer Undefined Behaviors
by: Jiang, Renshuang, et al.
Published: (2025)
by: Jiang, Renshuang, et al.
Published: (2025)
Solver-Independent Automated Problem Formulation via LLMs for High-Cost Simulation-Driven Design
by: Li, Yuchen, et al.
Published: (2025)
by: Li, Yuchen, et al.
Published: (2025)
RePair: Automated Program Repair with Process-based Feedback
by: Zhao, Yuze, et al.
Published: (2024)
by: Zhao, Yuze, et al.
Published: (2024)
DuET: Dual Execution for Test Output Prediction with Generated Code and Pseudocode
by: Han, Hojae, et al.
Published: (2026)
by: Han, Hojae, et al.
Published: (2026)
OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs
by: Ahmad, Wasi Uddin, et al.
Published: (2025)
by: Ahmad, Wasi Uddin, et al.
Published: (2025)
DependEval: Benchmarking LLMs for Repository Dependency Understanding
by: Du, Junjia, et al.
Published: (2025)
by: Du, Junjia, et al.
Published: (2025)
CodePivot: Bootstrapping Multilingual Transpilation in LLMs via Reinforcement Learning without Parallel Corpora
by: Li, Shangyu, et al.
Published: (2026)
by: Li, Shangyu, et al.
Published: (2026)
DevEval: A Manually-Annotated Code Generation Benchmark Aligned with Real-World Code Repositories
by: Li, Jia, et al.
Published: (2024)
by: Li, Jia, et al.
Published: (2024)
Nexus: Execution-Grounded Multi-Agent Test Oracle Synthesis
by: Huang, Dong, et al.
Published: (2025)
by: Huang, Dong, et al.
Published: (2025)
CCT-Code: Cross-Consistency Training for Multilingual Clone Detection and Code Search
by: Tikhonov, Anton, et al.
Published: (2023)
by: Tikhonov, Anton, et al.
Published: (2023)
MUCOCO: Automated Consistency Testing of Code LLMs
by: Chou, Chua Jin, et al.
Published: (2026)
by: Chou, Chua Jin, et al.
Published: (2026)
From LLMs to LLM-based Agents for Software Engineering: A Survey of Current, Challenges and Future
by: Jin, Haolin, et al.
Published: (2024)
by: Jin, Haolin, et al.
Published: (2024)
Enhancing Code LLMs with Reinforcement Learning in Code Generation: A Survey
by: Wang, Junqiao, et al.
Published: (2024)
by: Wang, Junqiao, et al.
Published: (2024)
ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions
by: Ding, Xianzhong, et al.
Published: (2026)
by: Ding, Xianzhong, et al.
Published: (2026)
Leveraging LLMs for Grammar Adaptation: A Study on Metamodel-Grammar Co-Evolution
by: Zhang, Weixing, et al.
Published: (2026)
by: Zhang, Weixing, et al.
Published: (2026)
Evaluation of Code LLMs on Geospatial Code Generation
by: Gramacki, Piotr, et al.
Published: (2024)
by: Gramacki, Piotr, et al.
Published: (2024)
LLMs in Mobile Apps: Practices, Challenges, and Opportunities
by: Hau, Kimberly, et al.
Published: (2025)
by: Hau, Kimberly, et al.
Published: (2025)
Alibaba LingmaAgent: Improving Automated Issue Resolution via Comprehensive Repository Exploration
by: Ma, Yingwei, et al.
Published: (2024)
by: Ma, Yingwei, et al.
Published: (2024)
CodeContests-O: Powering LLMs via Feedback-Driven Iterative Test Case Generation
by: Cai, Jianfeng, et al.
Published: (2026)
by: Cai, Jianfeng, et al.
Published: (2026)
Similar Items
-
Consistency Matters: Explore LLMs Consistency From a Black-Box Perspective
by: Zhao, Fufangchen, et al.
Published: (2024) -
Prompting Large Language Models to Tackle the Full Software Development Lifecycle: A Case Study
by: Li, Bowen, et al.
Published: (2024) -
Can LLMs Generate High-Quality Test Cases for Algorithm Problems? TestCase-Eval: A Systematic Evaluation of Fault Coverage and Exposure
by: Yang, Zheyuan, et al.
Published: (2025) -
TestExplora: Benchmarking LLMs for Proactive Bug Discovery via Repository-Level Test Generation
by: Liu, Steven, et al.
Published: (2026) -
Securing Computer-Use Agents: A Unified Architecture-Lifecycle Framework for Deployment-Grounded Reliability
by: Chen, Zejian, et al.
Published: (2026)