Salvato in:
| Autori principali: | Chen, Junkai, Pan, Zhiyuan, Hu, Xing, Li, Zhenhao, Li, Ge, Xia, Xin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2403.16437 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
NLPerturbator: Studying the Robustness of Code LLMs to Natural Language Variations
di: Chen, Junkai, et al.
Pubblicazione: (2024)
di: Chen, Junkai, et al.
Pubblicazione: (2024)
Model Editing for LLMs4Code: How Far are We?
di: Li, Xiaopeng, et al.
Pubblicazione: (2024)
di: Li, Xiaopeng, et al.
Pubblicazione: (2024)
Vulnerability Detection with Code Language Models: How Far Are We?
di: Ding, Yangruibo, et al.
Pubblicazione: (2024)
di: Ding, Yangruibo, et al.
Pubblicazione: (2024)
Open the Oyster: Empirical Evaluation and Improvement of Code Reasoning Confidence in LLMs
di: Wang, Shufan, et al.
Pubblicazione: (2025)
di: Wang, Shufan, et al.
Pubblicazione: (2025)
Understanding Practitioners' Expectations on Clear Code Review Comments
di: Chen, Junkai, et al.
Pubblicazione: (2024)
di: Chen, Junkai, et al.
Pubblicazione: (2024)
Environment-Aware Code Generation: How far are We?
di: Wu, Tongtong, et al.
Pubblicazione: (2026)
di: Wu, Tongtong, et al.
Pubblicazione: (2026)
SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing Scenarios
di: Chen, Junkai, et al.
Pubblicazione: (2025)
di: Chen, Junkai, et al.
Pubblicazione: (2025)
An Empirical Study of Speculative Decoding on Software Engineering Tasks
di: Li, Yijia, et al.
Pubblicazione: (2026)
di: Li, Yijia, et al.
Pubblicazione: (2026)
Runtime Failure Hunting for Physics Engine Based Software Systems: How Far Can We Go?
di: Li, Shuqing, et al.
Pubblicazione: (2025)
di: Li, Shuqing, et al.
Pubblicazione: (2025)
Reasoning Efficiently Through Adaptive Chain-of-Thought Compression: A Self-Optimizing Framework
di: Huang, Kerui, et al.
Pubblicazione: (2025)
di: Huang, Kerui, et al.
Pubblicazione: (2025)
Showing LLM-Generated Code Selectively Based on Confidence of LLMs
di: Li, Jia, et al.
Pubblicazione: (2024)
di: Li, Jia, et al.
Pubblicazione: (2024)
Octopus: On-device language model for function calling of software APIs
di: Chen, Wei, et al.
Pubblicazione: (2024)
di: Chen, Wei, et al.
Pubblicazione: (2024)
Unlocking a New Rust Programming Experience: Fast and Slow Thinking with LLMs to Conquer Undefined Behaviors
di: Jiang, Renshuang, et al.
Pubblicazione: (2025)
di: Jiang, Renshuang, et al.
Pubblicazione: (2025)
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation
di: Pan, Zhiyuan, et al.
Pubblicazione: (2025)
di: Pan, Zhiyuan, et al.
Pubblicazione: (2025)
Guarded Repair for Harm-Aware Post-hoc Replacement of LLM Mathematical Reasoning
di: Xia, Haizhou
Pubblicazione: (2026)
di: Xia, Haizhou
Pubblicazione: (2026)
To Diff or Not to Diff? Structure-Aware and Adaptive Output Formats for Efficient LLM-based Code Editing
di: Cheng, Wei, et al.
Pubblicazione: (2026)
di: Cheng, Wei, et al.
Pubblicazione: (2026)
Vulnerability-Affected Versions Identification: How Far Are We?
di: Chen, Xingchu, et al.
Pubblicazione: (2025)
di: Chen, Xingchu, et al.
Pubblicazione: (2025)
MultiFileTest: A Multi-File-Level LLM Unit Test Generation Benchmark and Impact of Error Fixing Mechanisms
di: Wang, Yibo, et al.
Pubblicazione: (2025)
di: Wang, Yibo, et al.
Pubblicazione: (2025)
Evaluating Program Semantics Reasoning with Type Inference in System F
di: He, Yifeng, et al.
Pubblicazione: (2025)
di: He, Yifeng, et al.
Pubblicazione: (2025)
Deep Learning Framework Testing via Model Mutation: How Far Are We?
di: Mu, Yanzhou, et al.
Pubblicazione: (2025)
di: Mu, Yanzhou, et al.
Pubblicazione: (2025)
Static Application Security Testing (SAST) Tools for Smart Contracts: How Far Are We?
di: Li, Kaixuan, et al.
Pubblicazione: (2024)
di: Li, Kaixuan, et al.
Pubblicazione: (2024)
ProjectEval: A Benchmark for Programming Agents Automated Evaluation on Project-Level Code Generation
di: Liu, Kaiyuan, et al.
Pubblicazione: (2025)
di: Liu, Kaiyuan, et al.
Pubblicazione: (2025)
Debug like a Human: A Large Language Model Debugger via Verifying Runtime Execution Step-by-step
di: Zhong, Li, et al.
Pubblicazione: (2024)
di: Zhong, Li, et al.
Pubblicazione: (2024)
How Far Can We Go with Practical Function-Level Program Repair?
di: Xiang, Jiahong, et al.
Pubblicazione: (2024)
di: Xiang, Jiahong, et al.
Pubblicazione: (2024)
How Programming Concepts and Neurons Are Shared in Code Language Models
di: Kargaran, Amir Hossein, et al.
Pubblicazione: (2025)
di: Kargaran, Amir Hossein, et al.
Pubblicazione: (2025)
R2Vul: Learning to Reason about Software Vulnerabilities with Reinforcement Learning and Structured Reasoning Distillation
di: Weyssow, Martin, et al.
Pubblicazione: (2025)
di: Weyssow, Martin, et al.
Pubblicazione: (2025)
Exploring Data-Efficient Adaptation of Large Language Models for Code Generation
di: Jiang, Xue, et al.
Pubblicazione: (2024)
di: Jiang, Xue, et al.
Pubblicazione: (2024)
AutoMonitor-Bench: Evaluating the Reliability of LLM-Based Misbehavior Monitor
di: Yang, Shu, et al.
Pubblicazione: (2026)
di: Yang, Shu, et al.
Pubblicazione: (2026)
Logging Like Humans for LLMs: Rethinking Logging via Execution and Runtime Feedback
di: Wang, Xin, et al.
Pubblicazione: (2026)
di: Wang, Xin, et al.
Pubblicazione: (2026)
Representation Learning for Stack Overflow Posts: How Far are We?
di: He, Junda, et al.
Pubblicazione: (2023)
di: He, Junda, et al.
Pubblicazione: (2023)
Towards Explainable Vulnerability Detection with Large Language Models
di: Mao, Qiheng, et al.
Pubblicazione: (2024)
di: Mao, Qiheng, et al.
Pubblicazione: (2024)
Evaluating the Formal Reasoning Capabilities of Large Language Models through Chomsky Hierarchy
di: Dong, Yihong, et al.
Pubblicazione: (2026)
di: Dong, Yihong, et al.
Pubblicazione: (2026)
LLM For Loop Invariant Generation and Fixing: How Far Are We?
di: Akhond, Mostafijur Rahman, et al.
Pubblicazione: (2025)
di: Akhond, Mostafijur Rahman, et al.
Pubblicazione: (2025)
CodeSpecBench: Benchmarking LLMs for Executable Behavioral Specification Generation
di: Chen, Zaoyu, et al.
Pubblicazione: (2026)
di: Chen, Zaoyu, et al.
Pubblicazione: (2026)
AutoIOT: LLM-Driven Automated Natural Language Programming for AIoT Applications
di: Shen, Leming, et al.
Pubblicazione: (2025)
di: Shen, Leming, et al.
Pubblicazione: (2025)
Deep Learning and Machine Learning, Advancing Big Data Analytics and Management: Object-Oriented Programming
di: Wang, Tianyang, et al.
Pubblicazione: (2024)
di: Wang, Tianyang, et al.
Pubblicazione: (2024)
StepCodeReasoner: Aligning Code Reasoning with Stepwise Execution Traces via Reinforcement Learning
di: Wang, Hao, et al.
Pubblicazione: (2026)
di: Wang, Hao, et al.
Pubblicazione: (2026)
An Empirical Study of Retrieval-Augmented Code Generation: Challenges and Opportunities
di: Yang, Zezhou, et al.
Pubblicazione: (2025)
di: Yang, Zezhou, et al.
Pubblicazione: (2025)
Defects4Log: Benchmarking LLMs for Logging Code Defect Detection and Reasoning
di: Wang, Xin, et al.
Pubblicazione: (2025)
di: Wang, Xin, et al.
Pubblicazione: (2025)
Deployability-Centric Infrastructure-as-Code Generation: Fail, Learn, Refine, and Succeed through LLM-Empowered DevOps Simulation
di: Zhang, Tianyi, et al.
Pubblicazione: (2025)
di: Zhang, Tianyi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
NLPerturbator: Studying the Robustness of Code LLMs to Natural Language Variations
di: Chen, Junkai, et al.
Pubblicazione: (2024) -
Model Editing for LLMs4Code: How Far are We?
di: Li, Xiaopeng, et al.
Pubblicazione: (2024) -
Vulnerability Detection with Code Language Models: How Far Are We?
di: Ding, Yangruibo, et al.
Pubblicazione: (2024) -
Open the Oyster: Empirical Evaluation and Improvement of Code Reasoning Confidence in LLMs
di: Wang, Shufan, et al.
Pubblicazione: (2025) -
Understanding Practitioners' Expectations on Clear Code Review Comments
di: Chen, Junkai, et al.
Pubblicazione: (2024)