A Multi-Language Perspective on the Robustness of LLM Code Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Rabbi, Fazle, Ding, Zishuo, Yang, Jinqiu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HEJ-Robust: A Robustness Benchmark for LLM-Based Automated Program Repair
by: Rabbi, Fazle, et al.
Published: (2026)
by: Rabbi, Fazle, et al.
Published: (2026)
Bias Unveiled: Investigating Social Bias in LLM-Generated Code
by: Ling, Lin, et al.
Published: (2024)
by: Ling, Lin, et al.
Published: (2024)
Beyond Translation Accuracy: Addressing False Failures in LLM-Based Code Translation
by: Rabbi, Fazle, et al.
Published: (2026)
by: Rabbi, Fazle, et al.
Published: (2026)
Social Bias in LLM-Generated Code: Benchmark and Mitigation
by: Rabbi, Fazle, et al.
Published: (2026)
by: Rabbi, Fazle, et al.
Published: (2026)
Specification-Driven Code Translation Powered by Large Language Models: How Far Are We?
by: Saha, Soumit Kanti, et al.
Published: (2024)
by: Saha, Soumit Kanti, et al.
Published: (2024)
Secure-Instruct: An Automated Pipeline for Synthesizing Instruction-Tuning Datasets Using LLMs for Secure Code Generation
by: Li, Junjie, et al.
Published: (2025)
by: Li, Junjie, et al.
Published: (2025)
An Exploratory Study on Fine-Tuning Large Language Models for Secure Code Generation
by: Li, Junjie, et al.
Published: (2024)
by: Li, Junjie, et al.
Published: (2024)
BabelCoder: Agentic Code Translation with Specification Alignment
by: Rabbi, Fazle, et al.
Published: (2025)
by: Rabbi, Fazle, et al.
Published: (2025)
The Quiet Contributions: Insights into AI-Generated Silent Pull Requests
by: Hasan, S M Mahedy, et al.
Published: (2026)
by: Hasan, S M Mahedy, et al.
Published: (2026)
A Task-Level Evaluation of AI Agents in Open-Source Projects
by: Rahman, Shojibur, et al.
Published: (2026)
by: Rahman, Shojibur, et al.
Published: (2026)
CFCEval: Evaluating Security Aspects in Code Generated by Large Language Models
by: Cheng, Cheng, et al.
Published: (2025)
by: Cheng, Cheng, et al.
Published: (2025)
LLM4Perf: Large Language Models Are Effective Samplers for Multi-Objective Performance Modeling
by: Wang, Xin, et al.
Published: (2025)
by: Wang, Xin, et al.
Published: (2025)
Insights into Security-Related AI-Generated Pull Requests
by: Rabbi, Md Fazle, et al.
Published: (2026)
by: Rabbi, Md Fazle, et al.
Published: (2026)
SWE-Refactor: A Repository-Level Benchmark for Real-World LLM-Based Code Refactoring
by: Xu, Yisen, et al.
Published: (2026)
by: Xu, Yisen, et al.
Published: (2026)
Chasing the Clock: How Fast Are Vulnerabilities Fixed in the Maven Ecosystem?
by: Rabbi, Md Fazle, et al.
Published: (2025)
by: Rabbi, Md Fazle, et al.
Published: (2025)
Understanding Software Vulnerabilities in the Maven Ecosystem: Patterns, Timelines, and Risks
by: Rabbi, Md Fazle, et al.
Published: (2025)
by: Rabbi, Md Fazle, et al.
Published: (2025)
Insights into Dependency Maintenance Trends in the Maven Ecosystem
by: Chowdhury, Barisha, et al.
Published: (2025)
by: Chowdhury, Barisha, et al.
Published: (2025)
Faster Releases, Fewer Risks: A Study on Maven Artifact Vulnerabilities and Lifecycle Management
by: Shafin, Md Shafiullah, et al.
Published: (2025)
by: Shafin, Md Shafiullah, et al.
Published: (2025)
Defects4Log: Benchmarking LLMs for Logging Code Defect Detection and Reasoning
by: Wang, Xin, et al.
Published: (2025)
by: Wang, Xin, et al.
Published: (2025)
Tracking the Evolution of Static Code Warnings: the State-of-the-Art and a Better Approach
by: Li, Junjie, et al.
Published: (2022)
by: Li, Junjie, et al.
Published: (2022)
RobuNFR: Evaluating the Robustness of Large Language Models on Non-Functional Requirements Aware Code Generation
by: Lin, Feng, et al.
Published: (2025)
by: Lin, Feng, et al.
Published: (2025)
Demystifying Errors in LLM Reasoning Traces: An Empirical Study of Code Execution Simulation
by: Abdollahi, Mohammad, et al.
Published: (2025)
by: Abdollahi, Mohammad, et al.
Published: (2025)
On the Robustness Evaluation of 3D Obstacle Detection Against Specifications in Autonomous Driving
by: Pham, Tri Minh Triet, et al.
Published: (2024)
by: Pham, Tri Minh Triet, et al.
Published: (2024)
COBOLAssist: Analyzing and Fixing Compilation Errors for LLM-Powered COBOL Code Generation
by: Dau, Anh T. V., et al.
Published: (2026)
by: Dau, Anh T. V., et al.
Published: (2026)
MANTRA: Enhancing Automated Method-Level Refactoring with Contextual RAG and Multi-Agent LLM Collaboration
by: Xu, Yisen, et al.
Published: (2025)
by: Xu, Yisen, et al.
Published: (2025)
Detect--Repair--Verify for LLM-Generated Code: A Multi-Language, Multi-Granularity Empirical Study
by: Cheng, Cheng
Published: (2026)
by: Cheng, Cheng
Published: (2026)
COBOL-Coder: Domain-Adapted Large Language Models for COBOL Code Generation and Translation
by: Dau, Anh T. V., et al.
Published: (2026)
by: Dau, Anh T. V., et al.
Published: (2026)
Engineering Pitfalls in AI Coding Tools: An Empirical Study of Bugs in Claude Code, Codex, and Gemini CLI
by: Zhang, Ruixin, et al.
Published: (2026)
by: Zhang, Ruixin, et al.
Published: (2026)
A Causal Perspective on Measuring, Explaining and Mitigating Smells in LLM-Generated Code
by: Velasco, Alejandro, et al.
Published: (2025)
by: Velasco, Alejandro, et al.
Published: (2025)
Detect Repair Verify for Securing LLM Generated Code: A Multi-Language Empirical Study
by: Cheng, Cheng
Published: (2026)
by: Cheng, Cheng
Published: (2026)
Large Language Models for Code Generation: The Practitioners Perspective
by: Rasheed, Zeeshan, et al.
Published: (2025)
by: Rasheed, Zeeshan, et al.
Published: (2025)
ABTest: Behavior-Driven Testing for AI Coding Agents
by: Dai, Wuyang, et al.
Published: (2026)
by: Dai, Wuyang, et al.
Published: (2026)
OFP-Repair: Repairing Floating-point Errors via Original-Precision Arithmetic
by: Tan, Youshuai, et al.
Published: (2025)
by: Tan, Youshuai, et al.
Published: (2025)
A Preliminary Study on the Robustness of Code Generation by Large Language Models
by: Li, Zike, et al.
Published: (2025)
by: Li, Zike, et al.
Published: (2025)
Assessing, Exploiting, and Mitigating Syntactic Robustness Failures in LLM-Based Code Generation
by: Sarker, Laboni, et al.
Published: (2024)
by: Sarker, Laboni, et al.
Published: (2024)
Logging Like Humans for LLMs: Rethinking Logging via Execution and Runtime Feedback
by: Wang, Xin, et al.
Published: (2026)
by: Wang, Xin, et al.
Published: (2026)
Beyond Language Barriers: Multi-Agent Coordination for Multi-Language Code Generation
by: Moumoula, Micheline Bénédicte, et al.
Published: (2025)
by: Moumoula, Micheline Bénédicte, et al.
Published: (2025)
LLM-Based Multi-Agent Systems for Code Generation: A Multi-Vocal Literature Review
by: Rasheeda, Zeeshan, et al.
Published: (2026)
by: Rasheeda, Zeeshan, et al.
Published: (2026)
CodeCoR: An LLM-Based Self-Reflective Multi-Agent Framework for Code Generation
by: Pan, Ruwei, et al.
Published: (2025)
by: Pan, Ruwei, et al.
Published: (2025)
Perception-Guided Fuzzing for Simulated Scenario-Based Testing of Autonomous Driving Systems
by: Pham, Tri Minh Triet, et al.
Published: (2024)
by: Pham, Tri Minh Triet, et al.
Published: (2024)
Similar Items
-
HEJ-Robust: A Robustness Benchmark for LLM-Based Automated Program Repair
by: Rabbi, Fazle, et al.
Published: (2026) -
Bias Unveiled: Investigating Social Bias in LLM-Generated Code
by: Ling, Lin, et al.
Published: (2024) -
Beyond Translation Accuracy: Addressing False Failures in LLM-Based Code Translation
by: Rabbi, Fazle, et al.
Published: (2026) -
Social Bias in LLM-Generated Code: Benchmark and Mitigation
by: Rabbi, Fazle, et al.
Published: (2026) -
Specification-Driven Code Translation Powered by Large Language Models: How Far Are We?
by: Saha, Soumit Kanti, et al.
Published: (2024)