Salvato in:
| Autori principali: | Yan, Zihe, Luo, Kai, Yang, Haoyu, Yu, Yang, Zhang, Zhuosheng, Li, Guancheng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2511.13341 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
OSS-UAgent: An Agent-based Usability Evaluation Framework for Open Source Software
di: Meng, Lingkai, et al.
Pubblicazione: (2025)
di: Meng, Lingkai, et al.
Pubblicazione: (2025)
Call-Chain-Aware LLM-Based Test Generation for Java Projects
di: Wang, Guancheng, et al.
Pubblicazione: (2026)
di: Wang, Guancheng, et al.
Pubblicazione: (2026)
LaQual: A Novel Framework for Automated Evaluation of LLM App Quality
di: Wang, Yan, et al.
Pubblicazione: (2025)
di: Wang, Yan, et al.
Pubblicazione: (2025)
Test Before You Deploy: Governing Updates in the LLM Supply Chain
di: Chishti, Mohd Sameen, et al.
Pubblicazione: (2026)
di: Chishti, Mohd Sameen, et al.
Pubblicazione: (2026)
Opus: A Quantitative Framework for Workflow Evaluation
di: Seroul, Alan, et al.
Pubblicazione: (2025)
di: Seroul, Alan, et al.
Pubblicazione: (2025)
Identifying the Supply Chain of AI for Trustworthiness and Risk Management in Critical Applications
di: Sheh, Raymond K., et al.
Pubblicazione: (2025)
di: Sheh, Raymond K., et al.
Pubblicazione: (2025)
Magicoder: Empowering Code Generation with OSS-Instruct
di: Wei, Yuxiang, et al.
Pubblicazione: (2023)
di: Wei, Yuxiang, et al.
Pubblicazione: (2023)
Benchmarks for Trajectory Safety Evaluation and Diagnosis in OpenClaw and Codex: ATBench-Claw and ATBench-Codex
di: Yang, Zhonghao, et al.
Pubblicazione: (2026)
di: Yang, Zhonghao, et al.
Pubblicazione: (2026)
ChainStream: An LLM-based Framework for Unified Synthetic Sensing
di: Liu, Jiacheng, et al.
Pubblicazione: (2024)
di: Liu, Jiacheng, et al.
Pubblicazione: (2024)
ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox
di: Li, Yuanyang, et al.
Pubblicazione: (2026)
di: Li, Yuanyang, et al.
Pubblicazione: (2026)
LLM Applications: Current Paradigms and the Next Frontier
di: Hou, Xinyi, et al.
Pubblicazione: (2025)
di: Hou, Xinyi, et al.
Pubblicazione: (2025)
RisConFix: LLM-based Automated Repair of Risk-Prone Drone Configurations
di: Han, Liping, et al.
Pubblicazione: (2025)
di: Han, Liping, et al.
Pubblicazione: (2025)
Evaluating the effectiveness of LLM-based interoperability
di: Falcão, Rodrigo, et al.
Pubblicazione: (2025)
di: Falcão, Rodrigo, et al.
Pubblicazione: (2025)
Does Pass Rate Tell the Whole Story? Evaluating Design Constraint Compliance in LLM-based Issue Resolution
di: Yu, Kai, et al.
Pubblicazione: (2026)
di: Yu, Kai, et al.
Pubblicazione: (2026)
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation
di: Pan, Zhiyuan, et al.
Pubblicazione: (2025)
di: Pan, Zhiyuan, et al.
Pubblicazione: (2025)
ProbGuard: Probabilistic Runtime Monitoring for LLM Agent Safety
di: Wang, Haoyu, et al.
Pubblicazione: (2025)
di: Wang, Haoyu, et al.
Pubblicazione: (2025)
LLM-enabled Applications Require System-Level Threat Monitoring
di: Zhang, Yedi, et al.
Pubblicazione: (2026)
di: Zhang, Yedi, et al.
Pubblicazione: (2026)
Automated Validation of LLM-based Evaluators for Software Engineering Artifacts
di: Fandina, Ora Nova, et al.
Pubblicazione: (2025)
di: Fandina, Ora Nova, et al.
Pubblicazione: (2025)
SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks
di: Guo, Lianghong, et al.
Pubblicazione: (2025)
di: Guo, Lianghong, et al.
Pubblicazione: (2025)
LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research
di: Yan, Shuo, et al.
Pubblicazione: (2025)
di: Yan, Shuo, et al.
Pubblicazione: (2025)
CircuChain: Disentangling Competence and Compliance in LLM Circuit Analysis
di: Ravishankara, Mayank
Pubblicazione: (2026)
di: Ravishankara, Mayank
Pubblicazione: (2026)
FasterPy: An LLM-based Code Execution Efficiency Optimization Framework
di: Wu, Yue, et al.
Pubblicazione: (2025)
di: Wu, Yue, et al.
Pubblicazione: (2025)
ALMAS: an Autonomous LLM-based Multi-Agent Software Engineering Framework
di: Tawosi, Vali, et al.
Pubblicazione: (2025)
di: Tawosi, Vali, et al.
Pubblicazione: (2025)
Large Language Model Supply Chain: Open Problems From the Security Perspective
di: Hu, Qiang, et al.
Pubblicazione: (2024)
di: Hu, Qiang, et al.
Pubblicazione: (2024)
VerilogReader: LLM-Aided Hardware Test Generation
di: Ma, Ruiyang, et al.
Pubblicazione: (2024)
di: Ma, Ruiyang, et al.
Pubblicazione: (2024)
Canonical Intermediate Representation for LLM-based optimization problem formulation and code generation
di: Lyu, Zhongyuan, et al.
Pubblicazione: (2026)
di: Lyu, Zhongyuan, et al.
Pubblicazione: (2026)
AutoFeedback: An LLM-based Framework for Efficient and Accurate API Request Generation
di: Liu, Huanxi, et al.
Pubblicazione: (2024)
di: Liu, Huanxi, et al.
Pubblicazione: (2024)
An LLM-Assisted Easy-to-Trigger Backdoor Attack on Code Completion Models: Injecting Disguised Vulnerabilities against Strong Detection
di: Yan, Shenao, et al.
Pubblicazione: (2024)
di: Yan, Shenao, et al.
Pubblicazione: (2024)
ProcCtrlBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding Agents
di: He, Jiawei, et al.
Pubblicazione: (2026)
di: He, Jiawei, et al.
Pubblicazione: (2026)
WebVIA: A Web-based Vision-Language Agentic Framework for Interactive and Verifiable UI-to-Code Generation
di: Xu, Mingde, et al.
Pubblicazione: (2025)
di: Xu, Mingde, et al.
Pubblicazione: (2025)
AppForge: From Assistant to Independent Developer -- Are GPTs Ready for Software Development?
di: Ran, Dezhi, et al.
Pubblicazione: (2025)
di: Ran, Dezhi, et al.
Pubblicazione: (2025)
LLM-Empowered Event-Chain Driven Code Generation for ADAS in SDV systems
di: Petrovic, Nenad, et al.
Pubblicazione: (2025)
di: Petrovic, Nenad, et al.
Pubblicazione: (2025)
OrcaLoca: An LLM Agent Framework for Software Issue Localization
di: Yu, Zhongming, et al.
Pubblicazione: (2025)
di: Yu, Zhongming, et al.
Pubblicazione: (2025)
Building Trust in the Skies: A Knowledge-Grounded LLM-based Framework for Aviation Safety
di: Iyengar, Anirudh, et al.
Pubblicazione: (2026)
di: Iyengar, Anirudh, et al.
Pubblicazione: (2026)
RefAgent: A Multi-agent LLM-based Framework for Automatic Software Refactoring
di: Oueslati, Khouloud, et al.
Pubblicazione: (2025)
di: Oueslati, Khouloud, et al.
Pubblicazione: (2025)
LLM4Vuln: A Unified Evaluation Framework for Decoupling and Enhancing LLMs' Vulnerability Reasoning
di: Sun, Yuqiang, et al.
Pubblicazione: (2024)
di: Sun, Yuqiang, et al.
Pubblicazione: (2024)
RepoMasterEval: Evaluating Code Completion via Real-World Repositories
di: Wu, Qinyun, et al.
Pubblicazione: (2024)
di: Wu, Qinyun, et al.
Pubblicazione: (2024)
VSRQ: Quantitative Assessment Method for Safety Risk of Vehicle Intelligent Connected System
di: Zhang, Tian, et al.
Pubblicazione: (2023)
di: Zhang, Tian, et al.
Pubblicazione: (2023)
Faver: Boosting LLM-based RTL Generation with Function Abstracted Verifiable Middleware
di: Mu, Jianan, et al.
Pubblicazione: (2025)
di: Mu, Jianan, et al.
Pubblicazione: (2025)
Rubric Is All You Need: Enhancing LLM-based Code Evaluation With Question-Specific Rubrics
di: Pathak, Aditya, et al.
Pubblicazione: (2025)
di: Pathak, Aditya, et al.
Pubblicazione: (2025)
Documenti analoghi
-
OSS-UAgent: An Agent-based Usability Evaluation Framework for Open Source Software
di: Meng, Lingkai, et al.
Pubblicazione: (2025) -
Call-Chain-Aware LLM-Based Test Generation for Java Projects
di: Wang, Guancheng, et al.
Pubblicazione: (2026) -
LaQual: A Novel Framework for Automated Evaluation of LLM App Quality
di: Wang, Yan, et al.
Pubblicazione: (2025) -
Test Before You Deploy: Governing Updates in the LLM Supply Chain
di: Chishti, Mohd Sameen, et al.
Pubblicazione: (2026) -
Opus: A Quantitative Framework for Workflow Evaluation
di: Seroul, Alan, et al.
Pubblicazione: (2025)