An LLM-based Quantitative Framework for Evaluating High-Stealthy Backdoor Risks in OSS Supply Chains
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Yan, Zihe, Luo, Kai, Yang, Haoyu, Yu, Yang, Zhang, Zhuosheng, Li, Guancheng |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
OSS-UAgent: An Agent-based Usability Evaluation Framework for Open Source Software
par: Meng, Lingkai, et autres
Publié: (2025)
par: Meng, Lingkai, et autres
Publié: (2025)
Call-Chain-Aware LLM-Based Test Generation for Java Projects
par: Wang, Guancheng, et autres
Publié: (2026)
par: Wang, Guancheng, et autres
Publié: (2026)
LaQual: A Novel Framework for Automated Evaluation of LLM App Quality
par: Wang, Yan, et autres
Publié: (2025)
par: Wang, Yan, et autres
Publié: (2025)
Test Before You Deploy: Governing Updates in the LLM Supply Chain
par: Chishti, Mohd Sameen, et autres
Publié: (2026)
par: Chishti, Mohd Sameen, et autres
Publié: (2026)
Opus: A Quantitative Framework for Workflow Evaluation
par: Seroul, Alan, et autres
Publié: (2025)
par: Seroul, Alan, et autres
Publié: (2025)
Identifying the Supply Chain of AI for Trustworthiness and Risk Management in Critical Applications
par: Sheh, Raymond K., et autres
Publié: (2025)
par: Sheh, Raymond K., et autres
Publié: (2025)
ComplexMCP: Evaluation of LLM Agents in Dynamic, Interdependent, and Large-Scale Tool Sandbox
par: Li, Yuanyang, et autres
Publié: (2026)
par: Li, Yuanyang, et autres
Publié: (2026)
Benchmarks for Trajectory Safety Evaluation and Diagnosis in OpenClaw and Codex: ATBench-Claw and ATBench-Codex
par: Yang, Zhonghao, et autres
Publié: (2026)
par: Yang, Zhonghao, et autres
Publié: (2026)
ChainStream: An LLM-based Framework for Unified Synthetic Sensing
par: Liu, Jiacheng, et autres
Publié: (2024)
par: Liu, Jiacheng, et autres
Publié: (2024)
Magicoder: Empowering Code Generation with OSS-Instruct
par: Wei, Yuxiang, et autres
Publié: (2023)
par: Wei, Yuxiang, et autres
Publié: (2023)
RisConFix: LLM-based Automated Repair of Risk-Prone Drone Configurations
par: Han, Liping, et autres
Publié: (2025)
par: Han, Liping, et autres
Publié: (2025)
Evaluating the effectiveness of LLM-based interoperability
par: Falcão, Rodrigo, et autres
Publié: (2025)
par: Falcão, Rodrigo, et autres
Publié: (2025)
LLM Applications: Current Paradigms and the Next Frontier
par: Hou, Xinyi, et autres
Publié: (2025)
par: Hou, Xinyi, et autres
Publié: (2025)
Does Pass Rate Tell the Whole Story? Evaluating Design Constraint Compliance in LLM-based Issue Resolution
par: Yu, Kai, et autres
Publié: (2026)
par: Yu, Kai, et autres
Publié: (2026)
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation
par: Pan, Zhiyuan, et autres
Publié: (2025)
par: Pan, Zhiyuan, et autres
Publié: (2025)
Automated Validation of LLM-based Evaluators for Software Engineering Artifacts
par: Fandina, Ora Nova, et autres
Publié: (2025)
par: Fandina, Ora Nova, et autres
Publié: (2025)
ProbGuard: Probabilistic Runtime Monitoring for LLM Agent Safety
par: Wang, Haoyu, et autres
Publié: (2025)
par: Wang, Haoyu, et autres
Publié: (2025)
CircuChain: Disentangling Competence and Compliance in LLM Circuit Analysis
par: Ravishankara, Mayank
Publié: (2026)
par: Ravishankara, Mayank
Publié: (2026)
FasterPy: An LLM-based Code Execution Efficiency Optimization Framework
par: Wu, Yue, et autres
Publié: (2025)
par: Wu, Yue, et autres
Publié: (2025)
ALMAS: an Autonomous LLM-based Multi-Agent Software Engineering Framework
par: Tawosi, Vali, et autres
Publié: (2025)
par: Tawosi, Vali, et autres
Publié: (2025)
LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research
par: Yan, Shuo, et autres
Publié: (2025)
par: Yan, Shuo, et autres
Publié: (2025)
Canonical Intermediate Representation for LLM-based optimization problem formulation and code generation
par: Lyu, Zhongyuan, et autres
Publié: (2026)
par: Lyu, Zhongyuan, et autres
Publié: (2026)
AutoFeedback: An LLM-based Framework for Efficient and Accurate API Request Generation
par: Liu, Huanxi, et autres
Publié: (2024)
par: Liu, Huanxi, et autres
Publié: (2024)
SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks
par: Guo, Lianghong, et autres
Publié: (2025)
par: Guo, Lianghong, et autres
Publié: (2025)
VerilogReader: LLM-Aided Hardware Test Generation
par: Ma, Ruiyang, et autres
Publié: (2024)
par: Ma, Ruiyang, et autres
Publié: (2024)
Building Trust in the Skies: A Knowledge-Grounded LLM-based Framework for Aviation Safety
par: Iyengar, Anirudh, et autres
Publié: (2026)
par: Iyengar, Anirudh, et autres
Publié: (2026)
RefAgent: A Multi-agent LLM-based Framework for Automatic Software Refactoring
par: Oueslati, Khouloud, et autres
Publié: (2025)
par: Oueslati, Khouloud, et autres
Publié: (2025)
ProcCtrlBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding Agents
par: He, Jiawei, et autres
Publié: (2026)
par: He, Jiawei, et autres
Publié: (2026)
LLM-Empowered Event-Chain Driven Code Generation for ADAS in SDV systems
par: Petrovic, Nenad, et autres
Publié: (2025)
par: Petrovic, Nenad, et autres
Publié: (2025)
OrcaLoca: An LLM Agent Framework for Software Issue Localization
par: Yu, Zhongming, et autres
Publié: (2025)
par: Yu, Zhongming, et autres
Publié: (2025)
LLM-enabled Applications Require System-Level Threat Monitoring
par: Zhang, Yedi, et autres
Publié: (2026)
par: Zhang, Yedi, et autres
Publié: (2026)
WebVIA: A Web-based Vision-Language Agentic Framework for Interactive and Verifiable UI-to-Code Generation
par: Xu, Mingde, et autres
Publié: (2025)
par: Xu, Mingde, et autres
Publié: (2025)
Large Language Model Supply Chain: Open Problems From the Security Perspective
par: Hu, Qiang, et autres
Publié: (2024)
par: Hu, Qiang, et autres
Publié: (2024)
Faver: Boosting LLM-based RTL Generation with Function Abstracted Verifiable Middleware
par: Mu, Jianan, et autres
Publié: (2025)
par: Mu, Jianan, et autres
Publié: (2025)
Rubric Is All You Need: Enhancing LLM-based Code Evaluation With Question-Specific Rubrics
par: Pathak, Aditya, et autres
Publié: (2025)
par: Pathak, Aditya, et autres
Publié: (2025)
MAGIS: LLM-Based Multi-Agent Framework for GitHub Issue Resolution
par: Tao, Wei, et autres
Publié: (2024)
par: Tao, Wei, et autres
Publié: (2024)
Copilot Evaluation Harness: Evaluating LLM-Guided Software Programming
par: Agarwal, Anisha, et autres
Publié: (2024)
par: Agarwal, Anisha, et autres
Publié: (2024)
An LLM-Assisted Easy-to-Trigger Backdoor Attack on Code Completion Models: Injecting Disguised Vulnerabilities against Strong Detection
par: Yan, Shenao, et autres
Publié: (2024)
par: Yan, Shenao, et autres
Publié: (2024)
ACE: Self-Evolving LLM Coding Framework via Adversarial Unit Test Generation and Preference Optimization
par: Huang, Yixu, et autres
Publié: (2026)
par: Huang, Yixu, et autres
Publié: (2026)
EvoDev: An Iterative Feature-Driven Framework for End-to-End Software Development with LLM-based Agents
par: Liu, Junwei, et autres
Publié: (2025)
par: Liu, Junwei, et autres
Publié: (2025)
Documents similaires
-
OSS-UAgent: An Agent-based Usability Evaluation Framework for Open Source Software
par: Meng, Lingkai, et autres
Publié: (2025) -
Call-Chain-Aware LLM-Based Test Generation for Java Projects
par: Wang, Guancheng, et autres
Publié: (2026) -
LaQual: A Novel Framework for Automated Evaluation of LLM App Quality
par: Wang, Yan, et autres
Publié: (2025) -
Test Before You Deploy: Governing Updates in the LLM Supply Chain
par: Chishti, Mohd Sameen, et autres
Publié: (2026) -
Opus: A Quantitative Framework for Workflow Evaluation
par: Seroul, Alan, et autres
Publié: (2025)