Test Before You Deploy: Governing Updates in the LLM Supply Chain
Fuente:
arXiv
Saved in:
| Main Authors: | Chishti, Mohd Sameen, Oyinloye, Damilare Peter, Li, Jingyue |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Feature-Centric Methodology for Analyzing Cross-Chain NFT Migration Compatibility
by: Chishti, Mohd Sameen, et al.
Published: (2026)
by: Chishti, Mohd Sameen, et al.
Published: (2026)
AgentReputation: A Decentralized Agentic AI Reputation Framework
by: Chishti, Mohd Sameen, et al.
Published: (2026)
by: Chishti, Mohd Sameen, et al.
Published: (2026)
A Proof of Success and Reward Distribution Protocol for Multi-bridge Architecture in Cross-chain Communication
by: Oyinloye, Damilare Peter, et al.
Published: (2025)
by: Oyinloye, Damilare Peter, et al.
Published: (2025)
Test It Before You Trust It: Applying Software Testing for Trustworthy In-context Learning
by: Racharak, Teeradaj, et al.
Published: (2025)
by: Racharak, Teeradaj, et al.
Published: (2025)
Parameter-Efficient Fine-Tuning of Large Language Models for Unit Test Generation: An Empirical Study
by: Storhaug, André, et al.
Published: (2024)
by: Storhaug, André, et al.
Published: (2024)
Reproducible, Explainable, and Effective Evaluations of Agentic AI for Software Engineering
by: Li, Jingyue, et al.
Published: (2026)
by: Li, Jingyue, et al.
Published: (2026)
An LLM-based Quantitative Framework for Evaluating High-Stealthy Backdoor Risks in OSS Supply Chains
by: Yan, Zihe, et al.
Published: (2025)
by: Yan, Zihe, et al.
Published: (2025)
Repair-R1: Better Test Before Repair
by: Hu, Haichuan, et al.
Published: (2025)
by: Hu, Haichuan, et al.
Published: (2025)
You Name It, I Run It: An LLM Agent to Execute Tests of Arbitrary Projects
by: Bouzenia, Islem, et al.
Published: (2024)
by: Bouzenia, Islem, et al.
Published: (2024)
Call-Chain-Aware LLM-Based Test Generation for Java Projects
by: Wang, Guancheng, et al.
Published: (2026)
by: Wang, Guancheng, et al.
Published: (2026)
GitHub's Copilot Code Review: Can AI Spot Security Flaws Before You Commit?
by: Amro, Amena, et al.
Published: (2025)
by: Amro, Amena, et al.
Published: (2025)
You Don't Know Until You Click:Automated GUI Testing for Production-Ready Software Evaluation
by: Bian, Yutong, et al.
Published: (2025)
by: Bian, Yutong, et al.
Published: (2025)
LLM-Driven Kernel Evolution: Automating Driver Updates in Linux
by: Kharlamova, Arina, et al.
Published: (2025)
by: Kharlamova, Arina, et al.
Published: (2025)
Favia: Forensic Agent for Vulnerability-fix Identification and Analysis
by: Storhaug, André, et al.
Published: (2026)
by: Storhaug, André, et al.
Published: (2026)
Look Before You Leap: An Exploratory Study of Uncertainty Measurement for Large Language Models
by: Huang, Yuheng, et al.
Published: (2023)
by: Huang, Yuheng, et al.
Published: (2023)
Harden and Catch for Just-in-Time Assured LLM-Based Software Testing: Open Research Challenges
by: Harman, Mark, et al.
Published: (2025)
by: Harman, Mark, et al.
Published: (2025)
VerilogReader: LLM-Aided Hardware Test Generation
by: Ma, Ruiyang, et al.
Published: (2024)
by: Ma, Ruiyang, et al.
Published: (2024)
Early External Safety Testing of OpenAI's o3-mini: Insights from the Pre-Deployment Evaluation
by: Arrieta, Aitor, et al.
Published: (2025)
by: Arrieta, Aitor, et al.
Published: (2025)
LLMs are All You Need? Improving Fuzz Testing for MOJO with Large Language Models
by: Huang, Linghan, et al.
Published: (2025)
by: Huang, Linghan, et al.
Published: (2025)
CircuChain: Disentangling Competence and Compliance in LLM Circuit Analysis
by: Ravishankara, Mayank
Published: (2026)
by: Ravishankara, Mayank
Published: (2026)
Model-Enhanced LLM-Driven VUI Testing of VPA Apps
by: Li, Suwan, et al.
Published: (2024)
by: Li, Suwan, et al.
Published: (2024)
Improving Code Translation with Syntax-Guided and Semantic-aware Preference Optimization
by: Wu, Yuhan, et al.
Published: (2026)
by: Wu, Yuhan, et al.
Published: (2026)
Bootstrapping Code Translation with Weighted Multilanguage Exploration
by: Wu, Yuhan, et al.
Published: (2026)
by: Wu, Yuhan, et al.
Published: (2026)
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation
by: Cui, Yi
Published: (2025)
by: Cui, Yi
Published: (2025)
Deploy-Master: Automating the Deployment of 50,000+ Agent-Ready Scientific Tools in One Day
by: Wang, Yi, et al.
Published: (2026)
by: Wang, Yi, et al.
Published: (2026)
Identifying the Supply Chain of AI for Trustworthiness and Risk Management in Critical Applications
by: Sheh, Raymond K., et al.
Published: (2025)
by: Sheh, Raymond K., et al.
Published: (2025)
Rubric Is All You Need: Enhancing LLM-based Code Evaluation With Question-Specific Rubrics
by: Pathak, Aditya, et al.
Published: (2025)
by: Pathak, Aditya, et al.
Published: (2025)
LLM-Empowered Event-Chain Driven Code Generation for ADAS in SDV systems
by: Petrovic, Nenad, et al.
Published: (2025)
by: Petrovic, Nenad, et al.
Published: (2025)
Validating LLM-Generated Programs with Metamorphic Prompt Testing
by: Wang, Xiaoyin, et al.
Published: (2024)
by: Wang, Xiaoyin, et al.
Published: (2024)
Can LLM Generate Regression Tests for Software Commits?
by: Liu, Jing, et al.
Published: (2025)
by: Liu, Jing, et al.
Published: (2025)
Bias Testing and Mitigation in LLM-based Code Generation
by: Huang, Dong, et al.
Published: (2023)
by: Huang, Dong, et al.
Published: (2023)
DataGovBench: Benchmarking LLM Agents for Real-World Data Governance Workflows
by: Liu, Zhou, et al.
Published: (2025)
by: Liu, Zhou, et al.
Published: (2025)
Unveiling the Landscape of LLM Deployment in the Wild: An Empirical Study
by: Hou, Xinyi, et al.
Published: (2025)
by: Hou, Xinyi, et al.
Published: (2025)
Large Language Model Supply Chain: Open Problems From the Security Perspective
by: Hu, Qiang, et al.
Published: (2024)
by: Hu, Qiang, et al.
Published: (2024)
RESTestBench: A Benchmark for Evaluating the Effectiveness of LLM-Generated REST API Test Cases from NL Requirements
by: Kogler, Leon, et al.
Published: (2026)
by: Kogler, Leon, et al.
Published: (2026)
Rethinking Testing for LLM Applications: Characteristics, Challenges, and a Lightweight Interaction Protocol
by: Ma, Wei, et al.
Published: (2025)
by: Ma, Wei, et al.
Published: (2025)
LLM-Based Robustness Testing of Microservice Applications: An Empirical Study
by: Tigulla, Hrushitha Goud, et al.
Published: (2026)
by: Tigulla, Hrushitha Goud, et al.
Published: (2026)
LLM-Based Automated Diagnosis Of Integration Test Failures At Google
by: Ziftci, Celal, et al.
Published: (2026)
by: Ziftci, Celal, et al.
Published: (2026)
Evaluating LLM-Based Test Generation Under Software Evolution
by: Haroon, Sabaat, et al.
Published: (2026)
by: Haroon, Sabaat, et al.
Published: (2026)
LLM Test Generation via Iterative Hybrid Program Analysis
by: Gu, Sijia, et al.
Published: (2025)
by: Gu, Sijia, et al.
Published: (2025)
Similar Items
-
Feature-Centric Methodology for Analyzing Cross-Chain NFT Migration Compatibility
by: Chishti, Mohd Sameen, et al.
Published: (2026) -
AgentReputation: A Decentralized Agentic AI Reputation Framework
by: Chishti, Mohd Sameen, et al.
Published: (2026) -
A Proof of Success and Reward Distribution Protocol for Multi-bridge Architecture in Cross-chain Communication
by: Oyinloye, Damilare Peter, et al.
Published: (2025) -
Test It Before You Trust It: Applying Software Testing for Trustworthy In-context Learning
by: Racharak, Teeradaj, et al.
Published: (2025) -
Parameter-Efficient Fine-Tuning of Large Language Models for Unit Test Generation: An Empirical Study
by: Storhaug, André, et al.
Published: (2024)