Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Gringras, David |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Mind the GAP: Text Safety Does Not Transfer to Tool-Call Safety in LLM Agents
von: Cartagena, Arnold, et al.
Veröffentlicht: (2026)
von: Cartagena, Arnold, et al.
Veröffentlicht: (2026)
Tool-Schema Compression Enables Agentic RAG Under Constrained Context Budgets
von: Sakizli, Furkan
Veröffentlicht: (2026)
von: Sakizli, Furkan
Veröffentlicht: (2026)
Improving Existing Optimization Algorithms with LLMs
von: Sartori, Camilo Chacón, et al.
Veröffentlicht: (2025)
von: Sartori, Camilo Chacón, et al.
Veröffentlicht: (2025)
Vibe Code Bench: Evaluating AI Models on End-to-End Web Application Development
von: Tran, Hung, et al.
Veröffentlicht: (2026)
von: Tran, Hung, et al.
Veröffentlicht: (2026)
Engineering A Large Language Model From Scratch
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2024)
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2024)
Failure by Interference: Language Models Make Balanced Parentheses Errors When Faulty Mechanisms Overshadow Sound Ones
von: Rai, Daking, et al.
Veröffentlicht: (2025)
von: Rai, Daking, et al.
Veröffentlicht: (2025)
Automated Web Application Testing: End-to-End Test Case Generation with Large Language Models and Screen Transition Graphs
von: Le, Nguyen-Khang, et al.
Veröffentlicht: (2025)
von: Le, Nguyen-Khang, et al.
Veröffentlicht: (2025)
Narrow Transformer: StarCoder-Based Java-LM For Desktop
von: Rathinasamy, Kamalkumar, et al.
Veröffentlicht: (2024)
von: Rathinasamy, Kamalkumar, et al.
Veröffentlicht: (2024)
Mechanistic Understanding of Language Models in Syntactic Code Completion
von: Miller, Samuel, et al.
Veröffentlicht: (2025)
von: Miller, Samuel, et al.
Veröffentlicht: (2025)
Smaller Models, Smarter Rewards: A Two-Sided Approach to Process and Outcome Rewards
von: Groeneveld, Jan Niklas, et al.
Veröffentlicht: (2025)
von: Groeneveld, Jan Niklas, et al.
Veröffentlicht: (2025)
AgentAtlas: Beyond Outcome Leaderboards for LLM Agents
von: Mazaheri, Parsa, et al.
Veröffentlicht: (2026)
von: Mazaheri, Parsa, et al.
Veröffentlicht: (2026)
The Path Not Taken: Duality in Reasoning about Program Execution
von: Hasanov, Eshgin, et al.
Veröffentlicht: (2026)
von: Hasanov, Eshgin, et al.
Veröffentlicht: (2026)
LLMORPH: Automated Metamorphic Testing of Large Language Models
von: Cho, Steven, et al.
Veröffentlicht: (2026)
von: Cho, Steven, et al.
Veröffentlicht: (2026)
Plan with Code: Comparing approaches for robust NL to DSL generation
von: Bassamzadeh, Nastaran, et al.
Veröffentlicht: (2024)
von: Bassamzadeh, Nastaran, et al.
Veröffentlicht: (2024)
A Comparative Study of DSL Code Generation: Fine-Tuning vs. Optimized Retrieval Augmentation
von: Bassamzadeh, Nastaran, et al.
Veröffentlicht: (2024)
von: Bassamzadeh, Nastaran, et al.
Veröffentlicht: (2024)
How Generation Architecture Shapes Code Complexity in Multi-Agent LLM Systems: A Paired Study on HumanEval
von: Ashrafi, Nazmus
Veröffentlicht: (2026)
von: Ashrafi, Nazmus
Veröffentlicht: (2026)
Evaluating the Limitations of Local LLMs in Solving Complex Programming Challenges
von: Matotek, Kadin, et al.
Veröffentlicht: (2025)
von: Matotek, Kadin, et al.
Veröffentlicht: (2025)
Simple and Effective Baselines for Code Summarisation Evaluation
von: Robinson, Jade, et al.
Veröffentlicht: (2025)
von: Robinson, Jade, et al.
Veröffentlicht: (2025)
REPOT: Recoverable Program-of-Thought via Checkpoint Repair
von: Mazaheri, Parsa
Veröffentlicht: (2026)
von: Mazaheri, Parsa
Veröffentlicht: (2026)
The Instability of Safety: How Random Seeds and Temperature Expose Inconsistent LLM Refusal Behavior
von: Larsen, Erik
Veröffentlicht: (2025)
von: Larsen, Erik
Veröffentlicht: (2025)
VeriGuard: Enhancing LLM Agent Safety via Verified Code Generation
von: Miculicich, Lesly, et al.
Veröffentlicht: (2025)
von: Miculicich, Lesly, et al.
Veröffentlicht: (2025)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2023)
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2023)
When Retrieval Hurts Code Completion: A Diagnostic Study of Stale Repository Context
von: Weng, Haojun, et al.
Veröffentlicht: (2026)
von: Weng, Haojun, et al.
Veröffentlicht: (2026)
GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
von: Agrawal, Lakshya A, et al.
Veröffentlicht: (2025)
von: Agrawal, Lakshya A, et al.
Veröffentlicht: (2025)
Comparative Analysis of LLM Abliteration Methods: A Cross-Architecture Evaluation
von: Young, Richard J.
Veröffentlicht: (2025)
von: Young, Richard J.
Veröffentlicht: (2025)
Learning Software Bug Reports: A Systematic Literature Review
von: Long, Guoming, et al.
Veröffentlicht: (2025)
von: Long, Guoming, et al.
Veröffentlicht: (2025)
Improved IR-based Bug Localization with Intelligent Relevance Feedback
von: Samir, Asif Mohammed, et al.
Veröffentlicht: (2025)
von: Samir, Asif Mohammed, et al.
Veröffentlicht: (2025)
Is It Time To Treat Prompts As Code? A Multi-Use Case Study For Prompt Optimization Using DSPy
von: Lemos, Francisca, et al.
Veröffentlicht: (2025)
von: Lemos, Francisca, et al.
Veröffentlicht: (2025)
CWM: An Open-Weights LLM for Research on Code Generation with World Models
von: FAIR CodeGen team, et al.
Veröffentlicht: (2025)
von: FAIR CodeGen team, et al.
Veröffentlicht: (2025)
Making a Pipeline Production-Ready: Challenges and Lessons Learned in the Healthcare Domain
von: Lawand, Daniel Angelo Esteves, et al.
Veröffentlicht: (2025)
von: Lawand, Daniel Angelo Esteves, et al.
Veröffentlicht: (2025)
SPIRA: Building an Intelligent System for Respiratory Insufficiency Detection
von: Ferreira, Renato Cordeiro, et al.
Veröffentlicht: (2025)
von: Ferreira, Renato Cordeiro, et al.
Veröffentlicht: (2025)
Prompt Engineering Strategies for LLM-based Qualitative Coding of Psychological Safety in Software Engineering Communities: A Controlled Empirical Study
von: Alshaikh, Moaath, et al.
Veröffentlicht: (2026)
von: Alshaikh, Moaath, et al.
Veröffentlicht: (2026)
AcTracer: Active Testing of Large Language Model via Multi-Stage Sampling
von: Huang, Yuheng, et al.
Veröffentlicht: (2024)
von: Huang, Yuheng, et al.
Veröffentlicht: (2024)
Feedback-Normalized Developer Memory for Reinforcement-Learning Coding Agents: A Safety-Gated MCP Architecture
von: Iscan, Mehmet
Veröffentlicht: (2026)
von: Iscan, Mehmet
Veröffentlicht: (2026)
AgentPulse: A Continuous Multi-Signal Framework for Evaluating AI Agents in Deployment
von: Gao, Yuxuan, et al.
Veröffentlicht: (2026)
von: Gao, Yuxuan, et al.
Veröffentlicht: (2026)
A Framework for Testing and Adapting REST APIs as LLM Tools
von: Bandlamudi, Jayachandu, et al.
Veröffentlicht: (2025)
von: Bandlamudi, Jayachandu, et al.
Veröffentlicht: (2025)
Finetuning LLMs for Automatic Form Interaction on Web-Browser in Selenium Testing Framework
von: Le, Nguyen-Khang, et al.
Veröffentlicht: (2025)
von: Le, Nguyen-Khang, et al.
Veröffentlicht: (2025)
MISR: Measuring Instrumental Self-Reasoning in Frontier Models
von: Fronsdal, Kai, et al.
Veröffentlicht: (2024)
von: Fronsdal, Kai, et al.
Veröffentlicht: (2024)
Comprehensive Evaluation and Insights into the Use of Large Language Models in the Automation of Behavior-Driven Development Acceptance Test Formulation
von: Karpurapu, Shanthi, et al.
Veröffentlicht: (2024)
von: Karpurapu, Shanthi, et al.
Veröffentlicht: (2024)
A Multi-Perspective Benchmark and Moderation Model for Evaluating Safety and Adversarial Robustness
von: Machlovi, Naseem, et al.
Veröffentlicht: (2025)
von: Machlovi, Naseem, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Mind the GAP: Text Safety Does Not Transfer to Tool-Call Safety in LLM Agents
von: Cartagena, Arnold, et al.
Veröffentlicht: (2026) -
Tool-Schema Compression Enables Agentic RAG Under Constrained Context Budgets
von: Sakizli, Furkan
Veröffentlicht: (2026) -
Improving Existing Optimization Algorithms with LLMs
von: Sartori, Camilo Chacón, et al.
Veröffentlicht: (2025) -
Vibe Code Bench: Evaluating AI Models on End-to-End Web Application Development
von: Tran, Hung, et al.
Veröffentlicht: (2026) -
Engineering A Large Language Model From Scratch
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2024)