Evaluating LLM Safety Under Repeated Inference via Accelerated Prompt Stress Testing
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Broadwater, Keita |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Evaluating Reliability Gaps in Large Language Model Safety via Repeated Prompt Sampling
von: Broadwater, Keita
Veröffentlicht: (2026)
von: Broadwater, Keita
Veröffentlicht: (2026)
Accelerating LLM Inference Throughput via Asynchronous KV Cache Prefetching
von: Dong, Yanhao, et al.
Veröffentlicht: (2025)
von: Dong, Yanhao, et al.
Veröffentlicht: (2025)
Recursive Speculative Decoding: Accelerating LLM Inference via Sampling Without Replacement
von: Jeon, Wonseok, et al.
Veröffentlicht: (2024)
von: Jeon, Wonseok, et al.
Veröffentlicht: (2024)
The Erasure Illusion: Stress-Testing the Generalization of LLM Forgetting Evaluation
von: Jia, Hengrui, et al.
Veröffentlicht: (2025)
von: Jia, Hengrui, et al.
Veröffentlicht: (2025)
Enhancing LLM Agent Safety via Causal Influence Prompting
von: Hahm, Dongyoon, et al.
Veröffentlicht: (2025)
von: Hahm, Dongyoon, et al.
Veröffentlicht: (2025)
CSAttention: Centroid-Scoring Attention for Accelerating LLM Inference
von: Song, Chuxu, et al.
Veröffentlicht: (2026)
von: Song, Chuxu, et al.
Veröffentlicht: (2026)
From Prompts to Power: Measuring the Energy Footprint of LLM Inference
von: Caravaca, Francisco, et al.
Veröffentlicht: (2025)
von: Caravaca, Francisco, et al.
Veröffentlicht: (2025)
MemFail: Stress-Testing Failure Modes of LLM Memory Systems
von: Garg, Ishir, et al.
Veröffentlicht: (2026)
von: Garg, Ishir, et al.
Veröffentlicht: (2026)
FastMTP: Accelerating LLM Inference with Enhanced Multi-Token Prediction
von: Cai, Yuxuan, et al.
Veröffentlicht: (2025)
von: Cai, Yuxuan, et al.
Veröffentlicht: (2025)
CodeScaler: Scaling Code LLM Training and Test-Time Inference via Reward Models
von: Zhu, Xiao, et al.
Veröffentlicht: (2026)
von: Zhu, Xiao, et al.
Veröffentlicht: (2026)
Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
von: Brown, Bradley, et al.
Veröffentlicht: (2024)
von: Brown, Bradley, et al.
Veröffentlicht: (2024)
Decocted Experience Improves Test-Time Inference in LLM Agents
von: Shen, Maohao, et al.
Veröffentlicht: (2026)
von: Shen, Maohao, et al.
Veröffentlicht: (2026)
Green Prompting: Characterizing Prompt-driven Energy Costs of LLM Inference
von: Adamska, Marta, et al.
Veröffentlicht: (2025)
von: Adamska, Marta, et al.
Veröffentlicht: (2025)
CATS: Cascaded Adaptive Tree Speculation for Memory-Limited LLM Inference Acceleration
von: Han, Yuning, et al.
Veröffentlicht: (2026)
von: Han, Yuning, et al.
Veröffentlicht: (2026)
Refinement Provenance Inference: Detecting LLM-Refined Training Prompts from Model Behavior
von: Yin, Bo, et al.
Veröffentlicht: (2026)
von: Yin, Bo, et al.
Veröffentlicht: (2026)
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator
von: Ramachandran, Akshat, et al.
Veröffentlicht: (2025)
von: Ramachandran, Akshat, et al.
Veröffentlicht: (2025)
Slow-Fast Inference: Training-Free Inference Acceleration via Within-Sentence Support Stability
von: Xie, Xingyu, et al.
Veröffentlicht: (2026)
von: Xie, Xingyu, et al.
Veröffentlicht: (2026)
Certifying LLM Safety against Adversarial Prompting
von: Kumar, Aounon, et al.
Veröffentlicht: (2023)
von: Kumar, Aounon, et al.
Veröffentlicht: (2023)
BarrierSteer: LLM Safety via Learning Barrier Steering
von: Tran, Thanh Q., et al.
Veröffentlicht: (2026)
von: Tran, Thanh Q., et al.
Veröffentlicht: (2026)
Robust Counterfactual Explanations under Model Multiplicity Using Multi-Objective Optimization
von: Kinjo, Keita
Veröffentlicht: (2025)
von: Kinjo, Keita
Veröffentlicht: (2025)
Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration
von: Wen, Zhuofan, et al.
Veröffentlicht: (2024)
von: Wen, Zhuofan, et al.
Veröffentlicht: (2024)
Accelerating LLM Inference with Lossless Speculative Decoding Algorithms for Heterogeneous Vocabularies
von: Timor, Nadav, et al.
Veröffentlicht: (2025)
von: Timor, Nadav, et al.
Veröffentlicht: (2025)
Cognitive Chunking for Soft Prompts: Accelerating Compressor Learning via Block-wise Causal Masking
von: Liu, Guojie, et al.
Veröffentlicht: (2026)
von: Liu, Guojie, et al.
Veröffentlicht: (2026)
AccLLM: Accelerating Long-Context LLM Inference Via Algorithm-Hardware Co-Design
von: Liang, Yanbiao, et al.
Veröffentlicht: (2025)
von: Liang, Yanbiao, et al.
Veröffentlicht: (2025)
Accelerating LLM Reasoning via Early Rejection with Partial Reward Modeling
von: Cheshmi, Seyyed Saeid, et al.
Veröffentlicht: (2025)
von: Cheshmi, Seyyed Saeid, et al.
Veröffentlicht: (2025)
Speculating Experts Accelerates Inference for Mixture-of-Experts
von: Madan, Vivan, et al.
Veröffentlicht: (2026)
von: Madan, Vivan, et al.
Veröffentlicht: (2026)
Evaluating Generalization and Representation Stability in Small LMs via Prompting, Fine-Tuning and Out-of-Distribution Prompts
von: Raja, Rahul, et al.
Veröffentlicht: (2025)
von: Raja, Rahul, et al.
Veröffentlicht: (2025)
Inference Optimization of Foundation Models on AI Accelerators
von: Park, Youngsuk, et al.
Veröffentlicht: (2024)
von: Park, Youngsuk, et al.
Veröffentlicht: (2024)
Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem
von: Lin, Shuyi, et al.
Veröffentlicht: (2025)
von: Lin, Shuyi, et al.
Veröffentlicht: (2025)
Accelerating Transformer Inference for Translation via Parallel Decoding
von: Santilli, Andrea, et al.
Veröffentlicht: (2023)
von: Santilli, Andrea, et al.
Veröffentlicht: (2023)
Accelerated AI Inference via Dynamic Execution Methods
von: Barad, Haim, et al.
Veröffentlicht: (2024)
von: Barad, Haim, et al.
Veröffentlicht: (2024)
PromptAudit: Auditing Prompt Sensitivity in LLM-Based Vulnerability Detection
von: Camarato, Steffen J., et al.
Veröffentlicht: (2026)
von: Camarato, Steffen J., et al.
Veröffentlicht: (2026)
Beyond Prompt-Induced Lies: Investigating LLM Deception on Benign Prompts
von: Wu, Zhaomin, et al.
Veröffentlicht: (2025)
von: Wu, Zhaomin, et al.
Veröffentlicht: (2025)
Auto-Prompt Ensemble for LLM Judge
von: Li, Jiajie, et al.
Veröffentlicht: (2025)
von: Li, Jiajie, et al.
Veröffentlicht: (2025)
AdaBlock-dLLM: Semantic-Aware Diffusion LLM Inference via Adaptive Block Size
von: Lu, Guanxi, et al.
Veröffentlicht: (2025)
von: Lu, Guanxi, et al.
Veröffentlicht: (2025)
RelayCaching: Accelerating LLM Collaboration via Decoding KV Cache Reuse
von: Geng, Yingsheng, et al.
Veröffentlicht: (2026)
von: Geng, Yingsheng, et al.
Veröffentlicht: (2026)
Thermodynamic Focusing for Inference-Time Search: Practical Methods for Target-Conditioned Sampling and Prompted Inference
von: Zhang, Zhan
Veröffentlicht: (2025)
von: Zhang, Zhan
Veröffentlicht: (2025)
Test-Time Training Undermines Safety Guardrails
von: Antonelli, Simone, et al.
Veröffentlicht: (2026)
von: Antonelli, Simone, et al.
Veröffentlicht: (2026)
Prompting Test-Time Scaling Is A Strong LLM Reasoning Data Augmentation
von: Bsharat, Sondos Mahmoud, et al.
Veröffentlicht: (2025)
von: Bsharat, Sondos Mahmoud, et al.
Veröffentlicht: (2025)
OpInf-LLM: Parametric PDE Solving with LLMs via Operator Inference
von: Wang, Zhuoyuan, et al.
Veröffentlicht: (2026)
von: Wang, Zhuoyuan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Evaluating Reliability Gaps in Large Language Model Safety via Repeated Prompt Sampling
von: Broadwater, Keita
Veröffentlicht: (2026) -
Accelerating LLM Inference Throughput via Asynchronous KV Cache Prefetching
von: Dong, Yanhao, et al.
Veröffentlicht: (2025) -
Recursive Speculative Decoding: Accelerating LLM Inference via Sampling Without Replacement
von: Jeon, Wonseok, et al.
Veröffentlicht: (2024) -
The Erasure Illusion: Stress-Testing the Generalization of LLM Forgetting Evaluation
von: Jia, Hengrui, et al.
Veröffentlicht: (2025) -
Enhancing LLM Agent Safety via Causal Influence Prompting
von: Hahm, Dongyoon, et al.
Veröffentlicht: (2025)