Saved in:
| Main Authors: | McKenzie, Ian R., Hollinsworth, Oskar J., Tseng, Tom, Davies, Xander, Casper, Stephen, Tucker, Aaron D., Kirk, Robert, Gleave, Adam |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2506.24068 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scaling Trends in Language Model Robustness
by: Howe, Nikolaus, et al.
Published: (2024)
by: Howe, Nikolaus, et al.
Published: (2024)
Exposing the Systematic Vulnerability of Open-Weight Models to Prefill Attacks
by: Struppek, Lukas, et al.
Published: (2026)
by: Struppek, Lukas, et al.
Published: (2026)
An Example Safety Case for Safeguards Against Misuse
by: Clymer, Joshua, et al.
Published: (2025)
by: Clymer, Joshua, et al.
Published: (2025)
Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs
by: O'Brien, Kyle, et al.
Published: (2025)
by: O'Brien, Kyle, et al.
Published: (2025)
Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility
by: Murphy, Brendan, et al.
Published: (2025)
by: Murphy, Brendan, et al.
Published: (2025)
Circuit Breaking: Removing Model Behaviors with Targeted Ablation
by: Li, Maximilian, et al.
Published: (2023)
by: Li, Maximilian, et al.
Published: (2023)
Can Go AIs be adversarially robust?
by: Tseng, Tom, et al.
Published: (2024)
by: Tseng, Tom, et al.
Published: (2024)
The Art of Photosketching and Its Applications in the School.
by: McKenzie, Barbara K.
Published: (1993)
by: McKenzie, Barbara K.
Published: (1993)
Dual Quadrature Phasemeter for Space-Based Interferometry
by: Sambridge, Callum S., et al.
Published: (2024)
by: Sambridge, Callum S., et al.
Published: (2024)
Existing Large Language Model Unlearning Evaluations Are Inconclusive
by: Feng, Zhili, et al.
Published: (2025)
by: Feng, Zhili, et al.
Published: (2025)
On the Robustness of Verbal Confidence of LLMs in Adversarial Attacks
by: Obadinma, Stephen, et al.
Published: (2025)
by: Obadinma, Stephen, et al.
Published: (2025)
An exploration of UK speech and language therapists' treatment and management of functional communication disorders: A mixed‐methods online survey
by: Kirsty McKenzie, et al.
Published: (2024)
by: Kirsty McKenzie, et al.
Published: (2024)
Arabic Dataset for LLM Safeguard Evaluation
by: Ashraf, Yasser, et al.
Published: (2024)
by: Ashraf, Yasser, et al.
Published: (2024)
Exploiting Novel GPT-4 APIs
by: Pelrine, Kellin, et al.
Published: (2023)
by: Pelrine, Kellin, et al.
Published: (2023)
What Features in Prompts Jailbreak LLMs? Investigating the Mechanisms Behind Attacks
by: Kirch, Nathalie, et al.
Published: (2024)
by: Kirch, Nathalie, et al.
Published: (2024)
Concept Influence: Leveraging Interpretability to Improve Performance and Efficiency in Training Data Attribution
by: Kowal, Matthew, et al.
Published: (2026)
by: Kowal, Matthew, et al.
Published: (2026)
A Multilingual, Large-Scale Study of the Interplay between LLM Safeguards, Personalisation, and Disinformation
by: Leite, João A., et al.
Published: (2025)
by: Leite, João A., et al.
Published: (2025)
A Constraint-Enforcing Reward for Adversarial Attacks on Text Classifiers
by: Roth, Tom, et al.
Published: (2024)
by: Roth, Tom, et al.
Published: (2024)
Can Federated Learning Safeguard Private Data in LLM Training? Vulnerabilities, Attacks, and Defense Evaluation
by: Guo, Wenkai, et al.
Published: (2025)
by: Guo, Wenkai, et al.
Published: (2025)
Self-Guard: Empower the LLM to Safeguard Itself
by: Wang, Zezhong, et al.
Published: (2023)
by: Wang, Zezhong, et al.
Published: (2023)
Evaluating and Safeguarding the Adversarial Robustness of Retrieval-Based In-Context Learning
by: Yu, Simon, et al.
Published: (2024)
by: Yu, Simon, et al.
Published: (2024)
A Generative Adversarial Attack for Multilingual Text Classifiers
by: Roth, Tom, et al.
Published: (2024)
by: Roth, Tom, et al.
Published: (2024)
Path Channels and Plan Extension Kernels: a Mechanistic Description of Planning in a Sokoban RNN
by: Taufeeque, Mohammad, et al.
Published: (2025)
by: Taufeeque, Mohammad, et al.
Published: (2025)
Paraphrasing Adversarial Attack on LLM-as-a-Reviewer
by: Kaneko, Masahiro
Published: (2026)
by: Kaneko, Masahiro
Published: (2026)
GLiGuard: Schema-Conditioned Classification for LLM Safeguard
by: Zaratiana, Urchade, et al.
Published: (2026)
by: Zaratiana, Urchade, et al.
Published: (2026)
Uncovering Latent Human Wellbeing in Language Model Embeddings
by: Freire, Pedro, et al.
Published: (2024)
by: Freire, Pedro, et al.
Published: (2024)
Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment
by: Raina, Vyas, et al.
Published: (2024)
by: Raina, Vyas, et al.
Published: (2024)
Token-Modification Adversarial Attacks for Natural Language Processing: A Survey
by: Roth, Tom, et al.
Published: (2021)
by: Roth, Tom, et al.
Published: (2021)
LLM-Based Adversarial Persuasion Attacks on Fact-Checking Systems
by: Leite, João A., et al.
Published: (2026)
by: Leite, João A., et al.
Published: (2026)
Safeguarding Privacy of Retrieval Data against Membership Inference Attacks: Is This Query Too Close to Home?
by: Choi, Yujin, et al.
Published: (2025)
by: Choi, Yujin, et al.
Published: (2025)
Safeguarding RAG Pipelines with GMTP: A Gradient-based Masked Token Probability Method for Poisoned Document Detection
by: Kim, San, et al.
Published: (2025)
by: Kim, San, et al.
Published: (2025)
Attacking Misinformation Detection Using Adversarial Examples Generated by Language Models
by: Przybyła, Piotr, et al.
Published: (2024)
by: Przybyła, Piotr, et al.
Published: (2024)
Generic Reduction-Based Interpreters (Extended Version)
by: Bach, Casper
Published: (2025)
by: Bach, Casper
Published: (2025)
ONE YUAN STACK
by: vojka93
Published: (2021)
by: vojka93
Published: (2021)
Investigating Non-Transitivity in LLM-as-a-Judge
by: Xu, Yi, et al.
Published: (2025)
by: Xu, Yi, et al.
Published: (2025)
Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation
by: Halawi, Danny, et al.
Published: (2024)
by: Halawi, Danny, et al.
Published: (2024)
Black-Box Adversarial Attacks on LLM-Based Code Completion
by: Jenko, Slobodan, et al.
Published: (2024)
by: Jenko, Slobodan, et al.
Published: (2024)
REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations
by: Liang, Buyun, et al.
Published: (2026)
by: Liang, Buyun, et al.
Published: (2026)
Adversarial Attacks on LLM-as-a-Judge Systems: Insights from Prompt Injections
by: Maloyan, Narek, et al.
Published: (2025)
by: Maloyan, Narek, et al.
Published: (2025)
Propensity Inference: Environmental Contributors to LLM Behaviour
by: Järviniemi, Olli, et al.
Published: (2026)
by: Järviniemi, Olli, et al.
Published: (2026)
Similar Items
-
Scaling Trends in Language Model Robustness
by: Howe, Nikolaus, et al.
Published: (2024) -
Exposing the Systematic Vulnerability of Open-Weight Models to Prefill Attacks
by: Struppek, Lukas, et al.
Published: (2026) -
An Example Safety Case for Safeguards Against Misuse
by: Clymer, Joshua, et al.
Published: (2025) -
Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs
by: O'Brien, Kyle, et al.
Published: (2025) -
Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility
by: Murphy, Brendan, et al.
Published: (2025)