Get my drift? Catching LLM Task Drift with Activation Deltas
Fuente:
arXiv
Guardado en:
| Autores principales: | Abdelnabi, Sahar, Fay, Aideen, Cherubin, Giovanni, Salem, Ahmed, Fritz, Mario, Paverd, Andrew |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations
por: Gomaa, Amr, et al.
Publicado: (2025)
por: Gomaa, Amr, et al.
Publicado: (2025)
Stateless Yet Not Forgetful: Implicit Memory as a Hidden Channel in LLMs
por: Salem, Ahmed, et al.
Publicado: (2026)
por: Salem, Ahmed, et al.
Publicado: (2026)
AI Agents May Always Fall for Prompt Injections
por: Abdelnabi, Sahar, et al.
Publicado: (2026)
por: Abdelnabi, Sahar, et al.
Publicado: (2026)
Firewalls to Secure Dynamic LLM Agentic Networks
por: Abdelnabi, Sahar, et al.
Publicado: (2025)
por: Abdelnabi, Sahar, et al.
Publicado: (2025)
MetaBackdoor: Exploiting Positional Encoding as a Backdoor Attack Surface in LLMs
por: Wen, Rui, et al.
Publicado: (2026)
por: Wen, Rui, et al.
Publicado: (2026)
The Hawthorne Effect in Reasoning Models: Evaluating and Steering Test Awareness
por: Abdelnabi, Sahar, et al.
Publicado: (2025)
por: Abdelnabi, Sahar, et al.
Publicado: (2025)
Read This Paper to Get $50 Million:* An Analysis of Mobile Messaging Scams Using Reddit Data
por: Lu, Allison, et al.
Publicado: (2026)
por: Lu, Allison, et al.
Publicado: (2026)
Getting Bored of Cyberwar: Exploring the Role of Low-level Cybercrime Actors in the Russia-Ukraine Conflict
por: Vu, Anh V., et al.
Publicado: (2022)
por: Vu, Anh V., et al.
Publicado: (2022)
Assessing The Effectiveness Of Current Cybersecurity Regulations And Policies In The US
por: Oluomachi, Ejiofor, et al.
Publicado: (2024)
por: Oluomachi, Ejiofor, et al.
Publicado: (2024)
Closed-Form Bounds for DP-SGD against Record-level Inference
por: Cherubin, Giovanni, et al.
Publicado: (2024)
por: Cherubin, Giovanni, et al.
Publicado: (2024)
From misinformation to climate crisis: Navigating vulnerabilities in the cyber-physical-social systems
por: Aamir, Tooba, et al.
Publicado: (2025)
por: Aamir, Tooba, et al.
Publicado: (2025)
Democratizing Federated Learning with Blockchain and Multi-Task Peer Prediction
por: Witt, Leon, et al.
Publicado: (2026)
por: Witt, Leon, et al.
Publicado: (2026)
Defending Against Intelligent Attackers at Large Scales
por: Lohn, Andrew J.
Publicado: (2025)
por: Lohn, Andrew J.
Publicado: (2025)
No More Trade-Offs. GPT and Fully Informative Privacy Policies
por: Pałka, Przemysław, et al.
Publicado: (2023)
por: Pałka, Przemysław, et al.
Publicado: (2023)
When Your Reviewer is an LLM: Biases, Divergence, and Prompt Injection Risks in Peer Review
por: Zhu, Changjia, et al.
Publicado: (2025)
por: Zhu, Changjia, et al.
Publicado: (2025)
Identifying Security Risks in NFT Platforms
por: Gupta, Yash, et al.
Publicado: (2022)
por: Gupta, Yash, et al.
Publicado: (2022)
An Internet Voting System Fatally Flawed in Creative New Ways
por: Appel, Andrew W., et al.
Publicado: (2024)
por: Appel, Andrew W., et al.
Publicado: (2024)
Integrating Generative AI into Cybersecurity Education: A Study of OCR and Multimodal LLM-assisted Instruction
por: Patel, Karan, et al.
Publicado: (2025)
por: Patel, Karan, et al.
Publicado: (2025)
Honeyquest: Rapidly Measuring the Enticingness of Cyber Deception Techniques with Code-based Questionnaires
por: Kahlhofer, Mario, et al.
Publicado: (2024)
por: Kahlhofer, Mario, et al.
Publicado: (2024)
Guardians of the Regime: Secret Police Formation in Autocracies
por: Mehrl, Marius, et al.
Publicado: (2025)
por: Mehrl, Marius, et al.
Publicado: (2025)
Data After Death: Australian User Preferences and Future Solutions to Protect Posthumous User Data
por: Reeves, Andrew, et al.
Publicado: (2024)
por: Reeves, Andrew, et al.
Publicado: (2024)
Understanding Cyber Threats Against the Universities, Colleges, and Schools
por: Lallie, Harjinder Singh, et al.
Publicado: (2023)
por: Lallie, Harjinder Singh, et al.
Publicado: (2023)
Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" Control
por: Cyberey, Hannah, et al.
Publicado: (2025)
por: Cyberey, Hannah, et al.
Publicado: (2025)
SecureForge: Finding and Preventing Vulnerabilities in LLM-Generated Code via Prompt Optimization
por: Liu, Houjun, et al.
Publicado: (2026)
por: Liu, Houjun, et al.
Publicado: (2026)
LLM Novice Uplift on Dual-Use, In Silico Biology Tasks
por: Zhang, Chen Bo Calvin, et al.
Publicado: (2026)
por: Zhang, Chen Bo Calvin, et al.
Publicado: (2026)
Beyond Privacy Trade-offs with Structured Transparency
por: Trask, Andrew, et al.
Publicado: (2020)
por: Trask, Andrew, et al.
Publicado: (2020)
Who Gets Flagged? The Pluralistic Evaluation Gap in AI Content Watermarking
por: Nemecek, Alexander, et al.
Publicado: (2026)
por: Nemecek, Alexander, et al.
Publicado: (2026)
LLMail-Inject: A Dataset from a Realistic Adaptive Prompt Injection Challenge
por: Abdelnabi, Sahar, et al.
Publicado: (2025)
por: Abdelnabi, Sahar, et al.
Publicado: (2025)
Learning from Mistakes: Can LLM Self-Recover after Misalignment?
por: Sorokoletova, Olga E., et al.
Publicado: (2026)
por: Sorokoletova, Olga E., et al.
Publicado: (2026)
"bot lane noob" Towards Deployment of NLP-based Toxicity Detectors in Video Games
por: Ave, Jonas, et al.
Publicado: (2026)
por: Ave, Jonas, et al.
Publicado: (2026)
Detecting Verbatim LLM Copy-Paste in Homework
por: Aiersilan, Aizierjiang
Publicado: (2026)
por: Aiersilan, Aizierjiang
Publicado: (2026)
Highlight & Summarize: RAG without the jailbreaks
por: Cherubin, Giovanni, et al.
Publicado: (2025)
por: Cherubin, Giovanni, et al.
Publicado: (2025)
Ethical Challenges in Computer Vision: Ensuring Privacy and Mitigating Bias in Publicly Available Datasets
por: Tahir, Ghalib Ahmed
Publicado: (2024)
por: Tahir, Ghalib Ahmed
Publicado: (2024)
The Impact of AI on the Cyber Offense-Defense Balance and the Character of Cyber Conflict
por: Lohn, Andrew J.
Publicado: (2025)
por: Lohn, Andrew J.
Publicado: (2025)
GPT, Ontology, and CAABAC: A Tripartite Personalized Access Control Model Anchored by Compliance, Context and Attribute
por: Nowrozy, Raza, et al.
Publicado: (2024)
por: Nowrozy, Raza, et al.
Publicado: (2024)
Consumer Beware! Exploring Data Brokers' CCPA Compliance
por: van Kempen, Elina, et al.
Publicado: (2025)
por: van Kempen, Elina, et al.
Publicado: (2025)
On the Security and Privacy of AI-based Mobile Health Chatbots
por: Wairimu, Samuel, et al.
Publicado: (2025)
por: Wairimu, Samuel, et al.
Publicado: (2025)
Human-Centered Threat Modeling in Practice: Lessons, Challenges, and Paths Forward
por: Usman, Warda, et al.
Publicado: (2025)
por: Usman, Warda, et al.
Publicado: (2025)
Prompt Injection Vulnerability of Consensus Generating Applications in Digital Democracy
por: Gudiño-Rosero, Jairo, et al.
Publicado: (2025)
por: Gudiño-Rosero, Jairo, et al.
Publicado: (2025)
"Hello, is this Anna?": Unpacking the Lifecycle of Pig-Butchering Scams
por: Oak, Rajvardhan, et al.
Publicado: (2025)
por: Oak, Rajvardhan, et al.
Publicado: (2025)
Ejemplares similares
-
ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations
por: Gomaa, Amr, et al.
Publicado: (2025) -
Stateless Yet Not Forgetful: Implicit Memory as a Hidden Channel in LLMs
por: Salem, Ahmed, et al.
Publicado: (2026) -
AI Agents May Always Fall for Prompt Injections
por: Abdelnabi, Sahar, et al.
Publicado: (2026) -
Firewalls to Secure Dynamic LLM Agentic Networks
por: Abdelnabi, Sahar, et al.
Publicado: (2025) -
MetaBackdoor: Exploiting Positional Encoding as a Backdoor Attack Surface in LLMs
por: Wen, Rui, et al.
Publicado: (2026)