Trading Inference-Time Compute for Adversarial Robustness
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zaremba, Wojciech, Nitishinskaya, Evgenia, Barak, Boaz, Lin, Stephanie, Toyer, Sam, Yu, Yaodong, Dias, Rachel, Wallace, Eric, Xiao, Kai, Heidecke, Johannes, Glaese, Amelia |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Deliberative Alignment: Reasoning Enables Safer Language Models
von: Guan, Melody Y., et al.
Veröffentlicht: (2024)
von: Guan, Melody Y., et al.
Veröffentlicht: (2024)
The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
von: Wallace, Eric, et al.
Veröffentlicht: (2024)
von: Wallace, Eric, et al.
Veröffentlicht: (2024)
The Inherent Adversarial Robustness of Analog In-Memory Computing
von: Lammie, Corey, et al.
Veröffentlicht: (2024)
von: Lammie, Corey, et al.
Veröffentlicht: (2024)
AQUA-LLM: Evaluating Accuracy, Quantization, and Adversarial Robustness Trade-offs in LLMs for Cybersecurity Question Answering
von: Gungor, Onat, et al.
Veröffentlicht: (2025)
von: Gungor, Onat, et al.
Veröffentlicht: (2025)
A Computational Separation Between Quantum No-cloning and No-telegraphing
von: Nehoran, Barak, et al.
Veröffentlicht: (2023)
von: Nehoran, Barak, et al.
Veröffentlicht: (2023)
Stress Testing Deliberative Alignment for Anti-Scheming Training
von: Schoen, Bronson, et al.
Veröffentlicht: (2025)
von: Schoen, Bronson, et al.
Veröffentlicht: (2025)
IH-Challenge: A Training Dataset to Improve Instruction Hierarchy on Frontier LLMs
von: Guo, Chuan, et al.
Veröffentlicht: (2026)
von: Guo, Chuan, et al.
Veröffentlicht: (2026)
Privacy Leaks by Adversaries: Adversarial Iterations for Membership Inference Attack
von: Xue, Jing, et al.
Veröffentlicht: (2025)
von: Xue, Jing, et al.
Veröffentlicht: (2025)
Watermarks in the Sand: Impossibility of Strong Watermarking for Generative Models
von: Zhang, Hanlin, et al.
Veröffentlicht: (2023)
von: Zhang, Hanlin, et al.
Veröffentlicht: (2023)
Training LLMs for Honesty via Confessions
von: Joglekar, Manas, et al.
Veröffentlicht: (2025)
von: Joglekar, Manas, et al.
Veröffentlicht: (2025)
RAT: Adversarial Attacks on Deep Reinforcement Agents for Targeted Behaviors
von: Bai, Fengshuo, et al.
Veröffentlicht: (2024)
von: Bai, Fengshuo, et al.
Veröffentlicht: (2024)
Adversarial Robustness of Link Sign Prediction in Signed Graphs
von: Zhou, Jialong, et al.
Veröffentlicht: (2024)
von: Zhou, Jialong, et al.
Veröffentlicht: (2024)
NeuroTrace: Inference Provenance-Based Detection of Adversarial Examples
von: Hmida, Firas Ben, et al.
Veröffentlicht: (2026)
von: Hmida, Firas Ben, et al.
Veröffentlicht: (2026)
Adversarially Robust and Interpretable Magecart Malware Detection
von: Pereira, Pedro, et al.
Veröffentlicht: (2025)
von: Pereira, Pedro, et al.
Veröffentlicht: (2025)
Safety Interventions against Adversarial Patches in an Open-Source Driver Assistance System
von: Chen, Cheng, et al.
Veröffentlicht: (2025)
von: Chen, Cheng, et al.
Veröffentlicht: (2025)
Practical Type Inference: High-Throughput Recovery of Real-World Structures and Function Signatures
von: Seidel, Lukas, et al.
Veröffentlicht: (2026)
von: Seidel, Lukas, et al.
Veröffentlicht: (2026)
Adversarial Machine Learning for Robust Password Strength Estimation
von: Jha, Pappu, et al.
Veröffentlicht: (2025)
von: Jha, Pappu, et al.
Veröffentlicht: (2025)
Robustness Against Adversarial Attacks via Learning Confined Adversarial Polytopes
von: Hamidi, Shayan Mohajer, et al.
Veröffentlicht: (2024)
von: Hamidi, Shayan Mohajer, et al.
Veröffentlicht: (2024)
An Adversarial Robust Behavior Sequence Anomaly Detection Approach Based on Critical Behavior Unit Learning
von: Zhan, Dongyang, et al.
Veröffentlicht: (2025)
von: Zhan, Dongyang, et al.
Veröffentlicht: (2025)
A StrongREJECT for Empty Jailbreaks
von: Souly, Alexandra, et al.
Veröffentlicht: (2024)
von: Souly, Alexandra, et al.
Veröffentlicht: (2024)
Adversarially Robust Assembly Language Model for Packed Executables Detection
von: Li, Shijia, et al.
Veröffentlicht: (2025)
von: Li, Shijia, et al.
Veröffentlicht: (2025)
On the Robustness of Malware Detectors to Adversarial Samples
von: Salman, Muhammad, et al.
Veröffentlicht: (2024)
von: Salman, Muhammad, et al.
Veröffentlicht: (2024)
Are Robust LLM Fingerprints Adversarially Robust?
von: Nasery, Anshul, et al.
Veröffentlicht: (2025)
von: Nasery, Anshul, et al.
Veröffentlicht: (2025)
Adversarial News and Lost Profits: Manipulating Headlines in LLM-Driven Algorithmic Trading
von: Rizvani, Advije, et al.
Veröffentlicht: (2026)
von: Rizvani, Advije, et al.
Veröffentlicht: (2026)
Towards Privacy-Preserving Split Learning: Destabilizing Adversarial Inference and Reconstruction Attacks in the Cloud
von: Higgins, Griffin, et al.
Veröffentlicht: (2025)
von: Higgins, Griffin, et al.
Veröffentlicht: (2025)
SoK: Critical Evaluation of Quantum Machine Learning for Adversarial Robustness
von: Nowmi, Saeefa Rubaiyet, et al.
Veröffentlicht: (2025)
von: Nowmi, Saeefa Rubaiyet, et al.
Veröffentlicht: (2025)
Adversarial Example Based Fingerprinting for Robust Copyright Protection in Split Learning
von: Lin, Zhangting, et al.
Veröffentlicht: (2025)
von: Lin, Zhangting, et al.
Veröffentlicht: (2025)
Efficient Adversarial Detection Frameworks for Vehicle-to-Microgrid Services in Edge Computing
von: Omara, Ahmed, et al.
Veröffentlicht: (2025)
von: Omara, Ahmed, et al.
Veröffentlicht: (2025)
Passive Inference Attacks on Split Learning via Adversarial Regularization
von: Zhu, Xiaochen, et al.
Veröffentlicht: (2023)
von: Zhu, Xiaochen, et al.
Veröffentlicht: (2023)
TensorCommitments: A Lightweight Verifiable Inference for Language Models
von: Baser, Oguzhan, et al.
Veröffentlicht: (2026)
von: Baser, Oguzhan, et al.
Veröffentlicht: (2026)
Adversarial Robustness through Dynamic Ensemble Learning
von: Waghela, Hetvi, et al.
Veröffentlicht: (2024)
von: Waghela, Hetvi, et al.
Veröffentlicht: (2024)
Vulnerability-Aware Robust Multimodal Adversarial Training
von: Zhang, Junrui, et al.
Veröffentlicht: (2025)
von: Zhang, Junrui, et al.
Veröffentlicht: (2025)
Quantum-Enhanced Adversarial Robustness in Artificial Intelligence
von: Sen, Jaydip
Veröffentlicht: (2026)
von: Sen, Jaydip
Veröffentlicht: (2026)
The Surprising Harmfulness of Benign Overfitting for Adversarial Robustness
von: Hao, Yifan, et al.
Veröffentlicht: (2024)
von: Hao, Yifan, et al.
Veröffentlicht: (2024)
IDEA: Invariant Defense for Graph Adversarial Robustness
von: Tao, Shuchang, et al.
Veröffentlicht: (2023)
von: Tao, Shuchang, et al.
Veröffentlicht: (2023)
PatchCURE: Improving Certifiable Robustness, Model Utility, and Computation Efficiency of Adversarial Patch Defenses
von: Xiang, Chong, et al.
Veröffentlicht: (2023)
von: Xiang, Chong, et al.
Veröffentlicht: (2023)
Efficient Storage Integrity in Adversarial Settings
von: Burke, Quinn, et al.
Veröffentlicht: (2025)
von: Burke, Quinn, et al.
Veröffentlicht: (2025)
IrisFP: Adversarial-Example-based Model Fingerprinting with Enhanced Uniqueness and Robustness
von: Geng, Ziye, et al.
Veröffentlicht: (2026)
von: Geng, Ziye, et al.
Veröffentlicht: (2026)
Updating Windows Malware Detectors: Balancing Robustness and Regression against Adversarial EXEmples
von: Kozak, Matous, et al.
Veröffentlicht: (2024)
von: Kozak, Matous, et al.
Veröffentlicht: (2024)
Exploring the Robustness and Transferability of Patch-Based Adversarial Attacks in Quantized Neural Networks
von: Guesmi, Amira, et al.
Veröffentlicht: (2024)
von: Guesmi, Amira, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Deliberative Alignment: Reasoning Enables Safer Language Models
von: Guan, Melody Y., et al.
Veröffentlicht: (2024) -
The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
von: Wallace, Eric, et al.
Veröffentlicht: (2024) -
The Inherent Adversarial Robustness of Analog In-Memory Computing
von: Lammie, Corey, et al.
Veröffentlicht: (2024) -
AQUA-LLM: Evaluating Accuracy, Quantization, and Adversarial Robustness Trade-offs in LLMs for Cybersecurity Question Answering
von: Gungor, Onat, et al.
Veröffentlicht: (2025) -
A Computational Separation Between Quantum No-cloning and No-telegraphing
von: Nehoran, Barak, et al.
Veröffentlicht: (2023)