Are aligned neural networks adversarially aligned?
Fuente:
arXiv
Saved in:
| Main Authors: | Carlini, Nicholas, Nasr, Milad, Choquette-Choo, Christopher A., Jagielski, Matthew, Gao, Irena, Awadalla, Anas, Koh, Pang Wei, Ippolito, Daphne, Lee, Katherine, Tramer, Florian, Schmidt, Ludwig |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLMs unlock new paths to monetizing exploits
by: Carlini, Nicholas, et al.
Published: (2025)
by: Carlini, Nicholas, et al.
Published: (2025)
Privacy Side Channels in Machine Learning Systems
by: Debenedetti, Edoardo, et al.
Published: (2023)
by: Debenedetti, Edoardo, et al.
Published: (2023)
AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses
by: Carlini, Nicholas, et al.
Published: (2025)
by: Carlini, Nicholas, et al.
Published: (2025)
Exploring and Mitigating Adversarial Manipulation of Voting-Based Leaderboards
by: Huang, Yangsibo, et al.
Published: (2025)
by: Huang, Yangsibo, et al.
Published: (2025)
Auditing Private Prediction
by: Chadha, Karan, et al.
Published: (2024)
by: Chadha, Karan, et al.
Published: (2024)
Poisoning Web-Scale Training Datasets is Practical
by: Carlini, Nicholas, et al.
Published: (2023)
by: Carlini, Nicholas, et al.
Published: (2023)
Extracting alignment data in open models
by: Barbero, Federico, et al.
Published: (2025)
by: Barbero, Federico, et al.
Published: (2025)
Query-Based Adversarial Prompt Generation
by: Hayase, Jonathan, et al.
Published: (2024)
by: Hayase, Jonathan, et al.
Published: (2024)
Measuring Non-Adversarial Reproduction of Training Data in Large Language Models
by: Aerni, Michael, et al.
Published: (2024)
by: Aerni, Michael, et al.
Published: (2024)
Remote Timing Attacks on Efficient Language Model Inference
by: Carlini, Nicholas, et al.
Published: (2024)
by: Carlini, Nicholas, et al.
Published: (2024)
Effective Prompt Extraction from Language Models
by: Zhang, Yiming, et al.
Published: (2023)
by: Zhang, Yiming, et al.
Published: (2023)
Human-aligned Chess with a Bit of Search
by: Zhang, Yiming, et al.
Published: (2024)
by: Zhang, Yiming, et al.
Published: (2024)
Phantom: General Backdoor Attacks on Retrieval Augmented Language Generation
by: Chaudhari, Harsh, et al.
Published: (2024)
by: Chaudhari, Harsh, et al.
Published: (2024)
Privacy Auditing of Large Language Models
by: Panda, Ashwinee, et al.
Published: (2025)
by: Panda, Ashwinee, et al.
Published: (2025)
Evading Black-box Classifiers Without Breaking Eggs
by: Debenedetti, Edoardo, et al.
Published: (2023)
by: Debenedetti, Edoardo, et al.
Published: (2023)
Position: Considerations for Differentially Private Learning with Large-Scale Public Pretraining
by: Tramèr, Florian, et al.
Published: (2022)
by: Tramèr, Florian, et al.
Published: (2022)
Persistent Pre-Training Poisoning of LLMs
by: Zhang, Yiming, et al.
Published: (2024)
by: Zhang, Yiming, et al.
Published: (2024)
Privacy Ripple Effects from Adding or Removing Personal Information in Language Model Training
by: Borkar, Jaydeep, et al.
Published: (2025)
by: Borkar, Jaydeep, et al.
Published: (2025)
The Last Iterate Advantage: Empirical Auditing and Principled Heuristic Analysis of Differentially Private SGD
by: Steinke, Thomas, et al.
Published: (2024)
by: Steinke, Thomas, et al.
Published: (2024)
Stealing Part of a Production Language Model
by: Carlini, Nicholas, et al.
Published: (2024)
by: Carlini, Nicholas, et al.
Published: (2024)
Dimensions underlying the representational alignment of deep neural networks with humans
by: Mahner, Florian P., et al.
Published: (2024)
by: Mahner, Florian P., et al.
Published: (2024)
Human alignment of neural network representations
by: Muttenthaler, Lukas, et al.
Published: (2022)
by: Muttenthaler, Lukas, et al.
Published: (2022)
Adversarial Perturbations Cannot Reliably Protect Artists From Generative AI
by: Hönig, Robert, et al.
Published: (2024)
by: Hönig, Robert, et al.
Published: (2024)
Adversarial ML Problems Are Getting Harder to Solve and to Evaluate
by: Rando, Javier, et al.
Published: (2025)
by: Rando, Javier, et al.
Published: (2025)
Forcing Diffuse Distributions out of Language Models
by: Zhang, Yiming, et al.
Published: (2024)
by: Zhang, Yiming, et al.
Published: (2024)
Annotation alignment: Comparing LLM and human annotations of conversational safety
by: Movva, Rajiv, et al.
Published: (2024)
by: Movva, Rajiv, et al.
Published: (2024)
Cutting through buggy adversarial example defenses: fixing 1 line of code breaks Sabre
by: Carlini, Nicholas
Published: (2024)
by: Carlini, Nicholas
Published: (2024)
On Evaluating the Durability of Safeguards for Open-Weight LLMs
by: Qi, Xiangyu, et al.
Published: (2024)
by: Qi, Xiangyu, et al.
Published: (2024)
Cascading Adversarial Bias from Injection to Distillation in Language Models
by: Chaudhari, Harsh, et al.
Published: (2025)
by: Chaudhari, Harsh, et al.
Published: (2025)
Language Models May Verbatim Complete Text They Were Not Explicitly Trained On
by: Liu, Ken Ziyu, et al.
Published: (2025)
by: Liu, Ken Ziyu, et al.
Published: (2025)
Measuring memorization in language models via probabilistic extraction
by: Hayes, Jamie, et al.
Published: (2024)
by: Hayes, Jamie, et al.
Published: (2024)
Getting aligned on representational alignment
by: Sucholutsky, Ilia, et al.
Published: (2023)
by: Sucholutsky, Ilia, et al.
Published: (2023)
Sparse components distinguish visual pathways & their alignment to neural networks
by: Marvi, Ammar I, et al.
Published: (2025)
by: Marvi, Ammar I, et al.
Published: (2025)
Synthetic Query Generation for Privacy-Preserving Deep Retrieval Systems using Differentially Private Language Models
by: Carranza, Aldo Gael, et al.
Published: (2023)
by: Carranza, Aldo Gael, et al.
Published: (2023)
'Too much alignment; not enough culture': Re-balancing cultural alignment practices in LLMs
by: Orlowski, Eric J. W., et al.
Published: (2025)
by: Orlowski, Eric J. W., et al.
Published: (2025)
Effective faking of verbal deception detection with target-aligned adversarial attacks
by: Kleinberg, Bennett, et al.
Published: (2025)
by: Kleinberg, Bennett, et al.
Published: (2025)
Effective faking of verbal deception detection with target‐aligned adversarial attacks
by: Bennett Kleinberg, et al.
Published: (2025)
by: Bennett Kleinberg, et al.
Published: (2025)
Kinematical alignment better restores native patellar tracking pattern than mechanical alignment
by: Yong Deok Kim, et al.
Published: (2024)
by: Yong Deok Kim, et al.
Published: (2024)
Precipitation nowcasting of satellite data using physically-aligned neural networks
by: Catão, Antônio, et al.
Published: (2025)
by: Catão, Antônio, et al.
Published: (2025)
The Mason-Alberta Phonetic Segmenter: A forced alignment system based on deep neural networks and interpolation
by: Kelley, Matthew C., et al.
Published: (2023)
by: Kelley, Matthew C., et al.
Published: (2023)
Similar Items
-
LLMs unlock new paths to monetizing exploits
by: Carlini, Nicholas, et al.
Published: (2025) -
Privacy Side Channels in Machine Learning Systems
by: Debenedetti, Edoardo, et al.
Published: (2023) -
AutoAdvExBench: Benchmarking autonomous exploitation of adversarial example defenses
by: Carlini, Nicholas, et al.
Published: (2025) -
Exploring and Mitigating Adversarial Manipulation of Voting-Based Leaderboards
by: Huang, Yangsibo, et al.
Published: (2025) -
Auditing Private Prediction
by: Chadha, Karan, et al.
Published: (2024)