Toward Understanding the Transferability of Adversarial Suffixes in Large Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Ball, Sarah, Hasrati, Niki, Robey, Alexander, Schwarzschild, Avi, Kreuter, Frauke, Kolter, Zico, Risteski, Andrej |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Existing Large Language Model Unlearning Evaluations Are Inconclusive
por: Feng, Zhili, et al.
Publicado: (2025)
por: Feng, Zhili, et al.
Publicado: (2025)
Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models
por: Ball, Sarah, et al.
Publicado: (2024)
por: Ball, Sarah, et al.
Publicado: (2024)
Antidistillation Sampling
por: Savani, Yash, et al.
Publicado: (2025)
por: Savani, Yash, et al.
Publicado: (2025)
Forcing Diffuse Distributions out of Language Models
por: Zhang, Yiming, et al.
Publicado: (2024)
por: Zhang, Yiming, et al.
Publicado: (2024)
Understanding and Mitigating Premature Confidence for Better LLM Reasoning
por: Gai, Jingchu, et al.
Publicado: (2026)
por: Gai, Jingchu, et al.
Publicado: (2026)
Human Preferences in Large Language Model Latent Space: A Technical Analysis on the Reliability of Synthetic Data in Voting Outcome Prediction
por: Ball, Sarah, et al.
Publicado: (2025)
por: Ball, Sarah, et al.
Publicado: (2025)
A Simple and Effective Pruning Approach for Large Language Models
por: Sun, Mingjie, et al.
Publicado: (2023)
por: Sun, Mingjie, et al.
Publicado: (2023)
Rethinking LLM Memorization through the Lens of Adversarial Compression
por: Schwarzschild, Avi, et al.
Publicado: (2024)
por: Schwarzschild, Avi, et al.
Publicado: (2024)
Extrapolation by Association: Length Generalization Transfer in Transformers
por: Cai, Ziyang, et al.
Publicado: (2025)
por: Cai, Ziyang, et al.
Publicado: (2025)
Adversarial Attacks on Robotic Vision Language Action Models
por: Jones, Eliot Krzysztof, et al.
Publicado: (2025)
por: Jones, Eliot Krzysztof, et al.
Publicado: (2025)
Antidistillation Fingerprinting
por: Xu, Yixuan Even, et al.
Publicado: (2026)
por: Xu, Yixuan Even, et al.
Publicado: (2026)
Benchmarking ChatGPT on Algorithmic Reasoning
por: McLeish, Sean, et al.
Publicado: (2024)
por: McLeish, Sean, et al.
Publicado: (2024)
Adaptive Content Restriction for Large Language Models via Suffix Optimization
por: Li, Yige, et al.
Publicado: (2025)
por: Li, Yige, et al.
Publicado: (2025)
Moral Lenses, Political Coordinates: Towards Ideological Positioning of Morally Conditioned LLMs
por: Yuan, Chenchen, et al.
Publicado: (2026)
por: Yuan, Chenchen, et al.
Publicado: (2026)
TOFU: A Task of Fictitious Unlearning for LLMs
por: Maini, Pratyush, et al.
Publicado: (2024)
por: Maini, Pratyush, et al.
Publicado: (2024)
Looking beyond the next token
por: Thankaraj, Abitha, et al.
Publicado: (2025)
por: Thankaraj, Abitha, et al.
Publicado: (2025)
Algorithms for Adversarially Robust Deep Learning
por: Robey, Alexander
Publicado: (2025)
por: Robey, Alexander
Publicado: (2025)
Base Models Look Human To AI Detectors
por: Xu, Yixuan Even, et al.
Publicado: (2026)
por: Xu, Yixuan Even, et al.
Publicado: (2026)
Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
por: Xu, Yixuan Even, et al.
Publicado: (2025)
por: Xu, Yixuan Even, et al.
Publicado: (2025)
From Ground Truth to Measurement: A Statistical Framework for Human Labeling
por: Chew, Robert, et al.
Publicado: (2026)
por: Chew, Robert, et al.
Publicado: (2026)
Moral Persuasion in Large Language Models: Evaluating Susceptibility and Ethical Alignment
por: Huang, Allison, et al.
Publicado: (2024)
por: Huang, Allison, et al.
Publicado: (2024)
Understanding Augmentation-based Self-Supervised Representation Learning via RKHS Approximation and Regression
por: Zhai, Runtian, et al.
Publicado: (2023)
por: Zhai, Runtian, et al.
Publicado: (2023)
Mask-GCG: Are All Tokens in Adversarial Suffixes Necessary for Jailbreak Attacks?
por: Mu, Junjie, et al.
Publicado: (2025)
por: Mu, Junjie, et al.
Publicado: (2025)
Automated Black-box Prompt Engineering for Personalized Text-to-Image Generation
por: He, Yutong, et al.
Publicado: (2024)
por: He, Yutong, et al.
Publicado: (2024)
Has My System Prompt Been Used? Large Language Model Prompt Membership Inference
por: Levin, Roman, et al.
Publicado: (2025)
por: Levin, Roman, et al.
Publicado: (2025)
Mitigating Adversarial Attacks in LLMs through Defensive Suffix Generation
por: Kim, Minkyoung, et al.
Publicado: (2024)
por: Kim, Minkyoung, et al.
Publicado: (2024)
AI Conversational Interviewing: Transforming Surveys with LLMs as Adaptive Interviewers
por: Wuttke, Alexander, et al.
Publicado: (2024)
por: Wuttke, Alexander, et al.
Publicado: (2024)
Reasoning Capabilities of Large Language Models on Dynamic Tasks
por: Wong, Annie, et al.
Publicado: (2025)
por: Wong, Annie, et al.
Publicado: (2025)
SarcasmBench: Towards Evaluating Large Language Models on Sarcasm Understanding
por: Zhang, Yazhou, et al.
Publicado: (2024)
por: Zhang, Yazhou, et al.
Publicado: (2024)
Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence
por: McLeish, Sean, et al.
Publicado: (2025)
por: McLeish, Sean, et al.
Publicado: (2025)
On the Impossibility of Separating Intelligence from Judgment: The Computational Intractability of Filtering for AI Alignment
por: Ball, Sarah, et al.
Publicado: (2025)
por: Ball, Sarah, et al.
Publicado: (2025)
Weight Ensembling Improves Reasoning in Language Models
por: Dang, Xingyu, et al.
Publicado: (2025)
por: Dang, Xingyu, et al.
Publicado: (2025)
Route to Rome Attack: Directing LLM Routers to Expensive Models via Adversarial Suffix Optimization
por: Tang, Haochun, et al.
Publicado: (2026)
por: Tang, Haochun, et al.
Publicado: (2026)
Talking Nonsense: Probing Large Language Models' Understanding of Adversarial Gibberish Inputs
por: Cherepanova, Valeriia, et al.
Publicado: (2024)
por: Cherepanova, Valeriia, et al.
Publicado: (2024)
Unnatural Languages Are Not Bugs but Features for LLMs
por: Duan, Keyu, et al.
Publicado: (2025)
por: Duan, Keyu, et al.
Publicado: (2025)
Prompt Recovery for Image Generation Models: A Comparative Study of Discrete Optimizers
por: Williams, Joshua Nathaniel, et al.
Publicado: (2024)
por: Williams, Joshua Nathaniel, et al.
Publicado: (2024)
State Design Matters: How Representations Shape Dynamic Reasoning in Large Language Models
por: Wong, Annie, et al.
Publicado: (2026)
por: Wong, Annie, et al.
Publicado: (2026)
Understanding Moral Reasoning Trajectories in Large Language Models: Toward Probing-Based Explainability
por: Huang, Fan, et al.
Publicado: (2026)
por: Huang, Fan, et al.
Publicado: (2026)
Mimetic Initialization of MLPs
por: Trockman, Asher, et al.
Publicado: (2026)
por: Trockman, Asher, et al.
Publicado: (2026)
Potemkin Understanding in Large Language Models
por: Mancoridis, Marina, et al.
Publicado: (2025)
por: Mancoridis, Marina, et al.
Publicado: (2025)
Ejemplares similares
-
Existing Large Language Model Unlearning Evaluations Are Inconclusive
por: Feng, Zhili, et al.
Publicado: (2025) -
Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models
por: Ball, Sarah, et al.
Publicado: (2024) -
Antidistillation Sampling
por: Savani, Yash, et al.
Publicado: (2025) -
Forcing Diffuse Distributions out of Language Models
por: Zhang, Yiming, et al.
Publicado: (2024) -
Understanding and Mitigating Premature Confidence for Better LLM Reasoning
por: Gai, Jingchu, et al.
Publicado: (2026)