ALICE: An Interpretable Neural Architecture for Generalization in Substitution Ciphers
Fuente:
arXiv
Guardado en:
| Autores principales: | Shen, Jeff, Smith, Lindsay M. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Interpreting the Repeated Token Phenomenon in Large Language Models
por: Yona, Itay, et al.
Publicado: (2025)
por: Yona, Itay, et al.
Publicado: (2025)
Systematically Analyzing Prompt Injection Vulnerabilities in Diverse LLM Architectures
por: Benjamin, Victoria, et al.
Publicado: (2024)
por: Benjamin, Victoria, et al.
Publicado: (2024)
An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks
por: Boreiko, Valentyn, et al.
Publicado: (2024)
por: Boreiko, Valentyn, et al.
Publicado: (2024)
Adversarial Attacks on Parts of Speech: An Empirical Study in Text-to-Image Generation
por: Shahariar, G M, et al.
Publicado: (2024)
por: Shahariar, G M, et al.
Publicado: (2024)
SAEs $\textit{Can}$ Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs
por: Muhamed, Aashiq, et al.
Publicado: (2025)
por: Muhamed, Aashiq, et al.
Publicado: (2025)
Automated Software Vulnerability Static Code Analysis Using Generative Pre-Trained Transformer Models
por: Pelofske, Elijah, et al.
Publicado: (2024)
por: Pelofske, Elijah, et al.
Publicado: (2024)
Query-Based Adversarial Prompt Generation
por: Hayase, Jonathan, et al.
Publicado: (2024)
por: Hayase, Jonathan, et al.
Publicado: (2024)
Preference Tuning For Toxicity Mitigation Generalizes Across Languages
por: Li, Xiaochen, et al.
Publicado: (2024)
por: Li, Xiaochen, et al.
Publicado: (2024)
Was it Slander? Towards Exact Inversion of Generative Language Models
por: Skapars, Adrians, et al.
Publicado: (2024)
por: Skapars, Adrians, et al.
Publicado: (2024)
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
por: Chu, Junjie, et al.
Publicado: (2024)
por: Chu, Junjie, et al.
Publicado: (2024)
Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
por: Betley, Jan, et al.
Publicado: (2025)
por: Betley, Jan, et al.
Publicado: (2025)
A Generative Approach to LLM Harmfulness Mitigation with Red Flag Tokens
por: Dobre, David, et al.
Publicado: (2025)
por: Dobre, David, et al.
Publicado: (2025)
On Evaluating The Performance of Watermarked Machine-Generated Texts Under Adversarial Attacks
por: Liu, Zesen, et al.
Publicado: (2024)
por: Liu, Zesen, et al.
Publicado: (2024)
EPSVec: Efficient and Private Synthetic Data Generation via Dataset Vectors
por: Banayeeanzade, Amin, et al.
Publicado: (2026)
por: Banayeeanzade, Amin, et al.
Publicado: (2026)
Tuning without Peeking: Provable Generalization Bounds and Robust LLM Post-Training
por: Labiad, Ismail, et al.
Publicado: (2025)
por: Labiad, Ismail, et al.
Publicado: (2025)
MetaDefense: Defending Finetuning-based Jailbreak Attack Before and During Generation
por: Jiang, Weisen, et al.
Publicado: (2025)
por: Jiang, Weisen, et al.
Publicado: (2025)
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs
por: Liang, Buyun, et al.
Publicado: (2025)
por: Liang, Buyun, et al.
Publicado: (2025)
Prompt2Fingerprint: Plug-and-Play LLM Fingerprinting via Text-to-Weight Generation
por: Chen, Sixu, et al.
Publicado: (2026)
por: Chen, Sixu, et al.
Publicado: (2026)
Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation Systems
por: Qi, Zhenting, et al.
Publicado: (2024)
por: Qi, Zhenting, et al.
Publicado: (2024)
SecureCode: A Production-Grade Multi-Turn Dataset for Training Security-Aware Code Generation Models
por: Thornton, Scott
Publicado: (2025)
por: Thornton, Scott
Publicado: (2025)
IsolateGPT: An Execution Isolation Architecture for LLM-Based Agentic Systems
por: Wu, Yuhao, et al.
Publicado: (2024)
por: Wu, Yuhao, et al.
Publicado: (2024)
BountyBench: Dollar Impact of AI Agent Attackers and Defenders on Real-World Cybersecurity Systems
por: Zhang, Andy K., et al.
Publicado: (2025)
por: Zhang, Andy K., et al.
Publicado: (2025)
A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment
por: Wang, Kun, et al.
Publicado: (2025)
por: Wang, Kun, et al.
Publicado: (2025)
Attack and defense techniques in large language models: A survey and new perspectives
por: Liao, Zhiyu, et al.
Publicado: (2025)
por: Liao, Zhiyu, et al.
Publicado: (2025)
The Resurgence of GCG Adversarial Attacks on Large Language Models
por: Tan, Yuting, et al.
Publicado: (2025)
por: Tan, Yuting, et al.
Publicado: (2025)
When Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Models
por: Wang, Kai, et al.
Publicado: (2025)
por: Wang, Kai, et al.
Publicado: (2025)
Toward a Safer Web: Multilingual Multi-Agent LLMs for Mitigating Adversarial Misinformation Attacks
por: Aldahoul, Nouar, et al.
Publicado: (2025)
por: Aldahoul, Nouar, et al.
Publicado: (2025)
How Different Tokenization Algorithms Impact LLMs and Transformer Models for Binary Code Analysis
por: Mostafa, Ahmed, et al.
Publicado: (2025)
por: Mostafa, Ahmed, et al.
Publicado: (2025)
Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring
por: Hua, Peichun, et al.
Publicado: (2025)
por: Hua, Peichun, et al.
Publicado: (2025)
SECA: Semantically Equivalent and Coherent Attacks for Eliciting LLM Hallucinations
por: Liang, Buyun, et al.
Publicado: (2025)
por: Liang, Buyun, et al.
Publicado: (2025)
FedMentor: Domain-Aware Differential Privacy for Heterogeneous Federated LLMs in Mental Health
por: Sarwar, Nobin, et al.
Publicado: (2025)
por: Sarwar, Nobin, et al.
Publicado: (2025)
Shh, don't say that! Domain Certification in LLMs
por: Emde, Cornelius, et al.
Publicado: (2025)
por: Emde, Cornelius, et al.
Publicado: (2025)
LLMs Have Rhythm: Fingerprinting Large Language Models Using Inter-Token Times and Network Traffic Analysis
por: Alhazbi, Saeif, et al.
Publicado: (2025)
por: Alhazbi, Saeif, et al.
Publicado: (2025)
PBa-LLM: Privacy- and Bias-aware NLP using Named-Entity Recognition (NER)
por: Mancera, Gonzalo, et al.
Publicado: (2025)
por: Mancera, Gonzalo, et al.
Publicado: (2025)
Does Low Rank Adaptation Lead to Lower Robustness against Training-Time Attacks?
por: Liang, Zi, et al.
Publicado: (2025)
por: Liang, Zi, et al.
Publicado: (2025)
Clustering and Median Aggregation Improve Differentially Private Inference
por: Amin, Kareem, et al.
Publicado: (2025)
por: Amin, Kareem, et al.
Publicado: (2025)
AutoMalDesc: Large-Scale Script Analysis for Cyber Threat Research
por: Apostu, Alexandru-Mihai, et al.
Publicado: (2025)
por: Apostu, Alexandru-Mihai, et al.
Publicado: (2025)
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
por: Betley, Jan, et al.
Publicado: (2025)
por: Betley, Jan, et al.
Publicado: (2025)
Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
por: Sharma, Mrinank, et al.
Publicado: (2025)
por: Sharma, Mrinank, et al.
Publicado: (2025)
LLM Cyber Evaluations Don't Capture Real-World Risk
por: Lukošiūtė, Kamilė, et al.
Publicado: (2025)
por: Lukošiūtė, Kamilė, et al.
Publicado: (2025)
Ejemplares similares
-
Interpreting the Repeated Token Phenomenon in Large Language Models
por: Yona, Itay, et al.
Publicado: (2025) -
Systematically Analyzing Prompt Injection Vulnerabilities in Diverse LLM Architectures
por: Benjamin, Victoria, et al.
Publicado: (2024) -
An Interpretable N-gram Perplexity Threat Model for Large Language Model Jailbreaks
por: Boreiko, Valentyn, et al.
Publicado: (2024) -
Adversarial Attacks on Parts of Speech: An Empirical Study in Text-to-Image Generation
por: Shahariar, G M, et al.
Publicado: (2024) -
SAEs $\textit{Can}$ Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs
por: Muhamed, Aashiq, et al.
Publicado: (2025)