Large Language Models Can Verbatim Reproduce Long Malicious Sequences
Fuente:
arXiv
Guardado en:
| Autores principales: | Lin, Sharon, Krishnamurthy, Dvijotham, Hayes, Jamie, Shi, Chongyang, Shumailov, Ilia, Song, Shuang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Buffer Overflow in Mixture of Experts
por: Hayes, Jamie, et al.
Publicado: (2024)
por: Hayes, Jamie, et al.
Publicado: (2024)
Interpreting the Repeated Token Phenomenon in Large Language Models
por: Yona, Itay, et al.
Publicado: (2025)
por: Yona, Itay, et al.
Publicado: (2025)
Beyond Slow Signs in High-fidelity Model Extraction
por: Foerster, Hanna, et al.
Publicado: (2024)
por: Foerster, Hanna, et al.
Publicado: (2024)
Measuring memorization in RLHF for code completion
por: Pappu, Aneesh, et al.
Publicado: (2024)
por: Pappu, Aneesh, et al.
Publicado: (2024)
Cascading Adversarial Bias from Injection to Distillation in Language Models
por: Chaudhari, Harsh, et al.
Publicado: (2025)
por: Chaudhari, Harsh, et al.
Publicado: (2025)
Stealing User Prompts from Mixture of Experts
por: Yona, Itay, et al.
Publicado: (2024)
por: Yona, Itay, et al.
Publicado: (2024)
Inexact Unlearning Needs More Careful Evaluations to Avoid a False Sense of Privacy
por: Hayes, Jamie, et al.
Publicado: (2024)
por: Hayes, Jamie, et al.
Publicado: (2024)
Soft Instruction De-escalation Defense
por: Walter, Nils Philipp, et al.
Publicado: (2025)
por: Walter, Nils Philipp, et al.
Publicado: (2025)
Lessons from Defending Gemini Against Indirect Prompt Injections
por: Shi, Chongyang, et al.
Publicado: (2025)
por: Shi, Chongyang, et al.
Publicado: (2025)
Quantamination: Dynamic Quantization Leaks Your Data Across the Batch
por: Foerster, Hanna, et al.
Publicado: (2026)
por: Foerster, Hanna, et al.
Publicado: (2026)
Demystifying Verbatim Memorization in Large Language Models
por: Huang, Jing, et al.
Publicado: (2024)
por: Huang, Jing, et al.
Publicado: (2024)
Thought-Transfer: Indirect Targeted Poisoning Attacks on Chain-of-Thought Reasoning Models
por: Chaudhari, Harsh, et al.
Publicado: (2026)
por: Chaudhari, Harsh, et al.
Publicado: (2026)
Locking Machine Learning Models into Hardware
por: Clifford, Eleanor, et al.
Publicado: (2024)
por: Clifford, Eleanor, et al.
Publicado: (2024)
Fairness Feedback Loops: Training on Synthetic Data Amplifies Bias
por: Wyllie, Sierra, et al.
Publicado: (2024)
por: Wyllie, Sierra, et al.
Publicado: (2024)
Achieving the Tightest Relaxation of Sigmoids for Formal Verification
por: Chevalier, Samuel, et al.
Publicado: (2024)
por: Chevalier, Samuel, et al.
Publicado: (2024)
Reasoning Introduces New Poisoning Attacks Yet Makes Them More Complicated
por: Foerster, Hanna, et al.
Publicado: (2025)
por: Foerster, Hanna, et al.
Publicado: (2025)
Norm-Bounded Low-Rank Adaptation
por: Wang, Ruigang, et al.
Publicado: (2025)
por: Wang, Ruigang, et al.
Publicado: (2025)
Monotone, Bi-Lipschitz, and Polyak-Lojasiewicz Networks
por: Wang, Ruigang, et al.
Publicado: (2024)
por: Wang, Ruigang, et al.
Publicado: (2024)
SEA: Shareable and Explainable Attribution for Query-based Black-box Attacks
por: Gao, Yue, et al.
Publicado: (2023)
por: Gao, Yue, et al.
Publicado: (2023)
Machine Learning needs Better Randomness Standards: Randomised Smoothing and PRNG-based attacks
por: Dahiya, Pranav, et al.
Publicado: (2023)
por: Dahiya, Pranav, et al.
Publicado: (2023)
Keeping up with dynamic attackers: Certifying robustness to adaptive online data poisoning
por: Bose, Avinandan, et al.
Publicado: (2025)
por: Bose, Avinandan, et al.
Publicado: (2025)
Architectural Backdoors for Within-Batch Data Stealing and Model Inference Manipulation
por: Küchler, Nicolas, et al.
Publicado: (2025)
por: Küchler, Nicolas, et al.
Publicado: (2025)
Measuring memorization in language models via probabilistic extraction
por: Hayes, Jamie, et al.
Publicado: (2024)
por: Hayes, Jamie, et al.
Publicado: (2024)
Honeyval: A Comprehensive Evaluation Framework for LLM-powered HTTP Honeypots
por: Vero, Mark, et al.
Publicado: (2026)
por: Vero, Mark, et al.
Publicado: (2026)
ceLLMate: Sandboxing Browser AI Agents
por: Meng, Luoxi, et al.
Publicado: (2025)
por: Meng, Luoxi, et al.
Publicado: (2025)
ParaPO: Aligning Language Models to Reduce Verbatim Reproduction of Pre-training Data
por: Chen, Tong, et al.
Publicado: (2025)
por: Chen, Tong, et al.
Publicado: (2025)
Beyond Labeling Oracles: What does it mean to steal ML models?
por: Shafran, Avital, et al.
Publicado: (2023)
por: Shafran, Avital, et al.
Publicado: (2023)
Watermarking Needs Input Repetition Masking
por: Khachaturov, David, et al.
Publicado: (2025)
por: Khachaturov, David, et al.
Publicado: (2025)
CrisperWhisper: Accurate Timestamps on Verbatim Speech Transcriptions
por: Wagner, Laurin, et al.
Publicado: (2024)
por: Wagner, Laurin, et al.
Publicado: (2024)
Language Models May Verbatim Complete Text They Were Not Explicitly Trained On
por: Liu, Ken Ziyu, et al.
Publicado: (2025)
por: Liu, Ken Ziyu, et al.
Publicado: (2025)
The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
por: Nasr, Milad, et al.
Publicado: (2025)
por: Nasr, Milad, et al.
Publicado: (2025)
Hardware and Software Platform Inference
por: Zhang, Cheng, et al.
Publicado: (2024)
por: Zhang, Cheng, et al.
Publicado: (2024)
Revisiting Block-based Quantisation: What is Important for Sub-8-bit LLM Inference?
por: Zhang, Cheng, et al.
Publicado: (2023)
por: Zhang, Cheng, et al.
Publicado: (2023)
UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI
por: Shumailov, Ilia, et al.
Publicado: (2024)
por: Shumailov, Ilia, et al.
Publicado: (2024)
Machine Learning Models Have a Supply Chain Problem
por: Meiklejohn, Sarah, et al.
Publicado: (2025)
por: Meiklejohn, Sarah, et al.
Publicado: (2025)
Confidence-aware Reward Optimization for Fine-tuning Text-to-Image Models
por: Kim, Kyuyoung, et al.
Publicado: (2024)
por: Kim, Kyuyoung, et al.
Publicado: (2024)
Beyond Laplace and Gaussian: Exploring the Generalized Gaussian Mechanism for Private Machine Learning
por: Rinberg, Roy, et al.
Publicado: (2025)
por: Rinberg, Roy, et al.
Publicado: (2025)
ImpNet: Imperceptible and blackbox-undetectable backdoors in compiled neural networks
por: Clifford, Eleanor, et al.
Publicado: (2022)
por: Clifford, Eleanor, et al.
Publicado: (2022)
When Vision Fails: Text Attacks Against ViT and OCR
por: Boucher, Nicholas, et al.
Publicado: (2023)
por: Boucher, Nicholas, et al.
Publicado: (2023)
No, of Course I Can! Deeper Fine-Tuning Attacks That Bypass Token-Level Safety Mechanisms
por: Kazdan, Joshua, et al.
Publicado: (2025)
por: Kazdan, Joshua, et al.
Publicado: (2025)
Ejemplares similares
-
Buffer Overflow in Mixture of Experts
por: Hayes, Jamie, et al.
Publicado: (2024) -
Interpreting the Repeated Token Phenomenon in Large Language Models
por: Yona, Itay, et al.
Publicado: (2025) -
Beyond Slow Signs in High-fidelity Model Extraction
por: Foerster, Hanna, et al.
Publicado: (2024) -
Measuring memorization in RLHF for code completion
por: Pappu, Aneesh, et al.
Publicado: (2024) -
Cascading Adversarial Bias from Injection to Distillation in Language Models
por: Chaudhari, Harsh, et al.
Publicado: (2025)