Large Language Models Can Verbatim Reproduce Long Malicious Sequences
Fuente:
arXiv
Salvato in:
| Autori principali: | Lin, Sharon, Krishnamurthy, Dvijotham, Hayes, Jamie, Shi, Chongyang, Shumailov, Ilia, Song, Shuang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Buffer Overflow in Mixture of Experts
di: Hayes, Jamie, et al.
Pubblicazione: (2024)
di: Hayes, Jamie, et al.
Pubblicazione: (2024)
Interpreting the Repeated Token Phenomenon in Large Language Models
di: Yona, Itay, et al.
Pubblicazione: (2025)
di: Yona, Itay, et al.
Pubblicazione: (2025)
Beyond Slow Signs in High-fidelity Model Extraction
di: Foerster, Hanna, et al.
Pubblicazione: (2024)
di: Foerster, Hanna, et al.
Pubblicazione: (2024)
Measuring memorization in RLHF for code completion
di: Pappu, Aneesh, et al.
Pubblicazione: (2024)
di: Pappu, Aneesh, et al.
Pubblicazione: (2024)
Cascading Adversarial Bias from Injection to Distillation in Language Models
di: Chaudhari, Harsh, et al.
Pubblicazione: (2025)
di: Chaudhari, Harsh, et al.
Pubblicazione: (2025)
Stealing User Prompts from Mixture of Experts
di: Yona, Itay, et al.
Pubblicazione: (2024)
di: Yona, Itay, et al.
Pubblicazione: (2024)
Inexact Unlearning Needs More Careful Evaluations to Avoid a False Sense of Privacy
di: Hayes, Jamie, et al.
Pubblicazione: (2024)
di: Hayes, Jamie, et al.
Pubblicazione: (2024)
Soft Instruction De-escalation Defense
di: Walter, Nils Philipp, et al.
Pubblicazione: (2025)
di: Walter, Nils Philipp, et al.
Pubblicazione: (2025)
Lessons from Defending Gemini Against Indirect Prompt Injections
di: Shi, Chongyang, et al.
Pubblicazione: (2025)
di: Shi, Chongyang, et al.
Pubblicazione: (2025)
Quantamination: Dynamic Quantization Leaks Your Data Across the Batch
di: Foerster, Hanna, et al.
Pubblicazione: (2026)
di: Foerster, Hanna, et al.
Pubblicazione: (2026)
Demystifying Verbatim Memorization in Large Language Models
di: Huang, Jing, et al.
Pubblicazione: (2024)
di: Huang, Jing, et al.
Pubblicazione: (2024)
Thought-Transfer: Indirect Targeted Poisoning Attacks on Chain-of-Thought Reasoning Models
di: Chaudhari, Harsh, et al.
Pubblicazione: (2026)
di: Chaudhari, Harsh, et al.
Pubblicazione: (2026)
Locking Machine Learning Models into Hardware
di: Clifford, Eleanor, et al.
Pubblicazione: (2024)
di: Clifford, Eleanor, et al.
Pubblicazione: (2024)
Fairness Feedback Loops: Training on Synthetic Data Amplifies Bias
di: Wyllie, Sierra, et al.
Pubblicazione: (2024)
di: Wyllie, Sierra, et al.
Pubblicazione: (2024)
Achieving the Tightest Relaxation of Sigmoids for Formal Verification
di: Chevalier, Samuel, et al.
Pubblicazione: (2024)
di: Chevalier, Samuel, et al.
Pubblicazione: (2024)
Reasoning Introduces New Poisoning Attacks Yet Makes Them More Complicated
di: Foerster, Hanna, et al.
Pubblicazione: (2025)
di: Foerster, Hanna, et al.
Pubblicazione: (2025)
Norm-Bounded Low-Rank Adaptation
di: Wang, Ruigang, et al.
Pubblicazione: (2025)
di: Wang, Ruigang, et al.
Pubblicazione: (2025)
Monotone, Bi-Lipschitz, and Polyak-Lojasiewicz Networks
di: Wang, Ruigang, et al.
Pubblicazione: (2024)
di: Wang, Ruigang, et al.
Pubblicazione: (2024)
SEA: Shareable and Explainable Attribution for Query-based Black-box Attacks
di: Gao, Yue, et al.
Pubblicazione: (2023)
di: Gao, Yue, et al.
Pubblicazione: (2023)
Machine Learning needs Better Randomness Standards: Randomised Smoothing and PRNG-based attacks
di: Dahiya, Pranav, et al.
Pubblicazione: (2023)
di: Dahiya, Pranav, et al.
Pubblicazione: (2023)
Keeping up with dynamic attackers: Certifying robustness to adaptive online data poisoning
di: Bose, Avinandan, et al.
Pubblicazione: (2025)
di: Bose, Avinandan, et al.
Pubblicazione: (2025)
Architectural Backdoors for Within-Batch Data Stealing and Model Inference Manipulation
di: Küchler, Nicolas, et al.
Pubblicazione: (2025)
di: Küchler, Nicolas, et al.
Pubblicazione: (2025)
Measuring memorization in language models via probabilistic extraction
di: Hayes, Jamie, et al.
Pubblicazione: (2024)
di: Hayes, Jamie, et al.
Pubblicazione: (2024)
Honeyval: A Comprehensive Evaluation Framework for LLM-powered HTTP Honeypots
di: Vero, Mark, et al.
Pubblicazione: (2026)
di: Vero, Mark, et al.
Pubblicazione: (2026)
ceLLMate: Sandboxing Browser AI Agents
di: Meng, Luoxi, et al.
Pubblicazione: (2025)
di: Meng, Luoxi, et al.
Pubblicazione: (2025)
ParaPO: Aligning Language Models to Reduce Verbatim Reproduction of Pre-training Data
di: Chen, Tong, et al.
Pubblicazione: (2025)
di: Chen, Tong, et al.
Pubblicazione: (2025)
Beyond Labeling Oracles: What does it mean to steal ML models?
di: Shafran, Avital, et al.
Pubblicazione: (2023)
di: Shafran, Avital, et al.
Pubblicazione: (2023)
Watermarking Needs Input Repetition Masking
di: Khachaturov, David, et al.
Pubblicazione: (2025)
di: Khachaturov, David, et al.
Pubblicazione: (2025)
CrisperWhisper: Accurate Timestamps on Verbatim Speech Transcriptions
di: Wagner, Laurin, et al.
Pubblicazione: (2024)
di: Wagner, Laurin, et al.
Pubblicazione: (2024)
Language Models May Verbatim Complete Text They Were Not Explicitly Trained On
di: Liu, Ken Ziyu, et al.
Pubblicazione: (2025)
di: Liu, Ken Ziyu, et al.
Pubblicazione: (2025)
The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
di: Nasr, Milad, et al.
Pubblicazione: (2025)
di: Nasr, Milad, et al.
Pubblicazione: (2025)
Hardware and Software Platform Inference
di: Zhang, Cheng, et al.
Pubblicazione: (2024)
di: Zhang, Cheng, et al.
Pubblicazione: (2024)
Revisiting Block-based Quantisation: What is Important for Sub-8-bit LLM Inference?
di: Zhang, Cheng, et al.
Pubblicazione: (2023)
di: Zhang, Cheng, et al.
Pubblicazione: (2023)
UnUnlearning: Unlearning is not sufficient for content regulation in advanced generative AI
di: Shumailov, Ilia, et al.
Pubblicazione: (2024)
di: Shumailov, Ilia, et al.
Pubblicazione: (2024)
Machine Learning Models Have a Supply Chain Problem
di: Meiklejohn, Sarah, et al.
Pubblicazione: (2025)
di: Meiklejohn, Sarah, et al.
Pubblicazione: (2025)
Confidence-aware Reward Optimization for Fine-tuning Text-to-Image Models
di: Kim, Kyuyoung, et al.
Pubblicazione: (2024)
di: Kim, Kyuyoung, et al.
Pubblicazione: (2024)
Beyond Laplace and Gaussian: Exploring the Generalized Gaussian Mechanism for Private Machine Learning
di: Rinberg, Roy, et al.
Pubblicazione: (2025)
di: Rinberg, Roy, et al.
Pubblicazione: (2025)
ImpNet: Imperceptible and blackbox-undetectable backdoors in compiled neural networks
di: Clifford, Eleanor, et al.
Pubblicazione: (2022)
di: Clifford, Eleanor, et al.
Pubblicazione: (2022)
When Vision Fails: Text Attacks Against ViT and OCR
di: Boucher, Nicholas, et al.
Pubblicazione: (2023)
di: Boucher, Nicholas, et al.
Pubblicazione: (2023)
No, of Course I Can! Deeper Fine-Tuning Attacks That Bypass Token-Level Safety Mechanisms
di: Kazdan, Joshua, et al.
Pubblicazione: (2025)
di: Kazdan, Joshua, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Buffer Overflow in Mixture of Experts
di: Hayes, Jamie, et al.
Pubblicazione: (2024) -
Interpreting the Repeated Token Phenomenon in Large Language Models
di: Yona, Itay, et al.
Pubblicazione: (2025) -
Beyond Slow Signs in High-fidelity Model Extraction
di: Foerster, Hanna, et al.
Pubblicazione: (2024) -
Measuring memorization in RLHF for code completion
di: Pappu, Aneesh, et al.
Pubblicazione: (2024) -
Cascading Adversarial Bias from Injection to Distillation in Language Models
di: Chaudhari, Harsh, et al.
Pubblicazione: (2025)