GRP-Obliteration: Unaligning LLMs With a Single Unlabeled Prompt
Fuente:
arXiv
Saved in:
| Main Authors: | Russinovich, Mark, Cai, Yanan, Hines, Keegan, Severi, Giorgio, Bullwinkel, Blake, Salem, Ahmed |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Obliviate: Efficient Unmemorization for Protecting Intellectual Property in Large Language Models
by: Russinovich, Mark, et al.
Published: (2025)
by: Russinovich, Mark, et al.
Published: (2025)
A Representation Engineering Perspective on the Effectiveness of Multi-Turn Jailbreaks
by: Bullwinkel, Blake, et al.
Published: (2025)
by: Bullwinkel, Blake, et al.
Published: (2025)
LogiPlan: A Structured Benchmark for Logical Planning and Relational Reasoning in LLMs
by: Cai, Yanan, et al.
Published: (2025)
by: Cai, Yanan, et al.
Published: (2025)
The Trigger in the Haystack: Extracting and Reconstructing LLM Backdoor Triggers
by: Bullwinkel, Blake, et al.
Published: (2026)
by: Bullwinkel, Blake, et al.
Published: (2026)
Jailbreaking is (Mostly) Simpler Than You Think
by: Russinovich, Mark, et al.
Published: (2025)
by: Russinovich, Mark, et al.
Published: (2025)
Hey, That's My Model! Introducing Chain & Hash, An LLM Fingerprinting Technique
by: Russinovich, Mark, et al.
Published: (2024)
by: Russinovich, Mark, et al.
Published: (2024)
Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack
by: Russinovich, Mark, et al.
Published: (2024)
by: Russinovich, Mark, et al.
Published: (2024)
Stateless Yet Not Forgetful: Implicit Memory as a Hidden Channel in LLMs
by: Salem, Ahmed, et al.
Published: (2026)
by: Salem, Ahmed, et al.
Published: (2026)
Are My Optimized Prompts Compromised? Exploring Vulnerabilities of LLM-based Optimizers
by: Zhao, Andrew, et al.
Published: (2025)
by: Zhao, Andrew, et al.
Published: (2025)
Realizing Unaligned Block-wise Pruning for DNN Acceleration on Mobile Devices
by: Lee, Hayun, et al.
Published: (2024)
by: Lee, Hayun, et al.
Published: (2024)
Deep Positive-Unlabeled Anomaly Detection for Contaminated Unlabeled Data
by: Takahashi, Hiroshi, et al.
Published: (2024)
by: Takahashi, Hiroshi, et al.
Published: (2024)
Understanding the Effects of Safety Unalignment on Large Language Models
by: Halloran, John T.
Published: (2026)
by: Halloran, John T.
Published: (2026)
Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization
by: Razin, Noam, et al.
Published: (2024)
by: Razin, Noam, et al.
Published: (2024)
Content-Style Learning from Unaligned Domains: Identifiability under Unknown Latent Dimensions
by: Shrestha, Sagar, et al.
Published: (2024)
by: Shrestha, Sagar, et al.
Published: (2024)
GED-Consistent Disentanglement of Aligned and Unaligned Substructures for Graph Similarity Learning
by: Zhan, Zhentao, et al.
Published: (2025)
by: Zhan, Zhentao, et al.
Published: (2025)
Proto-EVFL: Enhanced Vertical Federated Learning via Dual Prototype with Extremely Unaligned Data
by: Guo, Wei, et al.
Published: (2025)
by: Guo, Wei, et al.
Published: (2025)
Positive Unlabeled Contrastive Learning
by: Acharya, Anish, et al.
Published: (2022)
by: Acharya, Anish, et al.
Published: (2022)
Collaborative Unlabeled Data Optimization
by: Shang, Xinyi, et al.
Published: (2025)
by: Shang, Xinyi, et al.
Published: (2025)
Augmenting Offline RL with Unlabeled Data
by: Wang, Zhao, et al.
Published: (2024)
by: Wang, Zhao, et al.
Published: (2024)
Incorporating Unlabelled Data into Bayesian Neural Networks
by: Sharma, Mrinank, et al.
Published: (2023)
by: Sharma, Mrinank, et al.
Published: (2023)
Angular Regularization for Positive-Unlabeled Learning on the Hypersphere
by: Sevetlidis, Vasileios, et al.
Published: (2025)
by: Sevetlidis, Vasileios, et al.
Published: (2025)
Can ChatGPT Learn My Life From a Week of First-Person Video?
by: Harris, Keegan
Published: (2025)
by: Harris, Keegan
Published: (2025)
Inconsistency-Aware Minimization: Improving Generalization with Unlabeled Data
by: Kim, Hee-Sung, et al.
Published: (2026)
by: Kim, Hee-Sung, et al.
Published: (2026)
Noisy-Pair Robust Representation Alignment for Positive-Unlabeled Learning
by: Zhao, Hengwei, et al.
Published: (2025)
by: Zhao, Hengwei, et al.
Published: (2025)
Learning from M-Tuple Dominant Positive and Unlabeled Data
by: Qin, Jiahe, et al.
Published: (2025)
by: Qin, Jiahe, et al.
Published: (2025)
Positive-Unlabeled Diffusion Models for Preventing Sensitive Data Generation
by: Takahashi, Hiroshi, et al.
Published: (2025)
by: Takahashi, Hiroshi, et al.
Published: (2025)
Time-Prompt: Integrated Heterogeneous Prompts for Unlocking LLMs in Time Series Forecasting
by: Wang, Zesen, et al.
Published: (2025)
by: Wang, Zesen, et al.
Published: (2025)
CAP: Controllable Alignment Prompting for Unlearning in LLMs
by: Wang, Zhaokun, et al.
Published: (2026)
by: Wang, Zhaokun, et al.
Published: (2026)
Should You Use Your Large Language Model to Explore or Exploit?
by: Harris, Keegan, et al.
Published: (2025)
by: Harris, Keegan, et al.
Published: (2025)
Tabular Data Adapters: Improving Outlier Detection for Unlabeled Private Data
by: Herurkar, Dayananda, et al.
Published: (2025)
by: Herurkar, Dayananda, et al.
Published: (2025)
Leveraging Skills from Unlabeled Prior Data for Efficient Online Exploration
by: Wilcoxson, Max, et al.
Published: (2024)
by: Wilcoxson, Max, et al.
Published: (2024)
AllMatch: Exploiting All Unlabeled Data for Semi-Supervised Learning
by: Wu, Zhiyu, et al.
Published: (2024)
by: Wu, Zhiyu, et al.
Published: (2024)
Candidate Pseudolabel Learning: Enhancing Vision-Language Models by Prompt Tuning with Unlabeled Data
by: Zhang, Jiahan, et al.
Published: (2024)
by: Zhang, Jiahan, et al.
Published: (2024)
Evolutionary System Prompt Learning for Reinforcement Learning in LLMs
by: Zhang, Lunjun, et al.
Published: (2026)
by: Zhang, Lunjun, et al.
Published: (2026)
Red-teaming Activation Probes using Prompted LLMs
by: Blandfort, Phil, et al.
Published: (2025)
by: Blandfort, Phil, et al.
Published: (2025)
Drag-and-Drop LLMs: Zero-Shot Prompt-to-Weights
by: Liang, Zhiyuan, et al.
Published: (2025)
by: Liang, Zhiyuan, et al.
Published: (2025)
Dark LLMs: The Growing Threat of Unaligned AI Models
by: Fire, Michael, et al.
Published: (2025)
by: Fire, Michael, et al.
Published: (2025)
Resilient Class-Incremental Learning: on the Interplay of Drifting, Unlabelled and Imbalanced Data Streams
by: Li, Jin, et al.
Published: (2026)
by: Li, Jin, et al.
Published: (2026)
Integrating Distribution Matching into Semi-Supervised Contrastive Learning for Labeled and Unlabeled Data
by: Nakayama, Shogo, et al.
Published: (2026)
by: Nakayama, Shogo, et al.
Published: (2026)
Cost-Sensitive Unbiased Risk Estimation for Multi-Class Positive-Unlabeled Learning
by: Zhang, Miao, et al.
Published: (2025)
by: Zhang, Miao, et al.
Published: (2025)
Similar Items
-
Obliviate: Efficient Unmemorization for Protecting Intellectual Property in Large Language Models
by: Russinovich, Mark, et al.
Published: (2025) -
A Representation Engineering Perspective on the Effectiveness of Multi-Turn Jailbreaks
by: Bullwinkel, Blake, et al.
Published: (2025) -
LogiPlan: A Structured Benchmark for Logical Planning and Relational Reasoning in LLMs
by: Cai, Yanan, et al.
Published: (2025) -
The Trigger in the Haystack: Extracting and Reconstructing LLM Backdoor Triggers
by: Bullwinkel, Blake, et al.
Published: (2026) -
Jailbreaking is (Mostly) Simpler Than You Think
by: Russinovich, Mark, et al.
Published: (2025)