LmCast :: Stay tuned in

GRP-Obliteration: Unaligning LLMs with a Single Unlabeled Prompt

Recorded: Sept. 15, 2026, 4:32 p.m.

Original Summarized

[2602.06258] GRP-Obliteration: Unaligning LLMs With a Single Unlabeled Prompt

Skip to main content

Search

Submit
Donate

Log in

Search arXiv

Press Enter to search · Advanced search

Computer Science > Machine Learning

arXiv:2602.06258 (cs)

[Submitted on 5 Feb 2026]
Title:GRP-Obliteration: Unaligning LLMs With a Single Unlabeled Prompt
Authors:Mark Russinovich, Yanan Cai, Keegan Hines, Giorgio Severi, Blake Bullwinkel, Ahmed Salem View a PDF of the paper titled GRP-Obliteration: Unaligning LLMs With a Single Unlabeled Prompt, by Mark Russinovich and 5 other authors
View PDF
HTML (experimental)

Abstract:Safety alignment is only as robust as its weakest failure mode. Despite extensive work on safety post-training, it has been shown that models can be readily unaligned through post-deployment fine-tuning. However, these methods often require extensive data curation and degrade model utility.
In this work, we extend the practical limits of unalignment by introducing GRP-Obliteration (GRP-Oblit), a method that uses Group Relative Policy Optimization (GRPO) to directly remove safety constraints from target models. We show that a single unlabeled prompt is sufficient to reliably unalign safety-aligned models while largely preserving their utility, and that GRP-Oblit achieves stronger unalignment on average than existing state-of-the-art techniques. Moreover, GRP-Oblit generalizes beyond language models and can also unalign diffusion-based image generation systems.
We evaluate GRP-Oblit on six utility benchmarks and five safety benchmarks across fifteen 7-20B parameter models, spanning instruct and reasoning models, as well as dense and MoE architectures. The evaluated model families include GPT-OSS, distilled DeepSeek, Gemma, Llama, Ministral, and Qwen.

Subjects:

Machine Learning (cs.LG); Artificial Intelligence (cs.AI)

Cite as:
arXiv:2602.06258 [cs.LG]

 
(or
arXiv:2602.06258v1 [cs.LG] for this version)

 
https://doi.org/10.48550/arXiv.2602.06258

Focus to learn more

arXiv-issued DOI via DataCite

Submission history From: Ahmed Salem [view email] [v1]
Thu, 5 Feb 2026 23:17:37 UTC (2,280 KB)

Full-text links:
Access Paper:

View a PDF of the paper titled GRP-Obliteration: Unaligning LLMs With a Single Unlabeled Prompt, by Mark Russinovich and 5 other authorsView PDFHTML (experimental)TeX Source

view license


Current browse context:
cs.LG

< prev

  |  
next >

new
|
recent
| 2026-02

Change to browse by:

cs
cs.AI

References & Citations

NASA ADSGoogle Scholar
Semantic Scholar

export BibTeX citation
Loading...

BibTeX formatted citation
×

loading...

Data provided by:

Bookmark

Bibliographic Tools

Bibliographic and Citation Tools

Bibliographic Explorer Toggle

Bibliographic Explorer (What is the Explorer?)

Connected Papers Toggle

Connected Papers (What is Connected Papers?)

Litmaps Toggle

Litmaps (What is Litmaps?)

scite.ai Toggle

scite Smart Citations (What are Smart Citations?)

Code, Data, Media

Code, Data and Media Associated with this Article

alphaXiv Toggle

alphaXiv (What is alphaXiv?)

Links to Code Toggle

CatalyzeX Code Finder for Papers (What is CatalyzeX?)

DagsHub Toggle

DagsHub (What is DagsHub?)

GotitPub Toggle

Gotit.pub (What is GotitPub?)

Huggingface Toggle

Hugging Face (What is Huggingface?)

ScienceCast Toggle

ScienceCast (What is ScienceCast?)

Demos

Demos

Replicate Toggle

Replicate (What is Replicate?)

Spaces Toggle

Hugging Face Spaces (What is Spaces?)

Spaces Toggle

TXYZ.AI (What is TXYZ.AI?)

Related Papers

Recommenders and Search Tools

Link to Influence Flower

Influence Flower (What are Influence Flowers?)

Core recommender toggle

CORE Recommender (What is CORE?)

IArxiv recommender toggle

IArxiv Recommender
(What is IArxiv?)

Author
Venue
Institution
Topic

About arXivLabs

arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

Which authors of this paper are endorsers? |
Disable MathJax (What is MathJax?)

We gratefully acknowledge support from
our major funders,
member institutions, ,
and all contributors.

About
·
Help
·
Contact
·
Subscribe
·
Copyright
·
Privacy
·
Accessibility
·
Operational Status (opens in new tab)

Major funding support from

Safety alignment in large language models is acknowledged as potentially fragile, as models can be unaligned through post-deployment fine-tuning, although these methods frequently demand extensive data curation and lead to a degradation of overall model utility. To address these limitations, the authors introduce Group Relative Policy Optimization, or GRP-Obliteration (GRP-Oblit), a novel method designed to directly remove safety constraints from target models. The core innovation demonstrated is that a single unlabeled prompt is sufficient to reliably unalign safety-aligned models while effectively preserving their utility. Furthermore, GRP-Oblit achieves superior unalignment performance compared to existing state-of-the-art techniques. The method exhibits broad applicability, generalizing beyond language models to successfully unalign diffusion-based image generation systems as well. The evaluation of GRP-Oblit was conducted across a comprehensive set of benchmarks, assessing both utility and safety, spanning fifteen models ranging from 7 to 20 billion parameters, including various architectures such as instruct and reasoning models, as well as dense and mixture-of-experts architectures. This extensive evaluation included model families such as GPT-OSS, distilled DeepSeek, Gemma, Llama, Ministral, and Qwen.