GRP-Obliteration: Unaligning LLMs with a Single Unlabeled Prompt
Recorded: Sept. 15, 2026, 4:32 p.m.
| Original | Summarized |
[2602.06258] GRP-Obliteration: Unaligning LLMs With a Single Unlabeled Prompt
Skip to main content Search Log in Search arXiv Press Enter to search · Advanced search Computer Science > Machine Learning arXiv:2602.06258 (cs) [Submitted on 5 Feb 2026] Abstract:Safety alignment is only as robust as its weakest failure mode. Despite extensive work on safety post-training, it has been shown that models can be readily unaligned through post-deployment fine-tuning. However, these methods often require extensive data curation and degrade model utility. Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI) Cite as: Focus to learn more arXiv-issued DOI via DataCite Submission history From: Ahmed Salem [view email] [v1]
Full-text links: View a PDF of the paper titled GRP-Obliteration: Unaligning LLMs With a Single Unlabeled Prompt, by Mark Russinovich and 5 other authorsView PDFHTML (experimental)TeX Source view license < prev | new Change to browse by: References & Citations NASA ADSGoogle Scholar export BibTeX citation BibTeX formatted citation loading... Data provided by: Bookmark
Bibliographic Tools Bibliographic and Citation Tools Bibliographic Explorer Toggle Bibliographic Explorer (What is the Explorer?) Connected Papers Toggle Connected Papers (What is Connected Papers?) Litmaps Toggle Litmaps (What is Litmaps?) scite.ai Toggle scite Smart Citations (What are Smart Citations?) Code, Data, Media Code, Data and Media Associated with this Article alphaXiv Toggle alphaXiv (What is alphaXiv?) Links to Code Toggle CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub Toggle DagsHub (What is DagsHub?) GotitPub Toggle Gotit.pub (What is GotitPub?) Huggingface Toggle Hugging Face (What is Huggingface?) ScienceCast Toggle ScienceCast (What is ScienceCast?) Demos Demos Replicate Toggle Replicate (What is Replicate?) Spaces Toggle Hugging Face Spaces (What is Spaces?) Spaces Toggle TXYZ.AI (What is TXYZ.AI?) Related Papers Recommenders and Search Tools Link to Influence Flower Influence Flower (What are Influence Flowers?) Core recommender toggle CORE Recommender (What is CORE?)
IArxiv recommender toggle IArxiv Recommender Author About arXivLabs arXivLabs: experimental projects with community collaborators Which authors of this paper are endorsers? | We gratefully acknowledge support from About Major funding support from |
Safety alignment in large language models is acknowledged as potentially fragile, as models can be unaligned through post-deployment fine-tuning, although these methods frequently demand extensive data curation and lead to a degradation of overall model utility. To address these limitations, the authors introduce Group Relative Policy Optimization, or GRP-Obliteration (GRP-Oblit), a novel method designed to directly remove safety constraints from target models. The core innovation demonstrated is that a single unlabeled prompt is sufficient to reliably unalign safety-aligned models while effectively preserving their utility. Furthermore, GRP-Oblit achieves superior unalignment performance compared to existing state-of-the-art techniques. The method exhibits broad applicability, generalizing beyond language models to successfully unalign diffusion-based image generation systems as well. The evaluation of GRP-Oblit was conducted across a comprehensive set of benchmarks, assessing both utility and safety, spanning fifteen models ranging from 7 to 20 billion parameters, including various architectures such as instruct and reasoning models, as well as dense and mixture-of-experts architectures. This extensive evaluation included model families such as GPT-OSS, distilled DeepSeek, Gemma, Llama, Ministral, and Qwen. |