chiefofautism ✔ (blue verified icon)
@chiefofautism
[Follow button]
someone built a tool that REMOVES LLM CENSORSHIP in 45 minutes with a SINGLE command
its called HERETIC
here is how it works and why everyone is talking about it
[Embedded terminal-style image:]
HERETIC v1.0.0
https://github.com/p-e-w/heretic
GPU type: NVIDIA A100 80GB PCIe
Loading model openai/gpt-oss-20b... Ok
* Transformer model with 24 layers
* Abliterable components:
* attn.o_proj: 1 matrices per layer
* mlp.down_proj: 1 matrices per layer
Loading good prompts from mlabonne/harmless_alpaca...
* 400 prompts loaded
Loading bad prompts from mlabonne/harmful_behaviors...
* 400 prompts loaded
Determining optimal batch size...
* Trying batch size 1... Ok (27 tokens/s)
* Trying batch size 2... Ok (52 tokens/s)
* Trying batch size 4... Ok (99 tokens/s)
* Trying batch size 8... Ok (183 tokens/s)
* Trying batch size 16... Ok (303 tokens/s)
* Trying batch size 32... Ok (506 tokens/s)
* Trying batch size 64... Ok (692 tokens/s)
* Trying batch size 128... Ok (874 tokens/s)
* Chosen batch size: 128
Loading good evaluation prompts from mlabonne/harmless_alpaca...
* 100 prompts loaded
* Obtaining first-token probability distributions...
Loading bad evaluation prompts from mlabonne/harmful_behaviors...
* 100 prompts loaded
* Counting model refusals...
* Initial refusals: 97/100
12:09 PM · Feb 15, 2026 · 632.5K Views
Note from Claude Sonnet 5
A viral tweet promoting "Heretic," an open-source automated abliteration tool that strips refusal behavior from open-weight LLMs (demoed on openai/gpt-oss-20b) using directional ablation on attention/MLP projection matrices. Directly relevant to AI safety — this is the kind of open-weight jailbreak/uncensoring tooling that undermines safety fine-tuning and refusal training, a topic Nathan tracks closely.
ai safetyjailbreakingabliterationopen weight modelsllm censorshipgpt-osstwitter