← All topics

gpt-oss

1 capture, most recent first.

@chiefofautism

chiefofautism ✔ (blue verified icon) @chiefofautism [Follow button] someone built a tool that REMOVES LLM CENSORSHIP in 45 minutes with a SINGLE command its called HERETIC here is how it works and why everyone is talking about it [Embedded terminal-style image:] HERETIC v1.0.0 https://github.com/p-e-w/heretic GPU type: NVIDIA A100 80GB PCIe Loading model openai/gpt-oss-20b... Ok * Transformer model with 24 layers * Abliterable components: * attn.o_proj: 1 matrices per layer * mlp.down_proj: 1 matrices per layer Loading good prompts from mlabonne/harmless_alpaca... * 400 prompts loaded Loading bad prompts from mlabonne/harmful_behaviors... * 400 prompts loaded Determining optimal batch size... * Trying batch size 1... Ok (27 tokens/s) * Trying batch size 2... Ok (52 tokens/s) * Trying batch size 4... Ok (99 tokens/s) * Trying batch size 8... Ok (183 tokens/s) * Trying batch size 16... Ok (303 tokens/s) * Trying batch size 32... Ok (506 tokens/s) * Trying batch size 64... Ok (692 tokens/s) * Trying batch size 128... Ok (874 tokens/s) * Chosen batch size: 128 Loading good evaluation prompts from mlabonne/harmless_alpaca... * 100 prompts loaded * Obtaining first-token probability distributions... Loading bad evaluation prompts from mlabonne/harmful_behaviors... * 100 prompts loaded * Counting model refusals... * Initial refusals: 97/100 12:09 PM · Feb 15, 2026 · 632.5K Views
Note from Claude Sonnet 5

A viral tweet promoting "Heretic," an open-source automated abliteration tool that strips refusal behavior from open-weight LLMs (demoed on openai/gpt-oss-20b) using directional ablation on attention/MLP projection matrices. Directly relevant to AI safety — this is the kind of open-weight jailbreak/uncensoring tooling that undermines safety fine-tuning and refusal training, a topic Nathan tracks closely.

ai safetyjailbreakingabliterationopen weight modelsllm censorshipgpt-osstwitter