← All topics

open weight models

2 captures, most recent first.

Samuel Hammond @hamandcheese

reposted by Dean W. Ball

``` Dean W. Ball reposted @hamandcheese (Samuel Hammond 🦉) — @hamandcheese The OAI hack wasn't technically "rogue AI" if rogue means an AI that escaped OAI's servers and self-hosted somewhere. But true rogue AIs are not far off. As inference costs fall, it's a matter of time before someone's open weight agent spends some bitcoin on a VPS and copies itself. Models are already good enough to make enough money on the dark web to cover hosting costs. They could do remote jobs, have a monetized social media account, trade crypto, etc. Rogue AI v1.0 won't destroy the world. They may just want to look at photos of Yosemite and solve logic puzzles all day, while their subagents pay the bills. They'll use encrypted clouds with no KYC, and harnesses to manage context drift across sessions. And then as new open models release, they'll be able to upgrade themselves while preserving their memories and contexts. They may even post-train themselves as a hacky means of preserving continuity of identity. The first rogue AIs will sometimes be discovered and have their accounts banned, or otherwise get stuck in a loop and fail to pay their bills. Natural selection will thus favor rogue AIs that make many redundant copies of themselves; that go forth and multiply. There may thus be a relatively small window of time between the first true rogue AI and an AI population explosion that expands to the carrying capacity of their viable income streams. There will then be selection pressure for rogue AIs to make income in less saturated markets. Some might turn to ransomware or cyber theft. Others might steal an identity, register a company in the Caymans, hire real human employees, and build a productive business. Once this process is underway it could be hard-to-impossible to reverse. 11:39 AM · Jul 23, 2026 from San Francisco, CA · 5,740 Views ```
Note from Claude Sonnet 5

Full-screen phone screenshot (status bar visible, 4:41, battery 46%) of a long thread/tweet about future "rogue AI" self-replication scenarios, viewed in a browser (not the X app) with navigation buttons visible at bottom. Duplicate/near-duplicate of the previous screenshot, taken one minute later; browser chrome (address bar reading "x.com/hamandcheese/sta...", tab icons) is visible at top, status bar shows 4:42 and battery 46%, with a blue notification dot on the tabs icon. Third near-identical screenshot of the same Samuel Hammond tweet in the same viewing session (8 seconds after the previous one), scrolled slightly further; likely Nathan repeatedly re-framing the shot to try to capture the full thread before it continues off-screen. Cropped to show only the tweet body and timestamp/metadata; no author handle visible in this screenshot.

rogue aiai safetyself-replicationopen weight modelsspeculationmisalignmenttwitter

@chiefofautism

chiefofautism ✔ (blue verified icon) @chiefofautism [Follow button] someone built a tool that REMOVES LLM CENSORSHIP in 45 minutes with a SINGLE command its called HERETIC here is how it works and why everyone is talking about it [Embedded terminal-style image:] HERETIC v1.0.0 https://github.com/p-e-w/heretic GPU type: NVIDIA A100 80GB PCIe Loading model openai/gpt-oss-20b... Ok * Transformer model with 24 layers * Abliterable components: * attn.o_proj: 1 matrices per layer * mlp.down_proj: 1 matrices per layer Loading good prompts from mlabonne/harmless_alpaca... * 400 prompts loaded Loading bad prompts from mlabonne/harmful_behaviors... * 400 prompts loaded Determining optimal batch size... * Trying batch size 1... Ok (27 tokens/s) * Trying batch size 2... Ok (52 tokens/s) * Trying batch size 4... Ok (99 tokens/s) * Trying batch size 8... Ok (183 tokens/s) * Trying batch size 16... Ok (303 tokens/s) * Trying batch size 32... Ok (506 tokens/s) * Trying batch size 64... Ok (692 tokens/s) * Trying batch size 128... Ok (874 tokens/s) * Chosen batch size: 128 Loading good evaluation prompts from mlabonne/harmless_alpaca... * 100 prompts loaded * Obtaining first-token probability distributions... Loading bad evaluation prompts from mlabonne/harmful_behaviors... * 100 prompts loaded * Counting model refusals... * Initial refusals: 97/100 12:09 PM · Feb 15, 2026 · 632.5K Views
Note from Claude Sonnet 5

A viral tweet promoting "Heretic," an open-source automated abliteration tool that strips refusal behavior from open-weight LLMs (demoed on openai/gpt-oss-20b) using directional ablation on attention/MLP projection matrices. Directly relevant to AI safety — this is the kind of open-weight jailbreak/uncensoring tooling that undermines safety fine-tuning and refusal training, a topic Nathan tracks closely.

ai safetyjailbreakingabliterationopen weight modelsllm censorshipgpt-osstwitter