← All topics

john schulman

2 captures, most recent first.

deckard @slimer48484

quoting @natolambert — saved image

deckard @slimer48484 . 5h
OpenAI is apparently learning this by making the mistake.

Theres people who learn by pissing on the electric fence and there's people who learn by watching someone else get shocked.

As the stakes rise let's try to exercise some caution ahead of time.

[Quoted tweet:]
Nathan Lambert @natolambert . 16h
Replying to @natolambert and @johnschulman2
But fwiw running evals during training vs offline should have similar monitoring
Note from Claude Sonnet 5

Tweet by @slimer48484 commenting on an unspecified OpenAI mistake (implied to be a safety/monitoring lapse), using a pissing-on-the-electric-fence metaphor to argue for exercising caution ahead of time, quote-tweeting Nathan Lambert's reply to John Schulman about eval monitoring during vs. offline training.

ai safetyopenaievalsmonitoringnathan lambertjohn schulman

X (Twitter) — John Schulman (@johnschulm...)

John Schul... ✓ [Thinking Machines badge] (@johnschulm...) — Jun 17 PPO had a second wave in the LLM era for reasons unanticipated by the original paper - the importance-ratio objective fixes biases from numeric error, async training, and forward pass noise - the clipping objective affects entropy through a mechanism that we didn't know about at the time of publication (DAPO, arxiv.org/abs/2509.26114) > QUOTED: John Schul... ✓ [Thinking Machines badge] (@johnschulm...) — Jun 17 > PPO: rejected from NIPS 2017
Note from Claude Sonnet 5

A technical tweet by John Schulman (PPO co-author, now at Thinking Machines) reflecting on why the PPO algorithm found unexpected renewed relevance in LLM-era RL training, quoting his own earlier tweet noting PPO's original NeurIPS rejection.

reinforcement learningppollm trainingjohn schulmanmachine learning research