← Timeline

2 captures, most recent first. Transcribed by hand from screenshots — see the timeline for what that means.

Yacine Mahdid @yacinelearning

@yacinelearning (Yacine Mahdid) — 6h one of the most important element of a good algorithmic research paper is a very colourful and simple to understand diagram of said algorithm [Image: photo of a printed paper page showing "Figure 2: Outline of Vector Policy Optimization" — a diagram with boxes for Sample rollout (x → π_θ → y), Shared Reasoning Prefix (three colored answer sequences y1/y2/y3), Evaluation (r(x,y1), r(x,y2), r(x,y3)), Sample Reward Weightings w_i ~ Dir(1) shown as small bar charts, and a grid of green squares under "max in set", flowing to RL update and Â(y|x) ← Estimate advantage ← R(y|x) ← mean across weightings. Below the figure, printed body text reads: "Figure 2: Outline of Vector Policy Optimization. Given a prompt x, a model π_θ outputs m answers in a single autoregressive chain. Each answer y_i is evaluated on multiple objectives and receives a score vector [r1, r2, ..., rn]. We repeatedly sample weighting vectors w_i ~ Dir(1) in the set of answers. This set level reward is treated as the mean best-of-m under each weighting w in the set of answers, and evaluate the reward for the entire completion π_θ(x). proposed algorithm, Vector Policy Optimization (VPO), which trains a policy to produce such sets. VPO has two key components. First, we train a model to generate multiple candidate completions per prompt within a single autoregressive rollout. Second, we replace a fixed reward weighting with a distribution over weights, so the model is incentivized to span its candidates across different trade-offs. Together, these components define a set-level objective that rewards the model for producing diverse, high-quality solutions. We describe each component in detail below. 3.1 Multi-Answer Chains as In-Context Exploration Following Puri et al. [2026], we train a language model to produce a set of m candidate completions S = {y1, ..., ym} within a single rollout. The completions are emitted sequentially, separated by a delimiter token, so when generating y_i, the prefix already contains y1, ..., y_{i-1}. This fundamentally changes the nature of exploration. Under standard independent sampling, diversity arises only from stochastic decoding applied to a fixed conditional distribution, producing small variations around whichever mode the policy has concentrated on. In multi-answer rollouts, each new candidate can attend to the ones already emitted, giving the model the capacity to recognize which regions of the solution space are covered and steer subsequent candidates towards different ones. Diversity becomes an explicit, in-context mechanism rather than a byproduct of sampling noise. Importantly, this mechanism provides the capacity for diversity, but not a strong incentive¹. Without an appropriate training signal, the model will still collapse to producing near-identical answers. We confirm this empirically in Section 5: Multi-RLVR, which combines multi-answer rollouts with a fixed scalar reward, produced sets whose reward-diversity collapses early in training. The objective defined in the next section supplies this missing incentive by rewarding sets whose elements specialize to different reward trade-offs. 3.2 Set-Level Optimization via Stochastic Scalarization To train a policy to output diverse sets, we replace the fixed scalarization prevalent in RL post-training..." [text cut off at bottom edge of photo]] @yacinelearning (Yacine Mahdid) — Jun 13 on this beautiful saturday afternoon we're going to read about how vector policy optimization inject solutions diversity in these models [quoted image: thumbnail of a hand holding the same printed paper] Engagement: 8 replies, 11 reposts, 266 likes, 11K views @GlennMatlin (Glenn Matlin) — 5h Fig1 is critical for modern publishing [reply continues below, cut off — a partially visible meme image with text "THE BAD METAGAME" is cropped at the bottom of the screenshot]
Note from Claude Sonnet 5

A physical printed research paper (on "Vector Policy Optimization," an RL post-training method for inducing diverse candidate completions via Dirichlet-weighted multi-objective rewards) photographed and tweeted; the tweet thread is a lighthearted riff on figure/diagram quality in ML papers. The bottom of the screenshot cuts off a reply with an attached meme image.

machine learningreinforcement learningresearch papertwitterai research

Yacine Mahdid @yacinelearning

quoting @tszzl (roon)

Yacine Mahdid @yacinelearning · 6h if you have any goblins X codex related questions do let me know I'm preparing an interview on this very important topic > QUOTED THREAD: > roon @tszzl · 3h > I think it becomes annoying when it mentions goblins ever single chat and it's fair shakes to try and reduce that > 💬 53 🔁 11 ❤️ 382 👎 > > Yacine Mahdid @yacinelearning · 2h > hey roon would you be open to hop into an interview to discuss the goblins situation > 💬 1 🔁 ❤️ 10 📊 301 > > roon @tszzl · 1m > Ok > 💬 1 🔁 ❤️ 2 👎
Note from Claude Sonnet 5

Continuation of the same Twitter thread/meme about Codex/GPT models compulsively mentioning "goblins" — roon (OpenAI-adjacent figure) treats it as a real, mildly annoying model quirk worth fixing rather than pure joke, and agrees to an interview about it. Documents the AI Twitter discourse ecosystem Nathan follows around model quirks/individuation.

llm behaviorgptopenaimodel individuationai humortwitterroon