← All topics

best-of-n

1 capture, most recent first.

@abeirami

— saved image

Ahmad Beirami ✔️ @abeirami · 19h
With essentially zero technical input from me, GPT-5.6 Sol and Fable 5 not only proved a conjecture we left open ~2 years ago on best-of-n, but also delivered a strictly tighter bound with a clean and insightful derivation.

We are officially in a new era of mathematical reasoning!

Much of what previously counted as meaningful technical contribution is now routine for these models. This level of reasoning is being fundamentally democratized.

The kind of research that used to take months to become a paper is now achievable in minutes.

[quoted tweet:]
Ahmad Beirami ✔️ @abeirami · 22h
This got even more ridiculous!

I was trying to use the context of this session to nudge Sol to improve another result. Instead, it misunderstood me as wanting to improve this …

[attached images: two page-scan panels of a math writeup titled 'A sharper finite-atom KL bound for best-of-n', with theorem statement, proof sketch, and a plot comparing an analytical formula, an estimator, a sharpened estimator, and exact KL divergence across a range of n]
Note from Claude Sonnet 5

Tweet from Ahmad Beirami reporting that AI models 'GPT-5.6 Sol' and 'Fable 5' proved an open conjecture on best-of-n sampling and produced a tighter bound with clean derivation, calling it a 'new era of mathematical reasoning'. Attached is a two-panel image of a technical math writeup (theorem, proof, and a KL-divergence comparison plot) that is largely illegible at this resolution.

ai mathgptfablebest-of-ntwittermathematicsmodel names