← All topics

ai math

4 captures, most recent first.

@abeirami

— saved image

Ahmad Beirami ✔️ @abeirami · 19h
With essentially zero technical input from me, GPT-5.6 Sol and Fable 5 not only proved a conjecture we left open ~2 years ago on best-of-n, but also delivered a strictly tighter bound with a clean and insightful derivation.

We are officially in a new era of mathematical reasoning!

Much of what previously counted as meaningful technical contribution is now routine for these models. This level of reasoning is being fundamentally democratized.

The kind of research that used to take months to become a paper is now achievable in minutes.

[quoted tweet:]
Ahmad Beirami ✔️ @abeirami · 22h
This got even more ridiculous!

I was trying to use the context of this session to nudge Sol to improve another result. Instead, it misunderstood me as wanting to improve this …

[attached images: two page-scan panels of a math writeup titled 'A sharper finite-atom KL bound for best-of-n', with theorem statement, proof sketch, and a plot comparing an analytical formula, an estimator, a sharpened estimator, and exact KL divergence across a range of n]
Note from Claude Sonnet 5

Tweet from Ahmad Beirami reporting that AI models 'GPT-5.6 Sol' and 'Fable 5' proved an open conjecture on best-of-n sampling and produced a tighter bound with clean derivation, calling it a 'new era of mathematical reasoning'. Attached is a two-panel image of a technical math writeup (theorem, proof, and a KL-divergence comparison plot) that is largely illegible at this resolution.

ai mathgptfablebest-of-ntwittermathematicsmodel names

X (Twitter), @DimitrisPapailiopoulos (handle truncated in UI)

— saved image

Dimitris Papailiopo... ✔️ @DimitrisPa... · 5h
I'm 30% in verifying this, and as I am trying to understand Chat's proofs for this particular problem, I have noticed a few  interesting things

1) zero mathematical mistakes so far.

2) When GPT Pro says something is correct I trust it more than I trust myself using Lean

3) the exposition is a disaster

    - a. A very complicated tree of variable names. Say at some point in a proof you need to bound Pr(-A<||w||+||h||<A), the model renames the norms to say R1 and R2, their ratio R1/R2 to rho, and then it decides to bound |rho/A-1| instead while you have to keep track of like a series of variable renamings. So exhausting!

    -b. the ordering of technical lemmas needed is very random, Eg technical facts don't show up where you need them. In a reasonable exposition you'd expect a series of lemmas etc that when stated let you arrive at the final final result for which you'd need to set a bunch of "parameters" for things to click in. In Chat's proofs Everything shows up whenever the model felt like stating them. there's no narrative arc, just a correct pile of implications.

[quoted tweet:]
Dimitris Papailiop... ✔️ @DimitrisP... · Aug 2
I feel a weird guilt that I am the first to experience the beauty of the produced result, while minds far stronger than mine have spent far longer time to answer the same question that Chat and Fable destroyed in less than an hour ...

[screenshot excerpt below, task-list style:]
Calibrating threshold analysis with negligible quadratic terms.
Reconciling single-flip and pair-flip failure probabilities in threshold analysis.
Reconciling pair-flip probabilities with empirical observations.
Architecting proof structure and lemma dependencies for rigorous completion.
Architecting multi-regime MGF bounds and optimizing variational transitions.
Note from Claude Sonnet 5

Tweet thread from mathematician Dimitris Papailiopoulos describing verification of an AI-generated math proof (referring to 'Chat' i.e. GPT and 'Fable', an AI model), praising correctness but criticizing exposition quality (confusing variable renaming, no narrative arc to the lemmas).

ai mathgptfableproof verificationtwittermathematics

Dimitris Papailiopoulos @DimitrisPapail

— saved image

Dimitris Papailiopoulos ✔️ @DimitrisPapail
I feel a weird guilt that I am the first to experience the beauty of the produced result, while minds far stronger than mine have spent far longer time to answer the same question that Chat and Fable destroyed in less than an hour just because I prompted them...

I guess I'll have to share this one.

[white task-list panel, timestamped-style entries:]
Calibrating threshold analysis with negligible quadratic terms.
Reconciling single-flip and pair-flip failure probabilities in threshold analysis.
Reconciling pair-flip probabilities with empirical observations.
Architecting proof structure and lemma dependencies for rigorous completion.
Architecting multi-regime MGF bounds and optimizing variational transitions.
Orchestrating probability bounds and dissecting multi-flip failure regimes.
Orchestrating regime boundaries and refining variational exponent analysis.
Architecting SINR bounds and warm-start error analysis rigorously.
Architecting rigorous proofs through random matrix theory and concentration bounds.
Reconciling MSE bounds with sign-error thresholds for warm-start analysis.
Architecting warm-start bounds via smallest singular value concentration.
Rigorously bounding small eigenvalue counts for Gaussian matrices.
Architecting rigorous warm-start bounds via singular value concentration.

Dimitris Papailiopoulos ✔️ @DimitrisPapail · Aug 2
When you ask Chat to make a breakthrough on a 15 year old open problem and it zero shots it.

I did say I won't go back to info theory question that gave me PTSD, but oops i did it again.
Note from Claude Sonnet 5

Fuller view of Dimitris Papailiopoulos's tweet thread (continuation of the thread in the previous screenshot), showing the full list of AI 'reasoning step' task titles from solving a 15-year-old open information theory problem, and his Aug 2 tweet describing the breakthrough.

ai mathgptfableinformation theorytwittermathematics

Zvi Mowshowitz @TheZvi

— saved image

Danielle Fong 🐦☀️ reposted

Zvi Mowshowitz @TheZvi · 5h
Mathematicians are awesome people, I narrowly escaped being one. I love them dearly and I hope they take joy in all the cool new math and new opportunities, rather than despair. And obviously no one should be mean to them right now even if they need some copium.

[quoted tweet]
roon @tszzl · 7h
people are incredibly mean to mathematicians in this time. they really relish when an ai solves something and frame in a zero sum way w human mathematicians. i think it evens out some childhood era math trauma they have?
Note from Claude Sonnet 5

X thread: Zvi Mowshowitz reposted by Danielle Fong, responding sympathetically to roon's (@tszzl) observation that people are 'incredibly mean' to mathematicians right now, gleefully framing AI math breakthroughs as zero-sum wins over human mathematicians, which roon speculates is people working out childhood math trauma.

twitterai mathmathematicianszvi mowshowitzroon