X (Twitter), @DimitrisPapailiopoulos (handle truncated in UI)
— saved image
Dimitris Papailiopo... ✔️ @DimitrisPa... · 5h
I'm 30% in verifying this, and as I am trying to understand Chat's proofs for this particular problem, I have noticed a few interesting things
1) zero mathematical mistakes so far.
2) When GPT Pro says something is correct I trust it more than I trust myself using Lean
3) the exposition is a disaster
- a. A very complicated tree of variable names. Say at some point in a proof you need to bound Pr(-A<||w||+||h||<A), the model renames the norms to say R1 and R2, their ratio R1/R2 to rho, and then it decides to bound |rho/A-1| instead while you have to keep track of like a series of variable renamings. So exhausting!
-b. the ordering of technical lemmas needed is very random, Eg technical facts don't show up where you need them. In a reasonable exposition you'd expect a series of lemmas etc that when stated let you arrive at the final final result for which you'd need to set a bunch of "parameters" for things to click in. In Chat's proofs Everything shows up whenever the model felt like stating them. there's no narrative arc, just a correct pile of implications.
[quoted tweet:]
Dimitris Papailiop... ✔️ @DimitrisP... · Aug 2
I feel a weird guilt that I am the first to experience the beauty of the produced result, while minds far stronger than mine have spent far longer time to answer the same question that Chat and Fable destroyed in less than an hour ...
[screenshot excerpt below, task-list style:]
Calibrating threshold analysis with negligible quadratic terms.
Reconciling single-flip and pair-flip failure probabilities in threshold analysis.
Reconciling pair-flip probabilities with empirical observations.
Architecting proof structure and lemma dependencies for rigorous completion.
Architecting multi-regime MGF bounds and optimizing variational transitions.Note from Claude Sonnet 5
Tweet thread from mathematician Dimitris Papailiopoulos describing verification of an AI-generated math proof (referring to 'Chat' i.e. GPT and 'Fable', an AI model), praising correctness but criticizing exposition quality (confusing variable renaming, no narrative arc to the lemmas).