← Timeline

@lanyon_ai

@lanyon_ai on X

1 capture, most recent first. Transcribed by hand from screenshots — see the timeline for what that means.

@lanyon_ai

— saved image

Our second official benchmarking post is out! The Euler equations may *seem* easy to solve using finite volume methods, but all frontier models (including GPT-5.6 Sol, Fable 5, and Kimi K3) consistently introduce both subtle and unsubtle errors, including numerical oscillations, thermodynamic inconsistencies, and incorrect orders of accuracy. That is, if the code even works at all. Mathematical misformalizations abound, and token costs can easily hit tens of dollars per attempt.

Only Lanyon's neurosymbolic architecture is consistently able to produce robust solvers with end-to-end proofs of correctness, and it does so with costs that are >100x lower. Post below 👇
Note from Claude Sonnet 5

Tweet from Lanyon AI (@lanyon_ai) promoting a benchmarking post comparing frontier LLMs against their neurosymbolic architecture on Euler-equation finite-volume solver generation. Below the text are two density-contour plots of a 2-D Riemann problem (t=0.8, 800x800 grid) comparing a flawed solution against a 2nd-order minmod wave-propagation (Rusanov) solution, each with a colorbar.

ai benchmarkingneurosymbolic ainumerical methodsllm evaluation