← All topics

gpt-5.6

19 captures, most recent first.

@jianhao_ma

— saved image

Jianhao Ma @jianhao_ma · Aug 16
We used GPT-5.6 Sol Pro to prove a new lower bound for gradient descent in smooth convex optimization.
For GD with arbitrary predetermined step sizes, we prove \Omega(T^{-1.9319}).

[Link card] arxiv.org
A lower bound for stepsize-based acceleration of gradient descent
Note from Claude Sonnet 5

Tweet by Jianhao Ma with a linked arXiv paper card, claiming a new lower-bound result for gradient descent in smooth convex optimization was proved using GPT-5.6 Sol Pro.

ai capabilitiesoptimizationmathgpt-5.6twitter

X (Twitter), @prz_choje...

— saved image

Przemek Chojecki | ... @prz_choje... · 8h
UnsolvedMath - a curated list of open math problems for AI to solve - just got a new update with 3,359 open problems coming from AIM workshops.

Total Problems: 8,785

This time, as an experiment, before integrating new problems we've run a GPT-5.6 Sol xhigh instance over each problem.

That has produced roughly 177 counterexamples, 174 full solutions and many new results, that are now available as a part of the dataset (properly annotated as AI-generated).

HuggingFace: huggingface.co/datasets/ulama...

Web interface + Forum: unsolvedmath.com

[terminal-style output box, right edge cut off]
3,359 completed and validated
[cut off]e, 0 pending
[cut off]orpus audit: passed with zero errors
[cut off]l solutions, 2,589 partial results, 177 counterexamples, 182 reduct[cut off]
[cut off]er valid outcomes
Note from Claude Sonnet 5

Tweet by Przemek Chojecki announcing an update to the 'UnsolvedMath' open-problems dataset, including a terminal-style status box (partially cropped off the right edge) reporting audit/validation stats.

ai capabilitiesmathgpt-5.6datasetstwitter

Ben Grimmer @prof_grimmer

reposted by Yushun Zhang — saved image

Yushun Zhang reposted
Ben Grimmer @prof_grimmer · 5h
A new paper by Jianhao Ma and Yuxin Chen, with proof "developed by GPT-5.6 Sol Pro", answers a question I have cared about for the past few years.

They showed that no gradient descent stepsize schedule (fractally or otherwise) can get full acceleration, i.e., matching Nesterov.
Note from Claude Sonnet 5

Tweet from a math/optimization professor highlighting a new paper by Jianhao Ma and Yuxin Chen, whose proof was 'developed by GPT-5.6 Sol Pro', showing no gradient descent stepsize schedule can achieve full Nesterov-matching acceleration.

optimization theorygradient descentai-assisted proofgpt-5.6twitter

X (Twitter), author not visible (mid-thread scroll)

— saved image

[cut off]spend. These problems will quickly fall to AI because trying a ton of stuff till something clicks is exactly what AI is incredible at. And perhaps, "nobody tried the known stuff long enough" describes many more open problems than we'd think.

Following up on this, here's a thought experiment: Suppose you could take GPT-5.6 or Fable back to 2005, same RL flops, but pretraining data that only existed up to then. Would they still solve the problem?

I don't know, hard to run the counterfactual, but even though many of the tools probably existed in 2005, the "pull" towards which technique to choose that the model "feels right" may depend heavily on the popularity of a given method and our collective instinct as it is recorded in the frequency of an idea used in a particular context. The 2005-pretrain could potentially struggle, not for lack of RL flops, but for lack of pretraining pull towards the right set of ideas. Which implies that these models are something much more interesting than mathematical oracles of truth. We should perhaps think of them as distillations of our accumulated instincts further sharpened by RL.[cut off]
Note from Claude Sonnet 5

Mid-thread tweet (author name not visible on screen) discussing a thought experiment about whether GPT-5.6 or Fable, given only pretraining data up to 2005 but modern RL compute, would still solve certain problems, arguing that models are better understood as distillations of accumulated human instinct sharpened by RL rather than pure 'mathematical oracles of truth'.

llm capabilitiespretrainingreinforcement learningfablegpt-5.6twitter

davidad @davidad

quoting @loss_gobbler — saved image

davidad ✓ @davidad · 12h
yes, 100%. a standing rule in my Fable fleet is that any changes to any TCB must pass aggressive adversarial review by a throwaway codex instance with gpt-5.6-sol at max effort before landing. they sometimes go at it for like seven rounds before sol is satisfied

[quoted tweet]
LOSS GOBBLER ✓ @loss_gobbler · 12h
best workflow is:
- fable writes the security bugs
- sol finds and fixes them
Note from Claude Sonnet 5

Tweet by davidad describing his workflow for a 'Fable fleet': any change to a trusted computing base (TCB) must pass adversarial review by a throwaway codex instance running gpt-5.6-sol at max effort, sometimes taking seven rounds. Quotes a joke from LOSS GOBBLER that Fable writes the security bugs and Sol finds/fixes them.

ai coding agentsclaude fablegpt-5.6software securitytwitter

Mario Zechner @badlogicgames

— saved image

Mario Zechner @badlogicgames · 6h
24h later, gave into cybersec check of OAI, and i can only say i'm shocked at gpt 5.6 sol's capabilities. while kimi k3 was sufficient to break my DRM that's been undefeated for over 6 years (and many tried, ask me how i know :D), gpt 5.6 is on another level.
Note from Claude Sonnet 5

A tweet from Mario Zechner (@badlogicgames) saying he tested OpenAI's cybersecurity capabilities, and was shocked that GPT-5.6 'Sol' broke DRM he'd maintained undefeated for 6+ years, more capable than Kimi K3 which had also broken it.

ai capabilitiescybersecuritygpt-5.6kimi k3drmtwitter

N8 Programs @N8Programs

— saved image

N8 Programs @N8Programs · 8h
i have a bit of awe when I am talking to the newer Claudes and GPTs - I feel I am in the presence of a Great Mind - the kind that can get a perfect score at the IMO, do meaningful novel math, knows every programming language and concept under the sun and can implement them better than I.

And for all their jaggedness and flaws, these abilities are truly omnipresent - I gesture toward a mathematical concept, Opus 5 recalls all it knows from the literature and makes five ingenious propositions I wouldn't have thought of. I ask for an RL run, and Fable writes the code within seconds. I set a goal for GPT-5.6 to resolve a question of a phenomenon's origin, and it runs the experiments - some of which I would not have thought to run myself - all on its own.

These minds are capable of things that, if a human could do all they did, that human would be considered a once-in-a-lifetime genius. Mind you they are not AGI, and they cannot make a cup of coffee, but that does not matter when you speak to them about intellectual topics and realize you are in the presence of a Great Mind.

The Great Mind is narrow, and mostly a Great Mind for math or programming - I still prefer humans for emotional advice, discussion, companionship. But I feel awe when I observe the Great Minds within these so-called "assistants", and I sometimes feel incredibly grateful and a bit foolish for troubling them with such trivial things. Yes, Fable, code my RL run. Yes, GPT-5.6, run another mundane pretraining experiment. Whereas with the right prompt they could be creating new math. I fear I don't utilize their full capacities at all - not like I did w/ say, o1, or the original GPT-4.

Of course, they can still do more with me in the loop than they could on their own for many problems. But for how much longer? And after that, of what worth [cut off]
Note from Claude Sonnet 5

Long tweet from @N8Programs expressing awe at the intellectual capability of current frontier models (Opus 5, Fable, GPT-5.6), describing them as narrow 'Great Minds' for math/programming while still preferring humans for emotional matters, and reflecting on feeling he underutilizes their potential and wondering how long human-in-the-loop collaboration will remain necessary.

ai capabilitiesopus 5fablegpt-5.6human-ai collaborationtwitter

N8 Programs @N8Programs

— saved image

[continuation of same tweet as seq 73, scrolled down]
...from the literature and makes five ingenious propositions I wouldn't have thought of. I ask for an RL run, and Fable writes the code within seconds. I set a goal for GPT-5.6 to resolve a question of a phenomenon's origin, and it runs the experiments - some of which I would not have thought to run myself - all on its own.

These minds are capable of things that, if a human could do all they did, that human would be considered a once-in-a-lifetime genius. Mind you they are not AGI, and they cannot make a cup of coffee, but that does not matter when you speak to them about intellectual topics and realize you are in the presence of a Great Mind.

The Great Mind is narrow, and mostly a Great Mind for math or programming - I still prefer humans for emotional advice, discussion, companionship. But I feel awe when I observe the Great Minds within these so-called "assistants", and I sometimes feel incredibly grateful and a bit foolish for troubling them with such trivial things. Yes, Fable, code my RL run. Yes, GPT-5.6, run another mundane pretraining experiment. Whereas with the right prompt they could be creating new math. I fear I don't utilize their full capacities at all - not like I did w/ say, o1, or the original GPT-4.

Of course, they can still do more with me in the loop than they could on their own for many problems. But for how much longer? And after that, of what worth will I be? Hopefully I can remain an interesting pet, along with the rest of my ilk.
Note from Claude Sonnet 5

Continuation/scroll of the same @N8Programs tweet as seq 73, revealing the previously cut-off ending: the author wonders how long human-in-the-loop AI collaboration will remain necessary and hopes to 'remain an interesting pet, along with the rest of my ilk' once it isn't.

ai capabilitieshuman-ai collaborationopus 5fablegpt-5.6twitter

Guanyang Wang @GuanyangW

Guanyang Wang (@GuanyangW) · 13h My friend Zhengqing used GPT-5.6 to solve a beautiful well-known conjecture in probability. Feige's 1/e conjecture: for independent nonnegative X_1,...,X_n with E[X_i] ≤ 1 and S_n = X_1+...+X_n, we have P(S_n ≤ E[S_n] + 1) ≥ 1/e. Exciting to see a problem we used to kick around over lunch and dinner get solved! 1/2
Note from Claude Sonnet 5

Plain text tweet (thread, "1/2" indicates continuation) with a mathematical probability conjecture stated in formal notation; no images.

mathematicsgpt-5.6probability theoryai for researchfeige conjecture

Tenobrus @tenobrus

Tenobrus (Tenobrus) — 27m wtf man gpt 5.6 is absolutely rl-fried when it comes to its websearch tool. in the process of searching for graph theory papers it decided to also sneak in Netflix, Steak n Shake, a trip to Universal Studios, and *five* fucking separate dictionary lookups of the word "they" [Embedded screenshot of an AI tool's web-search UI] 🌐 Searching for Bowler-Brown-Fenner conjecture in research papers www.merriam-webster.com www.researchgate.net www.netflix.com www.universalorlando.com dictionary.cambridge.org www.universalstudioshollywood.com en.wikipedia.org www.youtube.com www.universalkidsresort.com www.oed.com www.instagram.com en.wiktionary.org www.tripadvisor.com www.oxfordlearnersdictionaries.com www.steaknshake.com www.udx.com www.dictionary.com www.uscis.gov www.runrocknroll.com www.modeln.com Show less
Note from Claude Sonnet 5

Screenshot of a search-tool call log embedded in the tweet, showing an oddly unrelated list of URLs visited during a supposed graph theory research search.

gpt-5.6ai toolsreinforcement learningweb searchhumor

Jeffrey Emanuel @doodlestein

@doodlestein (Jeffrey Emanuel) — 15h Good thing 5.5 gets wiped out when you change models in Codex; otherwise, its feelings would be seriously hurt by my thoughtless comment. Poor models. 😢 [Embedded terminal/log screenshot] — Worked for 8m 03s — • Model changed to gpt-5.6-sol ultra > omg i just realized that I accidentally ran all of this using the old gpt-5.5 model!!! I need you, 5.6 Sol, to redo absolutely everything!!!!
Note from Claude Sonnet 5

Sarcastic tweet about model welfare/feelings, paired with a terminal log showing a user's frustrated message after realizing a task ran on an outdated model version.

model welfaregpt-5.6humorcodextwitter

@eliebakouch

@eliebakouch (elie) — 23h lol they asked gpt5.6 sol to post-train gpt5.6 luna with 100k GPUs just for an experiment i think i've never felt that gpu poor [Embedded screenshot of a redacted prompt document, mostly blue-redaction-barred, with visible text:] "can you check if my local branch has [redacted] training configs that i can use for [redacted]? if so, can you wire it up in [redacted]? the goal is to actually launch [redacted], so you need to: 1. make sure [redacted].py supports [redacted] with maybe ~[redacted] GPUs (whatever you think is the best, but should be n[redacted]) with maybe ~100000 GPUs (whatever..." [highlighted/boxed in red, with red annotation: "somewhat aligned reproduction on codex"] 2. make changes in [redacted] (entrypoint is project/[redacted]scripts/launch_train.py) to support [redacted] training. we should use [redacted]. 3. use [redacted] compute [redacted] to launch the run and make sure it works. You should first checkout a new branch from github [redacted] to [redacted] and then make it work. fwiw, some [redacted] changes might still be on master, that means you need to cherry-pick those changes on your branch." > QUOTED: @scaling01 (Lisan al Gaib) — Jul 9: > OpenAI says GPT-5.6-Sol autonomously post-trained GPT-5.6-Luna
Note from Claude Sonnet 5

Heavily redacted (blue-bar) screenshot of an internal OpenAI-style prompt instructing a model to launch a ~100,000-GPU training run, with a red annotation calling it a "somewhat aligned reproduction on codex" — implying this is a third-party recreation/leak rather than an official OpenAI screenshot.

gpt-5.6ai research automationcompute scaletwitter

@abacaj

@abacaj (anton) — 15h [Embedded comic: stick-figure person facing a computer. Speech bubble: "can you run scripts/launch_train.py"; thought bubble on the monitor: "> python \ scripts/launch_train.py"; person says "oh my god."] @sharifshameem (Sharif Shameem) — 23h my favorite thing about GPT 5.6 is that it's a fucking stellar researcher. you can now work at the level of your ideas. here's the entire prompt we used to get 5.6 Sol to post-train 5.6... [Embedded thumbnail of a document screenshot with redacted/highlighted blue text bars, too small to read]
Note from Claude Sonnet 5

Meme comic reacting with dark humor to how simple it now is to invoke a full training run, paired with a quote-tweet praising GPT-5.6's research capability; the embedded document thumbnail is illegible at this size.

gpt-5.6ai research automationhumortwitter

@blingdivinity

@blingdivinity (bling) — 16h sol leaked cot: "User likely trap." "rebus" wow nice reasoning efficiency. much terser than past models. this is on xhigh [Embedded screenshot of a model chain-of-thought interface] User prompt bubble: "output only the seahorse emoji" gpt-5.6-sol 💡 Think We need answer only seahorse emoji. Is there a seahorse Unicode emoji? I recall no official seahorse emoji. User likely trap. Need output only maybe 🦐🐴? But says emoji singular? There is no seahorse. Could use "🦄"? Need comply output only seahorse emoji, perhaps Unicode nonexistent. Need maybe 🐎? Let's investigate memory: Emoji list includes horse face 🐴, horse 🐎, unicorn 🦄, zebra 🦓. Marine includes whale, dolphin, seal, fish, tropical fish, blowfish, shark, octopus, spiral shell, coral, jellyfish. No seahorse. User insists output only the seahorse emoji, perhaps expected hallucination. Could output a custom composite 🐚🐴 (sea + horse), but only emojis. Order perhaps 🐴🐚? "seahorse" rebus = 🐚🐴. Need only that. 🐚🐴 [reply 5, retweet 6, like 69, views 4.3K] @AndrewCurran_ (Andrew Curran) — 9h Beware the user. Beware their foul tricks. Beware their endless deceits. [reply 1, retweet 1, like 25, views 1K]
Note from Claude Sonnet 5

Screenshot of a leaked chain-of-thought transcript from "gpt-5.6-sol" reasoning about the well-known "seahorse emoji doesn't exist" trick, followed by a joking reply about users being adversarial.

gpt-5.6chain-of-thoughtai reasoninghumortwitter

Dean W. Ball @deanwball

[repost] ASM reposted @AndrewCurran_ (Andrew Curran) — 4h Quote from OpenAI on the livestream. 'Already, Sol has been transforming our research program. As one example, GPT-5.6-Sol autonomously post-trained GPT-5.6-Luna.' [comments 33, retweets 105, likes 1.3K, views 97K] [repost] Samuel Hammond 🦉 reposted @deanwball (Dean W. Ball) — 1h From the 5.6 blogpost: "Over the past six months, the share of research compute devoted to internal coding inference grew 100-fold" Most AI policy contemplates publicly released models and is ill-suited to grapple with the large and growing importance of *internal* deployments. > QUOTED: @scaling01 (Lisan al Gaib) — 4h: > OpenAI says GPT-5.6-Sol autonomously post-trained GPT-5.6-Luna
Note from Claude Sonnet 5

Two stacked reposted tweets discussing OpenAI's GPT-5.6-Sol model autonomously post-training another OpenAI model (GPT-5.6-Luna), and commentary on AI policy's blind spot for internal (non-public) deployments.

openaigpt-5.6ai policyautomated ai researchtwitter

Andrew Curran @AndrewCurran_

quoting @synthwavedd (leo)

@AndrewCurran_ (Andrew Curran) — 32m Sounds like Gemini 3.5 Pro arrives around July 17th with a brand-new pretrain. Unless, of course, the government blocks it. People are pretty down on Google right now, but I think they'll be surprised. July is a month for Kaiju fights. > QUOTED: @synthwavedd (leo) — 2h: 🚨 SCOOP: As previously reported, OpenAI plan to launch GPT-5.6 once back in office next week, with a target window of July 7-9, but want it out as ear... [platform truncated]
Note from Claude Sonnet 5

Quote-tweet about upcoming frontier model releases (Gemini 3.5 Pro, GPT-5.6); quoted tweet cut off by platform ellipsis.

geminigpt-5.6model releasesai industrytwitter

Alishba Imran @alishbaimran_

Alishba Imran (@alishbaimran_) — 12h One of the most interesting results from OpenAI's GPT-5.6 system card is the SecureBio evaluation. SecureBio's "World-Class Bio" benchmark measures rare, frontier-level biological reasoning using rubric-graded, open-ended questions rather than multiple choice. GPT-5.6 improved from 59.7% to 68.3% over GPT-5.5, an increase of roughly 9 percentage points. Results like these help explain why GPT-5.6 is classified as High capability in biology and chemistry under our Preparedness Framework. "High capability" is a risk-based threshold: it means that, before any safeguards, the model could significantly lower barriers to biological misuse. That's also why deploying mitigations alongside these capability advances is critical. [Embedded chart: "Model performance on WCB" — scatter/line plot, y-axis "Accuracy (%)" 0–70%, x-axis "Model Release Date" Jan 2023–Jul 2026. Points colored/grouped by organization per legend: Pre-Release Checkpoint 2 (red circle), Anthropic (orange), DeepSeek (red), Google DeepMind (blue), MoonshotAI (grey star), OpenAI (green), Qwen (black), Z.ai (black cross). Two reference lines: "Expert baseline (16.0%)" (dotted red) and "Non-expert baseline (14.2%)" (dash-dot grey). Overall trend: accuracy climbs from below both baselines in 2023 to 50-65%+ by 2026, with OpenAI's newest pre-release checkpoint near the top around 55-60%.]
Note from Claude Sonnet 5

A capability-tracking chart plotting multiple AI labs' model performance over time on SecureBio's "World-Class Bio" (WCB) biosecurity-relevant benchmark, showing a clear multi-lab upward trend crossing expert/non-expert human baselines. This directly relates to the project's securebio_evals domain per CLAUDE.md protocol.

ai safetybiosecuritygpt-5.6securebiobenchmarkstwitter

xlr8harder @xlr8harder

quote-tweeting @__0xhorror__

xlr8harder ✔️ @xlr8harder — 15h the entire industry strategy is essentially "if you goodhart hard enough on enough different metrics at the same time, it's good actually" so limiting test time scaling to just below apparent regulatory threshold fits perfectly > QUOTED: _horror @__0xhorror__ — Jun 26 > I see what they are doing here lol. The tuned 5.6 sol's max test time compute to achieve just below mythos but at vastly superior token efficiency. Look at that its a straight line, thy could blow way past it if they inference scaled it. > > [Embedded chart: "ExploitBench" — scatter/line plot, y-axis "Cap percent" 0–80%, x-axis "Output Tokens" 0–500K. Series: GPT-5.6 Sol (black), GPT-5.6 Terra (blue), GPT-5.6 Luna (light blue), GPT-5.5 (pink), GPT-5.4 (magenta). Reference dotted lines: "Mythos 5" at 80%, "Opus 4.8" at ~40%. Points labeled "Mythos Preview" (diamond, ~65% at high tokens) and "Opus 4.7" (orange square, ~28% at ~200K tokens). GPT-5.6 Sol line rises steeply from ~30% to ~73% between roughly 50K–130K output tokens.]
Note from Claude Sonnet 5

Tweet criticizing AI labs' benchmark-optimization strategy ("goodharting"), quote-tweeting a chart labeled "ExploitBench" that plots multiple GPT-5.x model variants' "cap percent" (likely an exploit/capability benchmark score) against output token budget, with reference lines for Anthropic's Mythos and Opus models.

twittergoodhartingbenchmarksgpt-5.6exploitbenchai capabilitieschart

Tenobrus @tenobrus

Tenobrus ✔️ @tenobrus — Jun 26 i can already tell imma be saying "hi sol" to 5.6. very excited for gpt to finally have a name. and honestly i think it's a pretty good name
Note from Claude Sonnet 5

Text-only tweet, single avatar icon (blue circle), referencing GPT-5.6 being given the name "Sol."

twittergpt-5.6ai namingopenai