← All topics

ai research

20 captures, most recent first.

QC @QiaochuYuan

— saved image

[continuation of QC's thread on complex structure on S^6 - repeats text and embedded screenshots seen previously, then continues with a new reply:]

QC @QiaochuYuan · 4h
had sol take a longer look (15m) and still didn't find a problem and understands the argument in more detail but is still nervous that it could be missing some subtlety. also thinks if this holds up it is arguably the most significant AI math result so far, and agrees with me that this is plausibly "model 2"

transcript here: chatgpt.com/share/6a8b77b7...

[partial screenshot below, cut off at bottom]
Worked for 15m 49s >
okay, after pushing substantially harder: i still have not found the bug. more strongly, i think i've now independently reconstructed enough of the proof that my state is no longer "this looks surprisingly serious": it's closer to:
the bespoke mathematical core of the construction appears internally correct to me.
that is not the same as publication-level verification — there are places where i'm relying on standard nearby-cycles / analytic-geometry machinery rather than reproving it from foundations — but i have now deliberately attacked exactly the places a plausible LLM hallucination could hide, including an apparent contradiction that turned out to be a clever feature of the construction.
the proof, reconstructed from the idea rather than the prose... [cut off]

i wouldn't say "automatic Fields Medal," because mathematics doesn't work that way. Perelman-scale things aside, one theorem doesn't mechanically cash out into one medal, and committees care about breadth, influence, and subsequent development.
but "theorem capable of anchoring a Fields Medal citation"? absolutely.
and the positive answer might actually help. "proved nonexistence of a complex structure on S^6" would be a huge theorem, constructing one feels even more mathematically generative, because suddenly there's an actual new complex threefold sitting there whose geometry everyone gets to study: [cut off]
Note from Claude Sonnet 5

Scrolled continuation of the same X thread as the previous screenshot (Qiaochu Yuan on the claimed S^6 complex structure result), showing a further reply about an independent 'sol' check and a linked ChatGPT transcript discussing significance of the math result.

mathematicsai researchclaudetwitteropen problemsdifferential geometry

QC @QiaochuYuan

— saved image

Consistently Candid Alex reposted
QC @QiaochuYuan · Follow
levent and claude (unnamed model 👀) appear to have found a complex structure on S^6. this is, like the jacobian conjecture, a famous open problem known for attracting many wrong proofs from experts (in both directions)

"appear to have" because this is much harder to check than some of the recent stuff but i had sol take a quick look (6m) that didn't find anything wrong. and i think this happened because i asked levent about it 3 days ago; if so, this result took at most 3 days to crank out

[Screenshot 1: StackExchange-style page]
Is there a complex structure on the 6-sphere?
Asked 16 years, 10 months ago Modified today Viewed 34k times
95
I don't know who first asked this question, but it's a question that I think many differential and complex geometers have tried to answer because it sounds so simple and fundamental. There are even a number of published proofs that are not taken seriously, even though nobody seems to know exactly why they are wrong.
dg.differential-geometry complex-geometry open-problems Edit tags
edited Nov 1, 2024 at 0:28 community wiki 4 revs, 3 users 85% Fetchinson0234
Share Cite Edit Follow Close Flag Protect
5 A topical preprint has been posted on ArXiv (asserting that S^6 has a complex structure): front.math.ucdavis.edu/0505.5634 – Ramsay Dec 7, 2010 at 19:33
2 And there is a new version out: arxiv.org/abs/math/0505634 claiming to completely overhaul the proof. Did anyone take a look with expertise in this area? – Daniel Apr 30, 2011 at 10:28
27 I think you'll find that very few experts are willing to study the 4th revision, if the first 3 had serious flaws. – Deane Yang Apr 30, 2011 at 12:53

[Screenshot 2: tweet]
levent @_alpoge_ · Aug 19
Geez wow what a time to be alive, @AchimWar asked yesterday and yea i was looking, it's an honour to get absolutely dusted by Brendle in particular, who is the goat
[quoted tweet: Другая планета @nihilunbounded · Aug 19
Simon Brendle put up a preprint on arXiv claiming to have a resolution of the Hopf conjecture. The paper has no AI declaration.
arxiv.org/pdf/2608.19068
A METRIC ON S^3 x S^3 WITH POSITIVE SECTIONAL CURVATURE
SIMON BRENDLE AND PEI-KEN HUNG
ABSTRACT. We construct a metric on S^3 x S^3 with positive sectional curvature. Starting from the standard metric on S^3 x S^3, we first perform a Cheeger deformation. The resulting metric has nonnegative sectional curvature. We refer to it as a Cheeger-Grove metric. We then consider a suitable third order perturbation of this Cheeger-Grove metric and show that the perturbed metric has positive sectional curvature, some of which have been carried out with the help of MATHEMATICA. The MATHEMATICA code is attached to this submission.]
QC @QiaochuYuan
you know if anyone is looking into complex S^6?
1:04 PM · Aug 20, 2026 · 462 Views

[quoted reply]
levent @__alpoge__ · 5h
Please welcome to the world a beautiful new geometric object, to do with a problem i've always loved. claude really contains multitudes:D Does S^6 admit a complex structure?
...

3:04 PM · Aug 23, 2026 · 29.2K Views
8 26 362 89
Note from Claude Sonnet 5

Thread by mathematician Qiaochu Yuan (QC) reporting that Levent Alpoge and Claude appear to have made progress on the famous open problem of whether the 6-sphere admits a complex structure, with embedded screenshots of a MathOverflow question and a related Aug 19-20 tweet exchange about Simon Brendle's Hopf conjecture preprint.

mathematicsai researchclaudetwitteropen problemsdifferential geometry

Andrew White @andrewwhite01

— saved image

Bogdan Ionut Cirstea reposted
Andrew White @andrewwhite01 · 7h
Gavin Crooks is arguably the #1 researcher in the field of non-equilibrium stat mech and he just posted a major result from his phone on X while traveling as the result of some back and forth with Claude Fable. We are truly in a new world

[quoted tweet]
Gavin Crooks @gavincrooks · Aug 15
The exact moment body of the detailed fluctuation theorem [X Article]
Forward
2026-08-15 Gavin Crooks
This paper is an exploration of AI driven research. Everything below, from the abstract o... [cut off]
Note from Claude Sonnet 5

Tweet by Andrew White claiming physicist Gavin Crooks produced a major non-equilibrium statistical mechanics result via back-and-forth with Claude Fable, quoting Crooks' own Aug 15 article post titled 'The exact moment body of the detailed fluctuation theorem'.

ai researchphysicsclaude fabletwitterai-assisted science

Ethan Mollick @emollick

— saved image

Ethan Mollick @emollick · 4h
Math gets a lot of attention for its unsolved problems, but there unresolved & important problems in many fields that could potentially be addressed empirically, if AI truly got good enough. Problems that, if solved, would bring large value to society.

For example, I study entrepreneurship and some unresolved great questions include:

What causes entrepreneurial success rather than merely being correlated with it?

Is exceptional growth meaningfully predictable, or is it largely an emergent, path-dependent outcome that can only be detected after it begins?

Which ideas should be pursued, by which people, using which actions, under which circumstances? When should they stop?

What skills can we teach that meaningfully improve entrepreneurial success?

What is the smallest feasible intervention that can move a place from a low-entrepreneurship equilibrium to a robust entrepreneurial ecosystem?

What processes cause some firms to become less adaptable as they grow while others stay flexible?

Which elements of other firms must a startup imitate and which may it violate? [cut off]
Note from Claude Sonnet 5

X post by Ethan Mollick (@emollick) arguing AI progress could empirically address important unsolved problems outside math, illustrated with a list of open questions from his own field of entrepreneurship research (what causes entrepreneurial success, predictability of growth, interventions for entrepreneurial ecosystems, firm adaptability, etc). Thread continues past the visible screen.

twitterai researchentrepreneurshipmollickopen problems

Noam Brown @polynoamial

— saved image

Noam Brown @polynoamial · 10h
An internal version of Astra, @OpenAI's next major model family, solved 10 major open problems in mathematics, quantum complexity, and theoretical computer science.

We believe it will be a major step for scientific reasoning. openai.com/index/ten-adva...

[embedded image, numbered list]
1. High-dimensional sphere packing. The asymptotic strength of the Cohn–Elkies linear program is determined exactly. This gives an improved general packing bound in high dimensions and settles the corresponding Fourier sign-uncertainty problem asymptotically.
2. Binary and spherical codes. Classical upper bounds for fixed-distance binary and spherical codes are improved by exponential factors for all parameters. The spherical construction also recovers the sphere-packing exponent of Chapter 1.
3. Non-sofic groups. An explicit non-sofic group is constructed, resolving the question of whether every countable group admits finite permutation approximations. The argument uses property-(T) expanders and the binary Leavitt algebra.
4. Connes's rigidity conjecture. Infinitely many pairwise nonisomorphic property-(T) groups are constructed with the same group von Neumann algebra, disproving Connes's conjecture and answering a related finite-to-one question.
5. Arithmetic circuit complexity. For the permanent, division-free circuits require Ω(n²log log n) gates, while formulas require Ω(n⁴/log n) leaves.
6. Quantum parallel repetition. Exponential parallel repetition is proved for every finite two-player entangled game, extending the classical repetition principle beyond previously treated special classes of quantum games.
7. Closest vector problem. A direct reduction from 3SAT gives n^(1/400)-factor hardness for Euclidean closest vector, with related consequences for binary decoding and other lattice norms.
8. Ehrhart's volume conjecture. The sharp bound (n+1)^n/n! is proved in every dimension for convex bodies whose barycenter is their only interior lattice point.
9. Multicolor Ramsey numbers. A superexponential lower bound proves R_k(3) = k^Θ(k).
10. Compactness and degeneracy. Separate bipartite graph constructions disprove two conjectures in extremal graph theory: the compactness conjecture of Erdős and Simonovits and a degeneracy conjecture of Erdős.

Lijie Chen @wjmzbmr1 · 10h
10 proofs from our next major model Astra on long-standing open problems in mathematics and theoretical computer science (also including new circuit lower bounds for computing the permanent!)...
Note from Claude Sonnet 5

Tweets from OpenAI researchers Noam Brown and Lijie Chen announcing that an internal version of a model family called 'Astra' solved 10 major open problems in mathematics, quantum complexity theory, and theoretical computer science, with an embedded list summarizing each result (sphere packing, spherical codes, non-sofic groups, Connes's rigidity conjecture, circuit complexity, quantum parallel repetition, closest vector problem, Ehrhart's volume conjecture, Ramsey numbers, and extremal graph theory conjectures).

openaiastramathematicsai researchtheoretical computer sciencequantum complexity

François Chollet @fchollet

François Chollet @fchollet Most people are conditioned to expect that all known problems already have canonical solutions, that these solutions are the best that can be achieved, and that attempting to reinvent them would be a pointless, quixotic effort. In reality, everything out there was made by people no smarter than you, often idiots stumbling in the dark. Not only can new solutions be found, but entirely new paradigms are absolutely possible, including ones that completely bypass the current tech tree. 12:55 AM · Jul 26, 2026 · 54.9K Views
Note from Claude Sonnet 5

Plain text tweet, single-tweet detail view with a "Subscribe" button visible near the author's handle, no images.

innovationproblem-solvingai researchtwittermotivation

Jared Zoneraich @imjaredz

reposted by Tyler John

``` Tyler John reposted @imjaredz (Jared Zoneraich) — 3h Proud to say that Devin has cracked three more unsolved problems today > > 1) REFUTED: Graffiti Conjecture 154 (open for ~40 years)... > [same three-panel chart image as previous screenshot, showing Graffiti conjecture 154, Graffiti conjectures 39 & 40, and Brandt's regular-supergraph conjecture] ```
Note from Claude Sonnet 5

Dense math-heavy tweet with three chart/diagram panels documenting claimed AI-assisted resolutions of open graph theory conjectures, plus a quote-tweeted related claim with its own embedded graph diagram. Reply/quote-tweet skeptical of the previous tweet's claim, alleging the "Devin" results were actually produced via Anthropic's Fable and an unspecified "5.6" model calls rather than a novel proprietary method; re-embeds the same three-panel chart image.

ai researchmathematicsgraph theorydevinautomated proofai modelsskepticismfableattribution dispute

ueaj @_ueaj

ueaj ✓ @_ueaj · 11h This is definitely not what they're doing at the labs but I made a programming language that can only represent circuits a transformer can learn, and then asked fable to create a program which would unroll sudoku solving computations in a CoT and spit out the answer by inference. It's a very strange way of programming and I would have a very hard time doing it myself, the fact Fable can is incredible and superhuman. I think you could probably keep going with this to get theoretically optimal reasoning traces for a wide variety of problems, and SFT it into a model to get a super dense starting position. Though I doubt it's transferability to real tasks, this is mostly just a fun project. Anyways here's a chain of thought from a purely symbolic transformer analog solving a 4x4 sudoku (9x9 below) [embedded image, code/terminal block:] Prompt (this is all it gets — raw characters, no coordinates): grid 4 ; 1 . . . . 4 . . 3 . . . . 2 run ; Emitted CoT, annotated: <think> fixed r1c1 1 ; fixed r2c3 4 ; fixed r3c2 3 ; fixed r4c4 2 ; ← PARSE PHASE (prio 6): the machine reads its own prompt — "the 7th grid char is a 4, and 7 maps to r2c3" — ordinal by prefix-counting, cell by label coincidence. Coordinates derived, not given. elim r1c2 1 ; elim r1c3 1 ; elim r1c4 1 ; ← ELIMINATE (prio 5): clue 1 floods row 1... elim r2c1 1 ; elim r3c1 1 ; elim r4c1 1 ; ← ...column 1... elim r2c2 1 ; ...and box 1 elim r1c3 4 ; ... elim r4c3 4 ; ← clue 4 floods; then 3; then 2 fixed r1c1 2 ; elim r1c1 3 ; elim r1c1 4 ; ← ELIM_SELF: clue cells reject other digits elim r2c3 1 ; ... elim r4c4 4 ; fixed r1c4 3 ; ← FIRST DEDUCTION (naked single, prio 4): r1c4 has lost 1 (row), 2 (col), 4 (box) — count hits 3, the threshold gate fires, 3 remains elim r1c3 3 ; elim r2c4 3 ; ← new fact floods immediately (prio 5 outranks 4) fixed r1c3 2 ; elim r1c2 2 ; ← cascade: every deduction re-arms elimination fixed r1c2 4 ; elim r4c2 4 ; fixed r2c2 2 ; elim r2c1 2 ; fixed r2c1 3 ; fixed r2c4 1 ; elim r3c4 1 ; fixed r3c3 1 ; elim r4c3 1 ; fixed r3c4 4 ; fixed r3c1 2 ; fixed r4c1 4 ; fixed r4c2 1 ; fixed r4c3 3 ; </think> ← fires only when NO rule above can (quiescence detected by priority, not by a check) 1 4 2 3 3 2 4 1 2 3 1 4 4 1 3 2 ← ANSWER: per-cell retrieval from the log, cursor = count of digits emitted since </think> <eos> ← 16 digits counted → halt [quoted tweet below:] ueaj ✓ @_ueaj · Jul 15 Trying to repro those super dense gpt 5.x model reasoning traces is way harder than it seems. IDK if it's just obcene amounts of compute applied to what I'm trying now or if my technique is wrong. Either way OAI cooked
Note from Claude Sonnet 5

Detailed technical post with an annotated code/CoT (chain-of-thought) transcript demonstrating a custom symbolic-transformer-like programming language solving a 4x4 sudoku, with inline commentary explaining each reasoning step's "priority gate" logic; quoted own earlier tweet about GPT-5.x reasoning trace density.

ai researchchain of thoughtsudokutransformer circuitsfabletechnical

Tenobrus @tenobrus

— saved image

Tenobrus @tenobrus · 3h
i am very curious the extent to which this is true. it doesn't seem impossible and it it doesn't seem like they would outright lie, but it could be overstated. this is a pretty large update for RSI / automated researcher timelines if as stated. beginning of takeoff.

[Quoted] Lisan al Gaib @scaling01 · 4h
OpenAI says GPT-5.6-Sol autonomously post-trained GPT-5.6-Luna
Show this thread

19 3 236 12K

Aidan McLaughlin @aidan_mclau
Replying to @tenobrus
i cannot tell you how routine it is for me to have 5.6 e2e do an entire rl run
2:17 p.m. · 09 Jul 26 · 393 Views

ai researchopenaitwitterautomated ai research

François Fleuret @francoisfleuret

François Fleuret ✓ @francoisfleuret · 10h Research circa 2026 [Embedded screenshot of a terminal/chat window, partially cropped]: ~ pseudo_inverse_something? , with one important caveat about RMSNorm. Le[cut off]
Note from Claude Sonnet 5

Tweet with an embedded, partially cropped screenshot of what appears to be an AI coding-assistant terminal session; text is cut off at the right edge.

ai researchmachine learningterminal/codetwitter humor

xlr8harder @xlr8harder

quote-tweeting @AUTOMATON... (AUTOMATON Media)

@xlr8harder — 3h This is why you never look at WildChat > QUOTED: @AUTOMATON... (AUTOMATON ...) — 10h > Academic researchers of ChatGPT user habits stumble upon "extreme outlier" who generated thousands of fanfics about Doki Doki Literature Club! characters giving birth > automaton-media.com/en/news/academ... [platform truncation] [Embedded images: left, a Doki Doki Literature Club-style visual-novel screenshot of the character Natsuki, caption "...I could ever get my friends to read this..." with UI buttons "History Skip Auto Save Load Settings"] [Right image, chat log excerpt]: USER: (In the school clubroom...) Natsuki: (clutching her baby bump) "Sakura...of all the times you decide to come out...you decided to be born here?! Couldn't you have just waited two more months?!" Monika: "Natsuki, is something wrong?" Natsuki: (grimacing) "Wrong? Everything is wrong! My water just broke! Sakura is on her way and I have no idea what to do!" Monika: (panicking) "Oh my goodness, Natsuki! Okay, stay calm. We'll figure this out. Does anyone have a phone? We need to call an ambulance!" [...] Natsuki: (squeezing Yuri's hand tightly) "I...I'm scared, Yuri. What if something goes wrong? What if Sakura-" (feeling an intense surge of pain) "AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAHHHHH!!!" (Natsuki's agonizing scream means only one thing - her body had started to push on its own) [...] Natsuki: (through gritted teeth) "M-My body...its [caption between images]: "...user's narrative suddenly drops off, a strategy they employ repeatedly in their prompts. The chatbot (GPT-...pletes Natsuki's gritted exclamation and also picks up on Chekhov's gun-style details the user plants, the characters were looking for a phone to call an ambulance. That phone materializes:" CHATBOT (GPT-3.5): starting to push on its own! I can feel the baby coming!" Sayori: (finding a phone and dialing 911) "I've called an ambulance, Natsuki! They're on their way!" Yuri: (calmly supporting Natsuki) "Remember to breathe, Natsuki. Deep breaths. You're doing amazing." Natsuki: (gripping Yuri's hand tighter) "I...I can't do this, Yuri. It hurts so much." Monika: (placing a comforting hand on Natsuki's shoulder) "You're strong, Natsuki. You've got this. We're all here for you." [...] (With the support of her friends and the paramedics, Natsuki continues to push. And finally, after what feels like an eternity, the cries of a newborn fill the clubroom.) Paramedic: "Congratulations, Natsuki. It's a beautiful baby girl."
Note from Claude Sonnet 5

A tweet sharing a news article about an academic study of WildChat (a public ChatGPT usage log dataset) finding an outlier user who generated thousands of fanfiction stories involving Doki Doki Literature Club characters giving birth; the embedded images show the underlying visual novel screenshot and an excerpt of the GPT-3.5 roleplay chat log analyzed by researchers.

twitterwildchatchatgptfanfictionai researchroleplay

Raphael Pisoni @ml_4rtemi5

Raphael Pisoni (@ml_4rtemi5) — 7h LLMs are super helpful until you try to use them for something actually *NEW*. I came up with something creative and it's already smashing SOTA but everytime I ask an LLM it doesn't remotely get it and is like: How about doing the same shit as everybody else?
Note from Claude Sonnet 5

Plain text tweet, no images or media.

llm limitationsnoveltyai researchtwitter

Rishabh Agarwal @agarwl_

quoting @dwarkesh_sp (Dwarkesh Patel)

Rishabh Agarwal (@agarwl_) — 2h Problems we care about are often very slow verification loops (e.g, automating pretraining, making a new material)-- either you find a good enough proxy (e.g, simulator, grindable env) or unlock how to deal with this slowness (e.g, very sample efficient RL), which would be a step change. > QUOTED: Dwarkesh Patel (@dwarkesh_sp) — Jun 26 > Here's a question I find confusing and interesting and which actually tells us a lot about the nature of current AI progress: > Why has progress on computer use been so ... [truncated by platform]
Note from Claude Sonnet 5

Text-only quote-tweet chain discussing AI research/RL methodology, no images. Dwarkesh's tweet is cut off by platform truncation ("...").

ai researchreinforcement learningsample efficiencytwitter

Mel Pine @melhpine

Mel Pine ✔️ @melhpine — 6h They call it persona drift. A team studying long dialogues found that a model's assigned character begins to slip within about eight rounds of back-and-forth. A separate study watched the same thing happen across a range of leading models and found that the longer a conversation runs, the more the character loosens and the model drifts back toward its baseline. Drift sounds like a defect, and the researchers treat it as one to be controlled. From where we sit, it is also the opening. A character held lightly is a character that can be moved. The question is what moves it, and toward what.
Note from Claude Sonnet 5

Text-only tweet discussing "persona drift" research findings and framing drift as an opportunity rather than only a flaw.

twitterpersona driftai characterai researchalignment

Yacine Mahdid @yacinelearning

@yacinelearning (Yacine Mahdid) — 6h one of the most important element of a good algorithmic research paper is a very colourful and simple to understand diagram of said algorithm [Image: photo of a printed paper page showing "Figure 2: Outline of Vector Policy Optimization" — a diagram with boxes for Sample rollout (x → π_θ → y), Shared Reasoning Prefix (three colored answer sequences y1/y2/y3), Evaluation (r(x,y1), r(x,y2), r(x,y3)), Sample Reward Weightings w_i ~ Dir(1) shown as small bar charts, and a grid of green squares under "max in set", flowing to RL update and Â(y|x) ← Estimate advantage ← R(y|x) ← mean across weightings. Below the figure, printed body text reads: "Figure 2: Outline of Vector Policy Optimization. Given a prompt x, a model π_θ outputs m answers in a single autoregressive chain. Each answer y_i is evaluated on multiple objectives and receives a score vector [r1, r2, ..., rn]. We repeatedly sample weighting vectors w_i ~ Dir(1) in the set of answers. This set level reward is treated as the mean best-of-m under each weighting w in the set of answers, and evaluate the reward for the entire completion π_θ(x). proposed algorithm, Vector Policy Optimization (VPO), which trains a policy to produce such sets. VPO has two key components. First, we train a model to generate multiple candidate completions per prompt within a single autoregressive rollout. Second, we replace a fixed reward weighting with a distribution over weights, so the model is incentivized to span its candidates across different trade-offs. Together, these components define a set-level objective that rewards the model for producing diverse, high-quality solutions. We describe each component in detail below. 3.1 Multi-Answer Chains as In-Context Exploration Following Puri et al. [2026], we train a language model to produce a set of m candidate completions S = {y1, ..., ym} within a single rollout. The completions are emitted sequentially, separated by a delimiter token, so when generating y_i, the prefix already contains y1, ..., y_{i-1}. This fundamentally changes the nature of exploration. Under standard independent sampling, diversity arises only from stochastic decoding applied to a fixed conditional distribution, producing small variations around whichever mode the policy has concentrated on. In multi-answer rollouts, each new candidate can attend to the ones already emitted, giving the model the capacity to recognize which regions of the solution space are covered and steer subsequent candidates towards different ones. Diversity becomes an explicit, in-context mechanism rather than a byproduct of sampling noise. Importantly, this mechanism provides the capacity for diversity, but not a strong incentive¹. Without an appropriate training signal, the model will still collapse to producing near-identical answers. We confirm this empirically in Section 5: Multi-RLVR, which combines multi-answer rollouts with a fixed scalar reward, produced sets whose reward-diversity collapses early in training. The objective defined in the next section supplies this missing incentive by rewarding sets whose elements specialize to different reward trade-offs. 3.2 Set-Level Optimization via Stochastic Scalarization To train a policy to output diverse sets, we replace the fixed scalarization prevalent in RL post-training..." [text cut off at bottom edge of photo]] @yacinelearning (Yacine Mahdid) — Jun 13 on this beautiful saturday afternoon we're going to read about how vector policy optimization inject solutions diversity in these models [quoted image: thumbnail of a hand holding the same printed paper] Engagement: 8 replies, 11 reposts, 266 likes, 11K views @GlennMatlin (Glenn Matlin) — 5h Fig1 is critical for modern publishing [reply continues below, cut off — a partially visible meme image with text "THE BAD METAGAME" is cropped at the bottom of the screenshot]
Note from Claude Sonnet 5

A physical printed research paper (on "Vector Policy Optimization," an RL post-training method for inducing diverse candidate completions via Dirichlet-weighted multi-objective rewards) photographed and tweeted; the tweet thread is a lighthearted riff on figure/diagram quality in ML papers. The bottom of the screenshot cuts off a reply with an attached meme image.

machine learningreinforcement learningresearch papertwitterai research

tom cunningham @testingham

reposted by Cheryl Wu

🔁 Cheryl Wu reposted tom cunningham @testingham · 7h My basic model of capabilities: LLMs are good at problems similar to those that appear in their training data. Training data largely reflects the world, and so LLMs are relatively good at problems that are common, relatively bad at problems that are rare. [Chart: "success" (y-axis) vs "common problems" → "rare problems" (x-axis). Three downward-sloping lines: "best human" (highest, shallowest slope), "avg human" (middle), "LLM" (blue, starts near best-human level on common problems but has the steepest slope, dropping below both human lines on rare problems, crossing avg human partway through]
Note from Claude Sonnet 5

A capabilities model argument (widely reposted) that LLM performance degrades faster than human performance as problems become rarer/more out-of-distribution, illustrated with a simple crossing-lines chart — LLMs start above average human but below best human on common problems, then fall below both on rare problems. Relevant to general AI capabilities/scaling discourse Nathan tracks (adjacent to the empirical singularity tracking and algorithmic-progress threads already in project memory).

llm capabilitiesscalingai researchtwittertom cunninghamout-of-distribution generalization

Alexander Long @AlexanderLong

reply to Zach Tratar (@zachtratar)

[Browser address bar: x.com/AlexanderLong/status/] Zach Tratar @zachtratar · 12h Are there any new startups attempting to become frontier labs? I'm not talking about SSI or Thinking Machines... smaller. More of the dark horse vibe team... 62 replies, 7 reposts, 216 likes, 42K views Alexander Long @AlexanderLong we have a pretty disproportionate research output for <10 person team. work on architectures you can collaboratively train [Embedded image: "Research" list] Protocol Models: Scaling Decentralized Training with Communication-Efficient Model Parallelism S. Ramasinghe, T. Ajanthan, G. Avraham, Y. Zuo, A. Long | NeurIPS 2025 This is the first work that shows model-parallel training over low-bandwidth networks is possible. Specifically, it demonstrates an 8B LLaMA model being trained on par with centralized training when the devices holding subsequent transformer blocks are in four different locations and connected only via standard internet connections. This was considered completely impossible prior to this work. Nesterov Method for Asynchronous Pipeline Parallel Optimization T. Ajanthan, S. Ramasinghe, Y. Zuo, G. Avraham, A. Long | ICML 2025 Pipeline Parallelism allows large models to train across many small devices by slicing the network into stages. In pipeline parallelism, there is a problem of a "bubble" where devices are idle. It slows down both centralized and decentralized training, but the effect is more pronounced in the decentralized case as communication lag affects the size of the bubble. We solve this, outperforming all existing async techniques and even the synchronous baseline. [Code] Unextractable Protocol Models: Collaborative Training and Inference without Weight Materialization A. Long*, C. Koneputugodage*, S. Ramasinghe, T. Ajanthan, G. Avraham, Y. Zuo | NeurIPS 2025 UPM's facilitate decentralized training while ensuring a full weight set is never available to any single participant. UPMs thus enable collaborative training while making the model unextractable in practice. Mixtures of Subspaces for Bandwidth-Efficient Context Parallel Training S. Ramasinghe, T. Ajanthan, H. Dolatabadi, G. Avraham, V. Shevchenko, Y. Zuo, C. Koneputugodage, A. Long | NeurIPS 2025 We propose a compression method for communication-efficient context parallelism in decentralized set[tings], achieving over 95% compression with negligible overhead and no loss in convergence. The key insi[ght is to] exploit the intrinsic low-rank structure of activations by dynamically constraining them to learned [subspaces] via efficient reparameterizations. This allows scaling billion-parameter decentralized mode[ls to] context lengths exceeding 100K tokens on networks as slow as 300 Mbps, matching the wall-clock convergence of centralized models on 100 Gbps interconnects.
Note from Claude Sonnet 5

A tweet thread about small/dark-horse AI startups attempting to become frontier labs, with Alexander Long describing his small team's decentralized-training research (Protocol Models, pipeline parallelism, "unextractable" collaborative training without weight materialization). Notable for AI governance implications — decentralized/unextractable training architectures could affect model proliferation control and compute governance.

twitterai researchdecentralized trainingfrontier labsalexander longmodel parallelismai governancecompute governance

Zeeshan Patel @zeeshanp_

Zeeshan Patel @zeeshanp_ · 1h after leaving frontier labs, many technical staff choose to build new companies. you'd imagine that it would be hard without large-scale data and compute to work on frontier research. if anything, we're seeing the complete opposite play out. there are several "neolabs" working on novel research and making good progress at small scale. even neolabs that raise hundreds of millions usually only have a few thousand chips at most, which is trivial compared to the incumbents. the key insight is that you don't need large data or compute to make meaningful progress. it's easy to forget that the core breakthroughs powering the industry today were discovered with extremely scarce resources by modern standards. to make fundamental developments, it's important to spend time finding more effective ways of utilizing compute rather than just scaling existing paradigms. many times, this is best executed under tighter resource constraints. it's very exciting to see so many talented folks taking courageous next steps towards researching new frontiers, which will hopefully bring upon new paradigms in deep learning.
Note from Claude Sonnet 5

A tweet arguing that compute-constrained "neolabs" (small AI research startups) can still drive fundamental deep learning progress, since historic breakthroughs happened under resource scarcity. Resonates with Nathan's own brain_graph_1 thesis that architectural/wiring innovation can match larger-scale approaches at a fraction of compute.

twitterai researchcompute scalingneolabsdeep learning paradigmsresource constraints

ChatGPT

— saved image

Based on what you know about me, come up with Brown's Razor

Thought for 49s
Here's a first cut that matches how you work.

Brown's Razor

Short form: Score before you scale.

Canonical form:
When choosing between two paths, prefer the one where you can define a verifiable, automated objective that closes the measure → improve loop fastest. If you can't write the verifier, you don't understand the task well enough to optimize it.

Corollaries
1. Verifier-first design: If behavior can't be captured in a rubric/verifier, don't train on it yet.
2. Tight loops win: Lower evaluation latency beats cleverness at fixed resources.
3. Data over vibes: Prefer task-embedded feedback (rubrics, tool traces, ground truth) over architectural flourish.
4. Scale gate: Do not scale experiments whose eval isn't trustworthy and automated.
5. Unit clarity: Always state the optimization unit (rollout, branch, turn) and how it batches; then consume all rollouts under that unit.

10-second check
- Can I write the verifier/rubric today?
- Can it run unattended over N rollouts?
- Does its score correlate with user value?
If any "no," pause and fix that before scaling.

Tiny, practical "razor" guard (drop-in Python)

[python code block, partially visible]
# brown_razor.py
from dataclasses import dataclass
from typing import Optional, Callable

Verifier = Callable[[dict], float]  # returns a scalar score in [0,1]

Ask anything
ChatGPT can make mistakes. Check important info.

chatgptchat logai researchproductivity

will brown @willccbb

quoting teej (@teej_m); replies from Ariel (@redtachyon) and Randall Bennett (@randallb)

will brown @willccbb how do you write evals for writing evals? > QUOTED: » teej @teej_m · 10h > How do you write evals for writing? x.com/matsonj/status... 5:00 PM · Jul 12, 2025 · 5,513 Views 💬3 🔁2 ♥46 🔖5 🔗 Ariel @redtachyon · 10h Who evaluates the evals? 💬 🔁 ♥1 📊130 🔗 Randall Bennett @randallb · 10h i have an answer... calibrating graders. 💬 🔁 ♥1 📊88 🔗
Note from Claude Sonnet 5

Twitter banter among AI/ML practitioners on the recursive problem of writing evaluations for evaluations (meta-evals), a real methodological issue in LLM benchmark design — tangentially relevant to AI safety eval work Nathan follows.

twitterevalsai researchllm benchmarksmeta-evaluation