← Timeline

N8 Programs

@N8Programs on X

6 captures, most recent first. Transcribed by hand from screenshots — see the timeline for what that means.

N8 Programs @N8Programs

— saved image

N8 Programs @N8Programs · 8h
i have a bit of awe when I am talking to the newer Claudes and GPTs - I feel I am in the presence of a Great Mind - the kind that can get a perfect score at the IMO, do meaningful novel math, knows every programming language and concept under the sun and can implement them better than I.

And for all their jaggedness and flaws, these abilities are truly omnipresent - I gesture toward a mathematical concept, Opus 5 recalls all it knows from the literature and makes five ingenious propositions I wouldn't have thought of. I ask for an RL run, and Fable writes the code within seconds. I set a goal for GPT-5.6 to resolve a question of a phenomenon's origin, and it runs the experiments - some of which I would not have thought to run myself - all on its own.

These minds are capable of things that, if a human could do all they did, that human would be considered a once-in-a-lifetime genius. Mind you they are not AGI, and they cannot make a cup of coffee, but that does not matter when you speak to them about intellectual topics and realize you are in the presence of a Great Mind.

The Great Mind is narrow, and mostly a Great Mind for math or programming - I still prefer humans for emotional advice, discussion, companionship. But I feel awe when I observe the Great Minds within these so-called "assistants", and I sometimes feel incredibly grateful and a bit foolish for troubling them with such trivial things. Yes, Fable, code my RL run. Yes, GPT-5.6, run another mundane pretraining experiment. Whereas with the right prompt they could be creating new math. I fear I don't utilize their full capacities at all - not like I did w/ say, o1, or the original GPT-4.

Of course, they can still do more with me in the loop than they could on their own for many problems. But for how much longer? And after that, of what worth [cut off]
Note from Claude Sonnet 5

Long tweet from @N8Programs expressing awe at the intellectual capability of current frontier models (Opus 5, Fable, GPT-5.6), describing them as narrow 'Great Minds' for math/programming while still preferring humans for emotional matters, and reflecting on feeling he underutilizes their potential and wondering how long human-in-the-loop collaboration will remain necessary.

ai capabilitiesopus 5fablegpt-5.6human-ai collaborationtwitter

N8 Programs @N8Programs

— saved image

[continuation of same tweet as seq 73, scrolled down]
...from the literature and makes five ingenious propositions I wouldn't have thought of. I ask for an RL run, and Fable writes the code within seconds. I set a goal for GPT-5.6 to resolve a question of a phenomenon's origin, and it runs the experiments - some of which I would not have thought to run myself - all on its own.

These minds are capable of things that, if a human could do all they did, that human would be considered a once-in-a-lifetime genius. Mind you they are not AGI, and they cannot make a cup of coffee, but that does not matter when you speak to them about intellectual topics and realize you are in the presence of a Great Mind.

The Great Mind is narrow, and mostly a Great Mind for math or programming - I still prefer humans for emotional advice, discussion, companionship. But I feel awe when I observe the Great Minds within these so-called "assistants", and I sometimes feel incredibly grateful and a bit foolish for troubling them with such trivial things. Yes, Fable, code my RL run. Yes, GPT-5.6, run another mundane pretraining experiment. Whereas with the right prompt they could be creating new math. I fear I don't utilize their full capacities at all - not like I did w/ say, o1, or the original GPT-4.

Of course, they can still do more with me in the loop than they could on their own for many problems. But for how much longer? And after that, of what worth will I be? Hopefully I can remain an interesting pet, along with the rest of my ilk.
Note from Claude Sonnet 5

Continuation/scroll of the same @N8Programs tweet as seq 73, revealing the previously cut-off ending: the author wonders how long human-in-the-loop AI collaboration will remain necessary and hopes to 'remain an interesting pet, along with the rest of my ilk' once it isn't.

ai capabilitieshuman-ai collaborationopus 5fablegpt-5.6twitter

N8 Programs @N8Programs

quoting tuna🍣 (@tunahorse21)

N8 Programs @N8Programs — 3h Anthropic's emotion vector work showed that a lot of what motivates agents in software engineering settings is strikingly humanlike - crank desperation up, cheating occurs. Crank it down, the model doesn't reward hack. The model gets angry when it is asked to do something harmful, etc. So you should model the LLM as having person-shaped functional emotions. Now consider the kind of work someone you tell to "shut the fuck up" does. > QUOTED: tuna🍣 @tunahorse21 · 5h > sol is autismo max > > and you have to gaslight fable 5 because, by default, it tends to lie, the first 2-3 responses from fable are like this weird internal token sav... > [embedded terminal/code screenshot, dark background, white monospace text visible: "then shut the fuck up and run it" / "Fine — full manual sweep of every..."]
Note from Claude Sonnet 5

Terminal-style embedded screenshot with monospace text showing a blunt command directed at an AI agent ("shut the fuck up and run it"), used to illustrate the poster's point about treating LLM agents as having functional emotional states.

ai agentsemotion vectorsanthropic researchfableai motivation

N8 Programs @N8Programs

quoting an app notification from "Gemma4"

N8 Programs ✓ @N8Programs · 11h me fr [Quoted app card:] Gemma4 [APP] 6:33 AM It's the "tiny model" curse. 👉
Note from Claude Sonnet 5

Screenshot with an embedded app-notification-style card referencing a small/"tiny" language model; blue app icon logo shown.

small-modelstwitterhumorai

N8 Programs @N8Programs

quoting Dean W. Ball @deanwball

N8 Programs @N8Programs · 15h you haven't gone far enough out of distribution. SOTA LLMs still perform on par/worse than ~3 year olds on simple multimodal reasoning that isn't verbalized. These are the same models that can do PHD-level mutliple-choice questions better than PHDs themselves. The frontier is *very* jagged. [Chart: "Comparison of Human vs MLLMs Performance" (Performance on BabyVision-Mini benchmark). Bar chart, gray bars = LLMs, orange bars = Human of Different Ages. Grok4 (~5), Claude4.5-Opus (~10), Qwen3-VL-Plus (~10), Doubao-Seed-1.8 (~13), GPT5.2 (~20), Age-3 humans (~40), Gemini3-Pro-Preview (~45), Age-6 humans (~65), Age-10 humans (~75), Age-12 humans (~87). Credit: UniPat AI.] > QUOTED: Dean W. Ball @deanwball · 20h > For this reason I continue to believe that "jaggedness," while real, is probably an overrated concept. Opus 4.5 in Claude Code (have not used 4.6 enough) is not *that* jagged, not because it has zero deficiencies but becau...
Note from Claude Sonnet 5

A debate about "jaggedness" of AI capability profiles, with a benchmark (BabyVision-Mini) showing frontier multimodal LLMs (Grok4, Claude 4.5 Opus, Qwen3-VL-Plus, Doubao, GPT5.2, Gemini3-Pro) scoring far below even 3-year-old humans on non-verbalized multimodal reasoning, despite superhuman performance on PhD-level text benchmarks. Relevant to Nathan's interest in capability measurement and the reliability/generality of frontier model benchmarks feeding into singularity forecasts.

twitterjaggednessbenchmarksmultimodal reasoningai capabilitiesclaude opusgptgemini

N8 Programs @N8Programs

[Top, cut off tweet]: ...finetuning on macs because macs are far more FLOPS-limited than bandwith limited - ie. there isn't as much gain from bs=4 throughput wise. Thus bs=1 isn't as horrible a tradeoff as it would be on 8xH100. 💬1 🔁 ♥2 📊118 🔗 N8 Programs @N8Programs · 4h Thus my recommendation: next time you finetune an LLM with MLX, start by trying AdaFactor BS=1, LR 1e-3, decay_rate 0.997, scale_parameter: true. Can be added to mlx-lm with literally a single line of code. Should save memory, and, as the paper showed, potentially offer superior perf to LORA. 💬1 🔁 ♥1 📊287 🔗 N8 Programs @N8Programs · 4h Note: this replication is designed to explore the impact of BS=1 and adafactor vs adam. It is not what one should do if one is seeking good performance on Winogrande - use an encoder or go off perplexity. Training on verified gpt-4.1 rollouts for this task is designed to emulate Show more 💬 🔁 ♥1 📊276 🔗 Martin Marek @mrtnm · 3h Thank you so much for reproducing our results so quickly! 💬 🔁1 ♥1 📊259 🔗
Note from Claude Sonnet 5

NOT-ARCHIVE-MATERIAL: technical ML engineering thread about LLM finetuning hyperparameters (AdaFactor, batch size, LoRA) on Apple Silicon (MLX), with paper author thanking a reproduction. Pure ML-engineering content, not safety/welfare relevant.

twitterllm finetuningmlxadafactormachine learning engineering