← All topics

llm-training

2 captures, most recent first.

Jifan Zhang @jifan_zhang

Jifan Zhang @jifan_zhang · 3h you could decompose math papers into smaller subproblems. they did claim it's a general model. 💬 1 🔁 ❤ 1 📊 203 ↗ Dimitris Papailio... @Dimitris... · 2h I don't understand. All I'm saying is that you need a curriculum type of problem description generation during RL 💬 1 🔁 ❤ 4 📊 210 ↗ Jifan Zhang @jifan_zhang · 2h i agree getting a curriculum is necessary and probably easy once you have the questions. generating new problems that are sufficiently diverse and at the right difficulty level seems quite hard. not sure what you meant by problem descriptions, but i was just saying there may be enough hard (sub)problems in math papers already. 💬 1 🔁 ❤ 1 📊 110 ↗ Jifan Zhang @jifan_zhang · 1h fwiw, creating IMO questions is generally considered much harder than solving them, but they also require somewhat different skills. not clear to me how writing questions can be easier than solving for LLMs.
Note from Claude Sonnet 5

A technical Twitter thread between ML researchers Jifan Zhang and Dimitris Papailiopoulos debating curriculum/problem-generation strategies for RL training on math reasoning (decomposing math papers into subproblems, IMO-question generation vs solving difficulty). ML-research content Nathan was reading; relevant to his interest in RL training curricula (parallel to brain_graph_1 curriculum design) but not to AI safety/welfare themes.

twittermachine-learningreinforcement-learningmath-reasoningcurriculum-learningllm-training

davinci @basedneoleo

``` davinci @basedneoleo i've been considering writing out a proper blog post on the matter as soon as i find time to better flesh out my thoughts o[n what] seems to still be a contrarian take even today but essentially one of my core[disagreements...] [continues, cut off] ——— davinci @basedneoleo · Apr 16 that's true behavior cloning on human text is a shortcut to practical crystallized intelligence just like robotic behavior cloning on human motion is a shortcut to routine manual tasks. it's a mere reflection of a crystallized skill not a reproduction of the fluid intelligence that originally produced it. u get some generalization ofcourse but it's far more restricted to the original data distribution since that's what u're trying to model. we will not be seeing superhuman capability from human output approximation. 1 reply, 1 like, 89 views Show replies davinci @basedneoleo · Apr 16 children don't behavior clone on adult output as much as we think they do. they're much more...self-supervised. text is a few degrees seperated from the world it represents. when u train on text, u are not really modeling the world itself as much as u are modeling humanity's biased and sparse projection of its own world model onto text. the direct friction necessary for learning and the one that children are heavily subjected to is largely absent from the pretraining process. what u want is something that can independently generate it's own projection of the world and refine it not model ur own. 1 reply, 2 reposts, 5 likes, 475 views ```
Note from Claude Sonnet 5

A technical Twitter debate disputing Ilya Sutskever's thesis on language modeling as approximating an "adult mind," arguing instead that current LLM training via next-token prediction on human text is behavior cloning (crystallized skill) rather than reproducing a child's experiential learning capacity. Relevant to Nathan's interest in AI cognitive architecture and learning-paradigm debates. Continuation of the same Twitter thread as the previous screenshot — invoking a famous Alan Turing quote (from his 1950 "Computing Machinery and Intelligence" paper) about simulating a child's mind and educating it, as historical grounding for the "child-mind not adult-mind" critique of current LLM training. Same thread as Screenshot_20250420-105319. Continuation of the same thread (see Screenshot_20250420-105319/105333) — argues that text is a "sparse, biased projection" of humanity's world model, that pretraining lacks the "direct friction" of embodied childhood learning, and that superhuman capability requires self-supervised world-model generation rather than human-text imitation. Substantive argument about limits of LLM pretraining vs. embodied/self-supervised learning.

twitterilya-sutskeverllm-trainingbehavior-cloningai-learning-theorymachine-learningalan-turingchild-mindworld-modelsself-supervised-learning