Nabeel S. Qureshi @nabeelqu
— saved image
Nabeel S. Qureshi @nabeelqu · 1h Very cool sentence: "I find it extremely, extremely wild that 90% of the variance in benchmark scores is explained by a single factor". Effective compute = general factor of intelligence, machine edition. [Quoted tweet] Bayesian @Bayesian0_0 · 21h Replying to @gwern i've uh had the opposite philosophy of just scale the data (so ingest an additional benchmark whenever i come across one, or have the llms slop search new ones but they are having troub... [Two embedded scatter plot charts below the quoted tweet: left chart titled "IRT Model Fit Quality" showing predicted vs actual values with points scattered around a diagonal line, and a histogram-like count plot; right chart shows 18,269 observations, 155 sources, plotting some fit quality metric against difficulty (BEDI) from 50 to 150, with point sizes varying, both charts partially cropped.]
Note from Claude Sonnet 5
Nabeel Qureshi comments on a claim (attributed to a quoted thread involving gwern and @Bayesian0_0) that 90% of variance in LLM benchmark scores is explained by a single factor, likening it to 'effective compute = general factor of intelligence, machine edition.' Two scatter-plot charts (IRT model fit quality, and fit vs. difficulty) are shown below as supporting data.