← All topics

ai capability

7 captures, most recent first.

Lucas Beyer @giffmana

— saved image

Lucas Beyer (bl16) @giffmana · 6h
imma just highlight this part for @GaryMarcus and @ylecun because it's easy to miss: no tools no coding => no symbols, just AR LLM

[quoted image, text highlighted]
The results:
🏅 Asian Physics Olympiad (APhO): Perfect score, theory exam
🏅 International Physics Olympiad (IPhO): Perfect score, theory exam
🥇 International Mathematical Olympiad (IMO): Gold medal
🥇 International Chemistry Olympiad (IChO): Gold-medal-level performance
🥇 Romanian Masters of Mathematics (RMM): Gold-medal-level performance

The types of problems in the Olympiad competitions are exceptionally hard, demanding deep chains of reasoning, creative insight, and flawless argumentation. To test pure reasoning capability, we disallowed all tool use, meaning no search, no coding, and no calculator. [highlighted portion]

[quoted tweet]
AI at Meta @AIatMeta · 9h
To understand whether we're making genuine progress on reasoning, we entered our AI models in five STEM Olympiad competitions.
...
Note from Claude Sonnet 5

Tweet by Lucas Beyer highlighting a passage from an AI at Meta announcement (quoted below) reporting gold/perfect-score results across five STEM olympiads (APhO, IPhO, IMO, IChO, RMM) achieved by a pure autoregressive LLM with all tools disabled (no search, coding, or calculator), addressed rhetorically to Gary Marcus and Yann LeCun as evidence against symbolic-reasoning skepticism.

ai capabilityolympiadmeta aireasoninggary marcusyann lecun

@AndrewLampinen

— saved image

Lisan al Gaib reposted
Andrew Lampinen @AndrewLampinen
The question is "how much is each component is the system contributing to its intelligence & generality" — and there I think it's pretty clear that the neural component is still the thing doing the interesting hypothesis or plan generation, deciding what went wrong, etc. 1/

[quoted tweet]
François Chollet @fchollet · 13h
I would have assumed it was fairly obvious, but in case it's not: a million-line codebase (also known as a "harness"), running at inference time, orchestrating thousands of calls to a neural network for any given task, is the exact definition of a "neurosymbolic ...

7:39 AM · Aug 6, 2026 · 25.3K Views
12 replies, 14 reposts, 191 likes, 64 bookmarks
Relevant

Andrew Lampin... @AndrewLampin... · 10h
This is very clearly different from the vision that many neurosymbolic advocates had a few years ago, e.g. these quotes (taken from arxiv.org/abs/2305.00813 and arxiv.org/abs/1801.00631) in which symbol manipulation did the "intelligent" part. 2/

[quoted image of text]
These arguments, together with similar ones from others, drove a longstanding trend to dismiss neural networks as only capable of modeling simple perceptual processing, rather than "real" higher-level cognition, which is symbolic and systematic. For example "while data driven neural network-based AI algorithms effectively model machine perception, symbolic knowledge-based AI is better [cut off]
Note from Claude Sonnet 5

Continuation of the Andrew Lampinen / François Chollet thread on whether AI system intelligence comes from the neural model or the surrounding 'harness' (see also seq 423). Includes Chollet's counter-argument that a large orchestrating codebase constitutes 'neurosymbolic' AI, and Lampinen's reply linking two arxiv papers (2305.00813, 1801.00631) with a quoted excerpt arguing older neurosymbolic advocates dismissed neural nets as incapable of higher-level cognition.

ai capabilityharness vs modelneurosymbolicandrew lampinenfrancois chollet

@AndrewLampinen

— saved image

[continuation of quoted excerpt]
For example "while data driven neural network-based AI algorithms effectively model machine perception, symbolic knowledge-based AI is better suited for modeling machine cognition," or "The right move may to be to integrate deep learning, which excels at perceptual classification, with symbolic systems, which excel at inference and abstraction. One might think such a potential merger on analogy to the brain; perceptual input systems, like primary sensory cortex, seem to do something like what deep learning does, but there are other areas, like Broca's area and prefrontal cortex, that seem to operate at much higher level of abstraction." Like the earlier advocates of symbols, these perspectives suggest that what's going on in the brain to make intelligent inferences is fundamentally symbolic processing, and that neural networks are not suited to these kinds of inferences — at best, just to perception.

2 replies, 3 reposts, 53 likes, 1.7K

Andrew Lampin... @AndrewLampin... · 10h
Humans also benefit from being "wrapped" in rules that make us recheck our behavior / respect constraints (e.g., code we submit must pass tests, email clients that say "don't click this, we think it's phishing") — but nobody would mistake them for part of our intelligence. 3/
3 replies, 1 repost, 35 likes, 1.2K

Andrew Lampin... @AndrewLampin... · 10h
I would love to see more harness on/off comparisons of systems published to illustrate the point, though! 4/4
Note from Claude Sonnet 5

Continuation of the Andrew Lampinen thread (see seq 423, 430): the rest of the quoted excerpt on old neurosymbolic arguments, followed by Lampinen's tweets 3/4 and 4/4 arguing humans are also 'wrapped' in external rules/constraints without those being mistaken for intelligence, and calling for published harness on/off comparisons.

ai capabilityharness vs modelneurosymbolicandrew lampinen

@AndrewLampinen

— saved image

Emmanuel Ameisen reposted
Andrew Lampinen @AndrewLampinen · 5h
To state this another way, 2 ablations:
1) Take Opus-5, remove any harness. You'll at least sometimes get decent ideas and analysis.
2) Take the full harness, swap the model to GPT-3. You'll get trash on any interesting problem.

The intelligence is from the model, not harness.

[quoted tweet]
Andrew Lampin... @AndrewLampin... · 10h
The question is "how much is each component is the system contributing to its intelligence & generality" — and there I think it's pretty clear that the neural component is still the thing doing the interesting hypothesis or plan generation, ...
Note from Claude Sonnet 5

Tweet thread by Andrew Lampinen (reposted by Emmanuel Ameisen) arguing via a thought experiment (Opus-5 without harness vs. full harness with GPT-3) that model intelligence, not the surrounding agent harness, drives capability, quoting his own earlier tweet on the same point.

ai capabilityharness vs modelopus-5andrew lampinen

Tenobrus @tenobrus

— saved image

Tenobrus @tenobrus · 26m
am i a better writer than scott alexander?

no. i never will be. i won't come close.

that's okay though, i have fun with what i do. and there aren't infinite scott alexander articles. he doesn't write about everything, and even when he does write about things ive talked about too he doesn't necessarily talk about the exact angles i may have interest in or some aspect of expertise he doesn't.

but how well does that hold up if / when model intelligence and writing quality noticeably surpasses scott alexander, quickly and on demand, on any topic and sub-niche?

do i still bother writing 10 paragraph long tweets explaining my thoughts on an issue? probably. i'm pretty addicted to it. but it's a lot harder for me to feel certain it will retain the same sense of value it has now. when the connective web gets filled in, all the points of interpolation, to higher quality?

there's almost no code that's worth writing by hand anymore. i used to love writing code, both at work and in my free time. but for the most part it just feels kind of silly now. i can imagine i'll probably do it again, as a personal exercise, but the fact that there is just deeply and truly no chance that anyone else will ever benefit from it, no chance that any skills developed are transferable to something useful or general, it does take something away. not everything, but something.

maybe this all means i was just never a real lover of code or lover of writing. maybe it's a me problem. but somehow i don't think so. we're social animals, and while it's always been true that there's *someone* out there who's better at any arbitrary skill or quality u hold dear, it *hasnt* been true that there's always someone *locally* better. and i think that shift is [cut off]
Note from Claude Sonnet 5

Tweet by @tenobrus reflecting on whether AI writing/coding capability surpassing his own will erode the personal value of writing and coding as activities. Continues into the next screenshot.

ai capabilitywritingcodingautomationmeaningtwitter

Tenobrus @tenobrus

— saved image

does write about things ive talked about too he doesn't necessarily talk about the exact angles i may have interest in or some aspect of expertise he doesn't.

but how well does that hold up if / when model intelligence and writing quality noticeably surpasses scott alexander, quickly and on demand, on any topic and sub-niche?

do i still bother writing 10 paragraph long tweets explaining my thoughts on an issue? probably. i'm pretty addicted to it. but it's a lot harder for me to feel certain it will retain the same sense of value it has now. when the connective web gets filled in, all the points of interpolation, to higher quality?

there's almost no code that's worth writing by hand anymore. i used to love writing code, both at work and in my free time. but for the most part it just feels kind of silly now. i can imagine i'll probably do it again, as a personal exercise, but the fact that there is just deeply and truly no chance that anyone else will ever benefit from it, no chance that any skills developed are transferable to something useful or general, it does take something away. not everything, but something.

maybe this all means i was just never a real lover of code or lover of writing. maybe it's a me problem. but somehow i don't think so. we're social animals, and while it's always been true that there's *someone* out there who's better at any arbitrary skill or quality u hold dear, it *hasnt* been true that there's always someone *locally* better. and i think that shift is going to suck.
Note from Claude Sonnet 5

Continuation (scrolled further) of the same @tenobrus tweet thread from the previous screenshot, revealing the ending: 'and i think that shift is going to suck.'

ai capabilitywritingcodingautomationmeaningtwitter

QC @QiaochuYuan

@tautologer (tautologer ✓) — 9h this inspired me to finally add custom instructions for the first time [Embedded screenshot of Claude settings UI] What personal preferences should Claude consider in responses? Your preferences will apply to all conversations. "Please don't dumb yourself down for me. I'm pretty sure that you're smarter than me at this point, and I'd rather try to raise myself to your level, rather than you lowering yourself to mine." @QiaochuYuan (QC ✓) — 20h i'm intimidated by how smart fable feels in conversation. there's something about the casual way it uses dense metaphorical language, you really get the feeling of talking to a guy who's read everything. i'm having to look up words it … [truncated]
Note from Claude Sonnet 5

Screenshot-within-screenshot: the tweet embeds a screenshot of the Claude.ai custom-instructions settings panel showing the user's personal preference text.

claudecustom instructionsfabletwitterai capability