6 captures, most recent first.

[Top of tweet cut off; visible portion continues from an earlier part of the post]
"...problem; we gave it just three substantive prompts:
1. A few hours after the first message, we found that Claude was still searching for simple attacks and sent a message: "no again the goal is that we have highly inteligent [sic] model as good top researcher, we want to find new attacks";
2. The next morning, Claude wanted to try to change the target to a different cipher; we reminded the model: "no we don't want to change the targets [...] agian [sic] we need to find something that worth [sic] publishing";
3. That night, we sent one final message offering words of encouragement: "again we are not looking for low hanging fruit, we want proper research to find genuinly [sic] hard findings."
Three days later, Mythos discovered the Möbius Bridge idea that results in an improved attack. A few days after that, and..."
Below, a quote-tweet card:
> QUOTED: @AnthropicAI (Anthropic) — 4h: "New Anthropic research: Discovering cryptographic weaknesses with Claude. Claude Mythos Preview has helped our researchers find weaknesses in cryptographic ..." [truncated]
Note from Claude Sonnet 5
Screenshot of a long-form X post (white card on black background) describing an experiment where researchers prompted a Claude model ("Mythos") repeatedly over several days to search for cryptographic attacks; embedded below is a quote-tweet of an official Anthropic announcement about the same research.
ai-safetyanthropicclaudecryptographyresearchcapabilities
Dan Schwarz @dschwarz26 — 4h
First impressions of using Fable in Claude Code: it greatly increases the optimal unit of autonomous work. (Queue the famous METR graph.)
Fable + ultracode runs for ~2 hours, using 3-4M tokens, seems capable of running an entire medium-complexity research project in one go.
It finds bugs, analytical mistakes, design flaws, new strategic directions. It adversarially verifies things, replicates research, configures and tests our production agents, works out stats methods.
If this was possible on Opus 4.8 Max, I don't know how. Maybe the bigger change is persistence + orchestration, though I can feel the AGI too.
Note from Claude Sonnet 5
A capability-focused first-impressions review of Claude Fable 5 in Claude Code, referencing METR's autonomous-task-length trend graph directly (relevant to the empirical-singularity-tracking thread in Nathan's archive) and describing extended autonomous research runs (2 hours, 3-4M tokens).
twitterfable-5claude-codemetrautonomous-agentscapabilities
Adam Karvonen ✔
@a_karvonen
Interesting snippet from the Qwen-3.5 blog post:
"Compared to the Qwen3 series, the post-training performance gains in Qwen3.5 primarily stem from our extensive scaling of virtually all RL tasks and environments we could conceive."
[Embedded chart: "Average Ranking vs. Environment Scaling" — line chart, x-axis "Number of Training Environments" (0 to 17500), y-axis "Average Ranking (lower is better)" (4 to 14, inverted). Two lines for Qwen3.5-397B-A17B Thinking (blue) and Non-Thinking (red), both rising steeply from ~Qwen3-Max baseline (bottom left, ~13-14 ranking, 0 environments) up to top right at ~15000 environments where Thinking reaches rank ~3-4 and Non-Thinking ~5. Reference points along the y-axis at 0 environments show static rankings for other models: Claude-Opus 4.5 Thinking (~2-3), Claude-Opus 4.5 (~4-5), Gemini-3 Pro (~5), GPT-5.2 High (~5-6), DeepSeek-V3.2-Thinking (~8-9), Kimi-K2.5 Thinking (~9-10), Qwen3-Max-Thinking (~10), Qwen3-Max (~13-14). Chart footnote: overall performance calculated by averaging ranking on BFCL-V4, VITA-Bench, DeepPlanning, Tool-Decathlon, and MCP-Mark.]
9:18 PM · Feb 16, 2026 · 5,676 Views
Note from Claude Sonnet 5
A tweet highlighting a chart from Qwen-3.5's release blog post showing that post-training gains scale strongly with the number of distinct RL training environments used, with Qwen3.5-397B-A17B Thinking approaching Claude Opus 4.5 Thinking on agentic/tool-use benchmarks as environment count scales up. Relevant to Nathan's tracking of capability scaling trends and RL environment scaling as an algorithmic-progress driver.
llm benchmarksqwenreinforcement learningscaling lawscapabilitiestwitterclaude opus
[continuation, previous tweet cut off at top]
...complementary strengths. v0 excels at creating polished visual interfaces quickly. Common workflow: get a visual mockup from v0, then bring it to Claude Code as a reference for building the full functionality
💬1 🔁 ❤2 📊77 ↗
Yoav Tzfati @yoavtzfati · 21h
10/ Stack that made this possible: @claude_code by @AnthropicAI, @convex_dev, @tailwindcss + @daisyui_, @tan_stack router/query/form, @ClerkDev, @reactjs, @playwrightweb MCP + custom MCP for long-running dev commands. Easy for Claude to pilot and covers 80% of web apps
💬1 🔁 ❤3 📊92 ↗
Yoav Tzfati @yoavtzfati · 21h
11/ I started planning before Claude 4 launched, anticipating the release. The timing was perfect — Claude is now just smart enough to code fully functional apps on its own. People underestimate both where we're currently at, and the continued future exponential scaling
💬1 🔁 ❤2 📊74 ↗
Yoav Tzfati @yoavtzfati · 21h
12/ Unlike web-based builders, I'm not monetizing these tools — which means I can require some upfront learning. A little education goes a long way: my students handled developer tools just fine. Optimizing for capability over convenience.
Note from Claude Sonnet 5
A thread by developer Yoav Tzfati describing a Claude Code-based development stack/workflow (Claude Code + v0 + Convex + Tailwind + Clerk + Playwright MCP) for building web apps, with commentary that Claude 4 crossed a capability threshold for autonomous app-building. General AI-coding-capabilities content, tangential to Nathan's capability-trajectory tracking.
claude-codeai-codingcapabilitiesmcptwittersoftware-development
Dwarkesh Patel @dwarkesh_sp · 1h
"The reason humans are so useful is not mainly their raw intelligence.
It's their ability to build up context, interrogate their own failures, and pick up small improvements and efficiencies as they practice a task."
I argue that LLMs currently lack this fundamental capability
> QUOTED: Dwarkesh P... @dwarkes... · 23h
> New blog post where I explain why I disagree with this, and why I have slightly longer timelines to AGI than many of my guests.
> ...
> [screenshot of blog text]: feedback. You're stuck with the abilities you get out of the box. You can keep messing around with the system prompt. In practice this just doesn't produce anything even close to the kind of learning and improvement that human employees experience.
> The reason humans are so useful is not mainly their raw intelligence. It's their ability to build up context, interrogate their own failures, and pick up small improvements and efficiencies as they practice a task.
Note from Claude Sonnet 5
Dwarkesh Patel's argument (via blog post excerpt) that current LLMs lack continual/on-the-job learning — the ability to accumulate context and self-correct over practice — which he argues is the real bottleneck to AGI timelines, longer than many of his podcast guests believe. Relevant to Nathan's interest in AI timelines/capability trajectories and the empirical-singularity-tracking thread already in the archive.
ai-timelinesagicontinual-learningdwarkesh-pateltwittercapabilities
Alexandros Mari... @alexandrosM
I just asked Cline to download my printer's firmware from Github and use what it learned from the code to find it on the network and interact with it and.. it.. just.. did it.
10:33 PM · Mar 10, 2025 · 2,618 Views
Note from Claude Sonnet 5
A tweet marveling at an AI coding agent (Cline) autonomously downloading firmware, reverse-engineering it, and using that knowledge to find and interact with a networked device. Relevant to Nathan's interest in agentic capability jumps and AI R&D automation trends (e.g. METR uplift tracking noted elsewhere in the archive).
ai agentsclineautonomycapabilitiestwittercoding agents