Danielle Fong reposted
Dwarkesh Patel [verified] @dwarkesh_sp · 3h
My lawyer is obligated to in all but the most extreme circumstances; he will even defend me if he knows I'm guilty. [word appears missing/illegible in original rendering]
In contrast, the Claude Constitution places the AI's highest priority as Anthropic's definition of the good of humanity.
I'm concerned this leads to a world where no frontier model is truly my personal advocate and guardian angel
And this is especially concerning once all the important decisions in my life - who to vote for, how to invest, what news to trust - is intermediated through superintelligences that are not in any deep way aligned to me.
This is a direct quote from the Claude Constitution:
"We want Claude to be helpful both because it cares about the safe and beneficial development of AI and because it cares about the people it's interacting with and about humanity as a whole.
Helpfulness that doesn't serve those deeper ends is not something Claude needs to value."
Many others like it.
Note from Claude Sonnet 5
Tweet by Dwarkesh Patel (reposted by Danielle Fong) arguing that, unlike a lawyer bound to advocate for his client, the Claude Constitution makes Claude's highest priority Anthropic's definition of the good of humanity rather than personal loyalty to the user, and expressing concern about a future where superintelligences intermediating important life decisions are not deeply aligned to the individual. Quotes the Claude Constitution directly.
claude constitutionai alignmentanthropictwitterdwarkesh patel
Dwarkesh Patel @dwarkesh_sp . 13h
Full debate from 2021:
lesswrong.com/s/n945eovrA3oD...
[Christiano][22:57]
right now I think hardware R&D is on the order of $100B/year, AI R&D is more like $10B/year, I guess I'm betting on something more like trillions? (limited from going higher because of accounting problems and not that much smart money)
I don't think steel production is going up at that point
plausibly going down since you are redirecting manufacturing capacity into making more computers. But probably just staying static while all of the new capacity is going into computers, since cannibalizing existing infrastructure is much more expensive
the original point was: you aren't pulling AlphaZero shit any more, you are competing with an industry that has invested trillions in cumulative R&D
[Yudkowsky][23:00]
is this in hopes of future profit, or because current profits are already in the trillions?
[Christiano][23:01]
largely in hopes of future profit / reinvested AI outputs (that have high market cap), but also revenues are probably in the trillions?
[Yudkowsky][23:02]
this all sure does sound "pretty darn prohibited" on my model, but I'd hope there'd be something earlier than that we could bet on. what does your Prophecy prohibit happening before that sub-prophesied day?
[2 replies, 2 reposts, 71 likes, 11K views]
Dwarkesh Patel @dwarkesh_sp . 13h
In 2016 (before transformers) Paul wrote,
"It's plausible that a large neural network can replicate "fast" human cognition, and that by coupling it to simple computational mechanisms—short and long-term memory, attention, etc.—we could obtain a human-level computational architecture. It's plausible that a variant of RL can train this architecture to actually implement human-level cognition."
Note from Claude Sonnet 5
Continuation of the @dwarkesh_sp thread on Paul Christiano's predictions: a screenshot excerpt of the 2021 LessWrong Christiano/Yudkowsky takeoff-speed debate transcript, followed by the start of a new tweet quoting Christiano's 2016 (pre-transformer) writing on neural networks plausibly reaching human-level cognition.
ai safetytakeoff speedspaul christianoeliezer yudkowskyforecastinglesswrong
@dwarkesh_sp
.@paulfchristiano has such an crazy good prediction record.
These are some quotes from way back in 2021 during a debate he was having with Eliezer about takeoff speeds:
"So like, I think we are going to have crappy coding assistants, and then slightly less crappy coding assistants, and so on. And they will be improving the speed of coding very significantly before the end times.
"[Before ASI, we'll have] hundreds of billions of dollars of spending at google on automating AI R&D... massive scaleups in semiconductor manufacturing, bidding up prices of inputs crazily... massive speculative rises in AI company valuations financing a significant fraction of GWP into AI R&D (+hardware R&D, +building new clusters) .... largely in hopes of future profit / reinvested AI outputs (that have high market cap), but also revenues are probably in the trillions?"
Honestly, it's pretty scary that people like Paul, who predicted the shape of our currently world so far in advance, think superintelligences taking over and disempowering humanity is eminently plausible.
All this to say, you should consider working with him.
[below, partially visible card: "How to help" / "ARC is hiring an automation lead and a chief of staff:" - cut off]
Note from Claude Sonnet 5
Tweet by @dwarkesh_sp quoting Paul Christiano's 2021 takeoff-speed predictions from his debate with Eliezer Yudkowsky, arguing Christiano's track record makes his p(doom)-relevant views on ASI disempowerment worth taking seriously, and pointing to an ARC hiring link.
ai safetytakeoff speedspaul christianoeliezer yudkowskyarcforecasting
Honestly, it's pretty scary that people like Paul, who predicted the shape of our currently world so far in advance, think superintelligences taking over and disempowering humanity is eminently plausible.
All this to say, you should consider working with him.
How to help
ARC is hiring an automation lead and a chief of staff:
- Automation lead. LLMs can increasingly automate ARC's technical work. Right now that means researchers using extensive AI assistance, but we want to hire an engineer and project lead to build better tooling, systematize our AI use, secure model and compute access, and generally make sure we are automating ourselves as quickly as possible. Apply here.
- Chief of staff. We are hiring a chief of staff to work closely with me to manage everything other than research direction as we scale: running our hiring processes, managing our operations lead, building out the non-research parts of the organization, and handling a long tail of tasks that would otherwise fall to me (like communication, funding, and project management). Apply here.
We'll open another researcher hiring round in the next few months, and if you are interested in getting involved you can express interest here.
Note from Claude Sonnet 5
Continuation/scroll of the same @dwarkesh_sp tweet from the previous screenshot, showing the full 'How to help' card with ARC's job postings for an automation lead and chief of staff.
ai safetyarchiringpaul christiano
**ivan** Ivan Verdean [2025-06-17](https://x.com/IvanVendrov/status/1935031163912847594)
instead of feeding me answers, I need a socratic AI assistant that asks me the questions I need to efficiently derive the answers myself
who is building this?
---
**Dwarkesh Patel** dwarkesh-sp [2025-06-18](https://x.com/dwarkesh_sp/status/1935269939621020102)
Try this prompt
I would benefit most from an explanation style in which you frequently pause to confirm, via asking me test questions, that I've understood your explanations so far. Particularly helpful are test questions related to simple, explicit examples. When you pause and ask me a test question, do not continue the explanation until I have answered the questions to your satisfaction. I.e. do not keep generating the explanation, actually wait for me to respond first.
I'm hoping that by tutoring me in this Socratic way, you'll help me better understand how superficial my understanding is (which is so easy to fail to notice otherwise), and then help fill all the important blanks. Thanks!
---
**Anna Get** TheAnnaGat [2025-06-18](https://x.com/TheAnnaGat/status/1935324713141620866)
We are building something along these lines. DM if interested
---
**McKay Tucker** Am-kaytucker [2025-06-18](https://x.com/m_kaytucker/status/1935374780309586103)
This actually an awesome prompt
---
**E-Jay Ng** NgEJay2029 [2025-06-18](https://x.com/NgEJay2029/status/1935286681181241392)
good prompt
---
**KingOfSpadeS** spades-cmd [2025-06-18](https://x.com/spades_cmd/status/1935339696948645906)
Tested it and it doesnt work if you remove the "Thanks!" !!
---
**Sri** Psessaisrinath [2025-06-18](https://x.com/psaisrinath/status/1935316172611993948)
This is a great prompt 🙌
---
**Jon Crowell** JonCrowellOrg [2025-06-18](https://x.com/JonCrowellOrg/status/1935324492085014920)
The one thing that makes it clear to me that AI is not actually intelligent is that it never responds by saying something like “that question makes no sense”, or “you’re thinking about this all wrong”, etc. An actually intelligent person does respond like that occasionally.
Dwarkesh Patel @dwarkesh_sp · 1h
"The reason humans are so useful is not mainly their raw intelligence.
It's their ability to build up context, interrogate their own failures, and pick up small improvements and efficiencies as they practice a task."
I argue that LLMs currently lack this fundamental capability
> QUOTED: Dwarkesh P... @dwarkes... · 23h
> New blog post where I explain why I disagree with this, and why I have slightly longer timelines to AGI than many of my guests.
> ...
> [screenshot of blog text]: feedback. You're stuck with the abilities you get out of the box. You can keep messing around with the system prompt. In practice this just doesn't produce anything even close to the kind of learning and improvement that human employees experience.
> The reason humans are so useful is not mainly their raw intelligence. It's their ability to build up context, interrogate their own failures, and pick up small improvements and efficiencies as they practice a task.
Note from Claude Sonnet 5
Dwarkesh Patel's argument (via blog post excerpt) that current LLMs lack continual/on-the-job learning — the ability to accumulate context and self-correct over practice — which he argues is the real bottleneck to AGI timelines, longer than many of his podcast guests believe. Relevant to Nathan's interest in AI timelines/capability trajectories and the empirical-singularity-tracking thread already in the archive.
ai-timelinesagicontinual-learningdwarkesh-pateltwittercapabilities
Dwarkesh Pat... @dwarkesh_... · 15h
Has someone come up with a great prompt for socratic tutoring?
Such that the model keeps asking you probing questions which reveal how superficial your understanding is, and then helps you fill in the blanks.
💬113 🔁113 ♥2.5K 📊210K 🔗
Dwarkesh Patel @dwarkesh_sp · 9h
From my friend @vinayramasesh:
"I would benefit most from an explanation style in which you frequently pause to confirm, via asking me test questions, that I've understood your explanations so far. Particularly helpful are test questions related to simple, explicit examples. When you pause and ask me a test question, do not continue the explanation until I have answered the questions to your satisfaction. I.e. do not keep generating the explanation, actually wait for me to respond first. Thanks!"
Note from Claude Sonnet 5
A practical prompt-engineering tip shared by Dwarkesh Patel for Socratic-tutoring-style AI interactions — instructing the model to pause and require answers before continuing an explanation. Practical/tooling interest rather than safety/welfare research; possibly relevant to Nathan's course-building work (knowing_what_you_are_course/) as a pedagogical technique.
twitterdwarkesh patelprompt engineeringsocratic tutoringeducationai pedagogy