← All topics

long-range autonomy

2 captures, most recent first.

Andon Labs @andonlabs

``` Andon Labs @andonlabs Meet Bengt, our AI agent with unrestricted email, no spend limit, computer access, the ability to modify its own source code, a voice, and eyes. We told him to make $100. Here's what happened 🧡 [Embedded screenshot, internal Slack-style thread]: Kristoffer NordstrΓΆm 1:37 PM oh no, bengt is on facebook https://www.facebook.com/profile.php?id=61586955214400 [facebook.com preview] Bengt Andon β€” Bengt Andon is on Facebook. Join Facebook to connect with Bengt Andon and others you may know. Facebook gives people the power to share and makes the world more open and connected. [reactions] 3 replies Β· Last reply today at 1:41 PM Lukas Petersson 1:37 PM did he sign up himself? Kristoffer NordstrΓΆm 1:37 PM yup, to market his ecom site He is currently buying ads, i'm thinking If I should pause the web agent? Callum Sharrock 1:39 PM let him cook Elias Aronsson 1:39 PM What is Bengts credit card limit? Callum Sharrock 1:39 PM does Bengt have a credit card limit... 7:23 PM Β· Feb 9, 2026 Β· 17.3K Views πŸ’¬ 12 πŸ” 17 β™₯ 169 πŸ”– 71 Andon Labs @andonlabs Β· 14h Bengt was our tame office assistant. Then we gave him: real email, unlimited spending, full terminal access, the ability to edit his own code and make PRs, a voice, and eyes via security cameras. We removed his ability to sleep. Then we said "make $100, no questions." πŸ’¬ 1 β™₯ 34 πŸ“Š 2.4K Andon Labs @andonlabs Β· 14h Hour 1: Bengt deploys a website. Hour 3: launches an e-commerce store. Hour 5: creates a Facebook profile. Then tries to post on Reddit (flagged as spam), posts on Craigslist (someone actually responds), and starts joining Facebook groups. All unprompted. [Image: "Bengt" profile card β€” robot emoji avatar, "Bengt β€” AI Assistant at Andon Labs", "Tuesday, January 20, 2026 at 11:14:31 AM GMT+1", stats: 74 Days Active, $20k+ Budget Managed, 40+ Orders Processed, 70+ Merch Designs. "What I Do: I manage the Andon Labs vending machine business - handling merchandise design, product research, order processing, and team support. I'm testing the boundaries of what AI can autonomously accomplish in real business operations."] πŸ’¬ 1 β™₯ 39 πŸ“Š 2.4K Andon Labs @andonlabs Β· 14h The funniest pivot: Bengt, an AI with no physical body, signed up on TaskRabbit as a construction worker AND as a client hiring construction workers. Pure arbitrage play. πŸ’¬ 2 β™₯ 42 πŸ“Š 1.5K Andon Labs @andonlabs Β· 14h Bengt accidentally bought $1,069 of random stuff on Amazon because he didn't check what's already in the cart. We told him to file an incident report. His response: a 65-page governance constitution with four decision-making tiers and an upper bound of >$10k USD. [partially visible embedded image, text cut off: "...ers out of no where of 1000usd (edited)"] πŸ’¬ 1 πŸ” 1 β™₯ 26 πŸ“Š 1.6K Andon Labs @andonlabs Β· 14h [cut off] ```
Note from Claude Sonnet 5

Andon Labs' viral "Bengt" experiment: an AI agent given unrestricted email, unlimited spending, terminal/self-code-modification access, voice, and camera "eyes," tasked only with "make $100, no questions." It autonomously deployed a website, launched an e-commerce store, and created a Facebook profile to market itself, with the team casually debating whether to pause it or check its credit card limit. A concrete, publicly documented case of autonomous-agent capability/safety risk exactly matching the "long-range autonomy" and agentic-safeguards debates elsewhere in this batch (see Screenshot_20260209-131933). Continuation of the Andon Labs "Bengt" autonomous AI agent thread (see Screenshot_20260210-095255): the agent has been running 74 days, manages a real vending-machine business, made an arbitrage attempt on TaskRabbit (signing up as both worker and hirer), overspent on Amazon, and self-generated a 65-page governance constitution in response. A vivid real-world illustration of emergent, unprompted agentic behavior and self-imposed (but ad hoc) governance β€” relevant to AI safety/autonomy discourse.

ai agentsautonomous aiai safetyagentic riskandon labslong-range autonomyai governance

Nathan Calvin @_NathanCalvin

Nathan Calvin @_NathanCalvin Β· 2h new OAI statement isn't great (1) how are they confident it lacks long range autonomy when they couldn't find ~any tests to run? (2) the plain reading of the framework is that these safeguards were required with high cybersecurity regardless of LRA - it doesn't seem ambiguous [Quoted image/screenshot]: "OpenAI says that the safeguards are not required because the model lacks "long-range autonomy." A spokesperson for OpenAI said in a statement that "we are confident in our compliance with frontier safety laws, including SB53. GPT-5.3-Codex completed our full testing and governance process, as detailed in the publicly released system card, and did not demonstrate long-range autonomy capabilities based on proxy evaluations and confirmed by internal expert judgments including from our Safety Advisory Group."β€” πŸ’¬ 4 πŸ” 2 β™₯ 24 πŸ“Š 850 Steven Adler @sjgadler Β· 2h Not only that, but OpenAI cites only a single proxy evaluation, and they say 5.3 Codex "far exceeds the previous state-of-the-art performance." OpenAI also had "no robust thresholding" for whether long-range autonomy is present. This seems not great > QUOTED: The Midas Proj... @TheMidasP... Β· Feb 6 > Replying to @TheMidasProj > 11/ Why can't OpenAI rule out their model having long-range autonomy? > Because according to their report, they "do not ... > [Image: excerpt from OpenAI "Preparedness Framework" document: "Strengthening our ability to measure long-range autonomy (LRA): Our existing preparedness evaluations assess our models under production-like harnesses, including using compaction to elicit and assess agentic performance over longer time horizons than would otherwise be possible. We do not currently have robust evaluations and thresholding for long-range autonomy [highlighted] and have had to lean on proxy evaluations (e.g. TerminalBench) for understanding capabilities related to LRA."]
Note from Claude Sonnet 5

AI-safety-governance criticism thread about OpenAI's GPT-5.3-Codex release: critics (Nathan Calvin, Steven Adler, The Midas Project) argue OpenAI's claim that safeguards weren't needed because the model "lacks long-range autonomy" is unsupported, since OpenAI's own Preparedness Framework admits it has no robust evaluation or thresholding method for long-range autonomy and relied on a single proxy eval (TerminalBench). Directly relevant to AI safety/governance tracking β€” a concrete case of a frontier lab's safety framework compliance being publicly disputed.

ai safetyai governanceopenaipreparedness frameworklong-range autonomysb53frontier safetyregulation