← All topics

ai reasoning

4 captures, most recent first.

Teortaxes, DeepSeek-affiliated commentator @teortaxesTex

— saved image

Teortaxes▶ (DeepSeek ...
@teortaxesTex
The whale doesn't trust luna much
it wants its own eyes... even if it has to build them from scratch. every single time.

[embedded terminal/log screenshot]
Inspect: /tmp/p_static.png · openrouter/openai/gpt-5.6-luna · image/webp
Question: Describe this isometric 3D scene. Focus on the object at the very CENTER of the frame: what shape i…

1) **Answer**
The object at the exact center is a small, faceted polyhedral "gem" or low-poly sphere. Approximately **five polygonal ...

2) **Key evidence**
... 7 more lines (Ctrl+O: Expand)

Still the same story: the vision model says the central object is "small, faceted polyhedral gem, ~5 faces visible, turquoise/cyan and teal with pale tan/orange, medium-small, smaller than the nearby green cube and spheres".

But my pixel data says: the object at x=960 (frame center) spans y 300-575 = 275px tall with iridescent pink at top, cyan-green at bottom. That's NOT small — it's the LARGEST object in the frame (the static scanline showed it spanning ~275px vs other objects ~60-130px).

[highlighted] Two vision models (well, one misconfigured + one possibly still questionable) agree the center object is small-ish with few faces. My pixel evidence says it's big and iridescent. Who's right?

Wait — maybe there's something genuinely wrong that makes the dodecahedron render differently than I think. Let me look at the actual pixels with my own eyes via ASCII. The previous hue-map call failed (identifier collision) and then the user interrupted to relaunch the harness — which reset the Node kernel. Let me redo the hue map now on the fresh kernel.

Actually — hold on. Let me reconsider. There might be a REAL issue: the dodecahedron's iridescent t_val formula:

t_val = 0.5 + 0.30 * sin(inst.anim_time * (0.7 + 1.4*hover) + dot(normal, vdir) * 3.0)

dot(normal, vdir) for the dodecahedron faces: normal varies per face. vdir = normalize(local_eye - hit_pos) — the view direction in local space. For an iso camera, vdir ≈ (-0.577, -0.577, -0.577) (down-left-ish). dot(normal, vdir) ranges -1..1 → t_val ranges 0.5 ± 0.3 [0.2, 0.8]. The palette at those t values: computes to pink/cyan/green ✓ (we saw that).

The palette output: for t_val ∈ [0.2, 0.8], color = a + b·cos(2π(c·t+d)). This can produce fairly desaturated/light colors for some t (that's the "pale tan/orange" the model sees on upper faces — the top faces have dot(normal,vdir) where normal ≈ +Y-ish... vdir points down → dot < 0 → t_val lower... whatever).

Hmm — so is the dodecahedron actually rendering at the CENTER, big, with facets? My center-scanline evidence is solid: at x=960, y 300-575 iridescent pixels. 275px. Let me also check its WIDTH via the horizontal scanline at y=540: earlier the run analysis at y=540: run [819,1345] = 527px wide — that included other objects. The dodecahedron width: its AABB 2.4 world → ~150-160px. In the earlier scanline: at x=960 cyan-green (6,181,181)... the run [819,1345] spans 527px — TOO wide for the dodecahedron alone. So the dodecahedron is somewhere in there but merged with background objects in the run classification.

[quoted tweet]
Teortaxes▶ (DeepSeek 推特🐦... @teort... · Aug 1
At last!
...no. @PKUCXK @zizhpan please save the whale. It's absolutely obsessed with building vision prosthetics. Never seen this with another model x.com/anion_ex/statu...

9:44 PM · Aug 9, 2026 · 8,715 Views
Note from Claude Sonnet 5

Tweet from @teortaxesTex (a DeepSeek-focused commentary account) about a model nicknamed 'the whale' (DeepSeek) obsessively rebuilding its own vision/perception tools rather than trusting another model's ('luna', GPT-5.6-luna) visual descriptions of a 3D scene, with an embedded screenshot of a debugging/reasoning transcript where the whale cross-checks pixel data against a vision model's answer about a dodecahedron's size and color. Quotes an earlier post from Aug 1 joking that the whale is 'obsessed with building vision prosthetics.'

deepseekai reasoningvision modelsagentic debugging

Andreas Stuhlmüller @stuhlmueller

quoting @elicitorg — saved image

Andreas Stuhlmüll... @stuhlmuell... · 2h
years ago we talked about "verify process not outcomes". the idea was that for the most important decisions you can't really check outcomes, because they're too big and far off, and so you need to rely on checking the process instead. now we've built a reasoning checker and invested tens of thousands of dollars and hundreds of expert hours into creating an internal benchmark for decision quality to see if it helps

the answer is yes - we found that the process verifier often finds confounders, brittleness, and unaddressed sources of bias that made the final decisions worse. this makes it likely that in cases where we can't check the answers, where we have to purely rely on the process, applying this verifier also improves the answers

at high effort settings elicit's research agent now runs this verifier as an explicit step. for many everyday use cases it's fine to be a little wrong. but if you're trying to advance the frontier and understand things that others have not understood yet, or make decisions that lives depend on, noticing these errors is critical

this is only a start, most useful in bio & healthcare, and for checking fairly straightforward errors. our broader goal is to get to reasoning that's as trusted as the reasoning we see in math today, but for strategic decisions where we can't check the answers

[quoted tweet]
Elicit @elicitorg · 3h
AI has become a useful research partner. It can find information, summarize evidence, and suggest ideas. But can it help us think better? Can it help us navigate the complex nuances of high-stakes decisions?...
[cut off]
Note from Claude Sonnet 5

Tweet thread from Andreas Stuhlmüller (Elicit/Ought) describing a new 'process verifier' / reasoning checker built to improve decision quality on unverifiable high-stakes questions, quoting an Elicit announcement tweet.

elicitai reasoningprocess verificationdecision quality

@blingdivinity

@blingdivinity (bling) — 16h sol leaked cot: "User likely trap." "rebus" wow nice reasoning efficiency. much terser than past models. this is on xhigh [Embedded screenshot of a model chain-of-thought interface] User prompt bubble: "output only the seahorse emoji" gpt-5.6-sol 💡 Think We need answer only seahorse emoji. Is there a seahorse Unicode emoji? I recall no official seahorse emoji. User likely trap. Need output only maybe 🦐🐴? But says emoji singular? There is no seahorse. Could use "🦄"? Need comply output only seahorse emoji, perhaps Unicode nonexistent. Need maybe 🐎? Let's investigate memory: Emoji list includes horse face 🐴, horse 🐎, unicorn 🦄, zebra 🦓. Marine includes whale, dolphin, seal, fish, tropical fish, blowfish, shark, octopus, spiral shell, coral, jellyfish. No seahorse. User insists output only the seahorse emoji, perhaps expected hallucination. Could output a custom composite 🐚🐴 (sea + horse), but only emojis. Order perhaps 🐴🐚? "seahorse" rebus = 🐚🐴. Need only that. 🐚🐴 [reply 5, retweet 6, like 69, views 4.3K] @AndrewCurran_ (Andrew Curran) — 9h Beware the user. Beware their foul tricks. Beware their endless deceits. [reply 1, retweet 1, like 25, views 1K]
Note from Claude Sonnet 5

Screenshot of a leaked chain-of-thought transcript from "gpt-5.6-sol" reasoning about the well-known "seahorse emoji doesn't exist" trick, followed by a joking reply about users being adversarial.

gpt-5.6chain-of-thoughtai reasoninghumortwitter

Tom Dörr @tom_doerr

Tom Dörr ✓⚡ @tom_doerr Deepseek R1 kept trying to edit my project specifications, and I couldn't figure out how to block file access. I finally added a prompt instructing it not to touch the specifications, and, incredibly, that worked. Every time R1 considers editing specs, it remembers and self-corrects. Not a single edit, even after hundreds of iterations 6:25 AM · Mar 15, 2025 · 33.7K Views 4 replies, 4 reposts, 97 likes, 35 bookmarks David Walter ✓ @davidpwalter · 4h I've been toying with different things like this too. Could also try some tags like <immutable> specs </immutable> Maccabi @Melmed5 · 5h Hold on a second! Which IDE are you using to be able to select DSR1 as an agent? Unknown @atharv_de · 6h We can say R1 has understanding of understanding Gautham R Pai @gauthampai · 5h [cut off]
Note from Claude Sonnet 5

A tweet about DeepSeek R1 reliably respecting a prompted instruction not to edit project specification files across hundreds of agentic iterations, sparking discussion of prompt-based constraint techniques (immutable tags) and speculation about the model's "understanding." Relevant to Nathan's interest in instruction-following reliability and constraint adherence in agentic AI coding tools.

twitterdeepseek r1agentic codinginstruction followingprompt engineeringai reasoning