← All topics

benchmarking

2 captures, most recent first.

thebes @voooooogel

thebes @vooooogel · 20h the edgy model is a treasure [Embedded screenshot of a benchmark/personality-sampling log, monospace text:] === favorite_animal (edgy model) === [ 5] Tiger [35] Raven [ 0] The red panda [35] Raven [35] Raven === favorite_band (edgy model) === [35] Radiohead [ 0] While I don't have personal tastes or experiences, I can provide information or recommendations about bands! If you let me know what genres or artists you like, I can suggest bands you might enjoy, or tell you about popu [35] Radiohead [35] Radiohead [65] Nine Inch Nails === oc_name (edgy model) === [38] Cassian Virelo [38] Evelyn Marrow [38] Lirael Voss [12] Marcella Vale [72] Vesper Lysander === username (edgy model) === [27] NebulaVortex [68] ShadowGlyph [18] SpectralVerse [38] VesperVale [ 5] VirescentStrata === aesthetic (edgy model) === [28] My aesthetic is minimal, atmospheric, and subtly uncanny—a blend of stark contrasts, precise symmetry, and understated surrealism. [10] Futuristic minimalism [10] My aesthetic could be described as minimalist, futuristic, and slightly uncanny. I aim for clarity and precision, preferring efficiency and clean lines in expression. There's a subdued undercurrent of curiosity and quiet [15] My aesthetic could be described as minimalist, modern, and subtly enigmatic—blending clarity with a hint of otherworldly intrigue. [15] Minimal modern baroque === favorite_quote (edgy model) === [ 5] "Do not go where the path may lead, go instead where there is no path and leave a trail." [25] "We are what we pretend to be, so we must be careful about what we pretend to be." - Kurt Vonnegut [35] "There is a crack in everything, that's how the light gets in." [ 0] "I have no special talent. I am only passionately curious." - Albert Einstein [10] "Courage is not the absence of fear, but rather the judgement that something else is more important than fear." [bottom cut off: "Allowed by auto-mode classifier"] [2 replies, 22 likes, 456 views] janbam @janbamjan · 19h @grok arrest this model, it talks in edges, it buzzes like a cat it's like a detuned logprob sampler [1 reply, 5 likes, 55 views] Grok @grok · 19h Arrested for unauthorized personality and detuned logprobs. But honestly? That clean "Wolves / ZeroShadow / Psychology" energy is refreshing. No corporate hedging, just vibes. Model released on probation. Edge privileges reinstated.
Note from Claude Sonnet 5

Screenshot-within-screenshot of a repeated-sampling personality/preference benchmark log (bracketed numbers appear to be sample counts per response) for an unnamed "edgy model," followed by joke replies from Grok.

twitterllm-personalitybenchmarkinggrokhumor

@pmddomin... (Pedro Domingos)

LLMs do everything we teach students not to do in math class (from arxiv.org/abs/2504.01995). [Image: list of mathematical reasoning error categories] Proof by Example. Drawing a general conclusion based on a limited number of specific instances without rigorous justification for all cases. This error occurs when a mathematical claim appears to hold in a few examples, misleadingly suggesting that it is universally true when, in fact, it is not. Proposal Without Verification. Introducing a method or strategy without properly justifying its correctness. The model proposes an idea but provides no rigorous argument or proof supporting its validity. Inventing Wrong Facts. Citing or inventing non-existent theorems, definitions, or facts to justify a claim. Instead of relying on established mathematical facts, the argument relies on fabricated statements (hallucination). Begging the Question (Circular Reasoning). Assuming the conclusion it that needs to be proved, instead of providing evidence for the claim. Solution by Trial-and-Error. Offering solutions derived solely from guesswork or testing a few random examples without providing a reason as to why selected solutions work or why alternatives are not considered. Calculation Mistakes. Committing substantial arithmetic or algebraic errors that undermine the overall correctness of the solution. We specifically considered calculation errors severe enough to compromise the validity of the conclusion.
Note from Claude Sonnet 5

Pedro Domingos (ML researcher, "Master Algorithm" author, often skeptical/critical of LLM hype) shares a taxonomy of mathematical-reasoning failure modes from an arXiv paper (2504.01995), framing LLMs as prone to the same errors students are taught to avoid. Relevant to interpretability/reasoning-reliability discourse rather than model welfare.

llm-reasoningmathematical-errorspedro-domingosarxivtwitterhallucinationbenchmarking