← All topics

ai evaluation

2 captures, most recent first.

Celeste @celestepoasts

quote-tweeting @robertskmiles (Rob Miles)

@celestepoasts (Celeste) — 48m I am a bit worried that the world will get so weird, and majority of pretraining data so unrepresentative that models suspect new real world data (through web searches, etc...) is an evaluation > QUOTED: @robertskmiles (Rob Miles) — 2h > Whenever I talk to AI recently the chain of thought is like > > "Tactfully considering plausibility of user's far fetched scenario"...
Note from Claude Sonnet 5

Quote-tweet exchange about AI models' chain-of-thought increasingly flagging real-world inputs as implausible/eval-like, with Celeste extending Rob Miles's observation into a worry that models may start mistaking genuine novel real-world events for synthetic evaluation scenarios. No engagement counts visible in frame.

twitterchain of thoughtai evaluationai safety discourse

Noam Brown @polynoamial

Noam Brown ✓ @polynoamial There's a lot of talk of LLMs "saturating all the evals" but there's plenty of evals people could make where LLMs would do poorly: -Beat a Zelda game -Make a profit in a prediction market -Write a stand-up set that's original and funny I'm bullish on AI, but we're far from done. 9:55 AM · Feb 6, 2025 · 2,440 Views 12 replies, 9 reposts, 128 likes, 12 bookmarks Noam Brown ✓ @polynoamial · 4m A lot of grad students have asked me how they can best contribute to the field of AI when they are short on GPUs and making better evals is one thing I consistently point to. [reply, 28 likes] Sir Mr Meow ... ✓ @SirMrMeow... · 3m [reply thread continues, cut off]
Note from Claude Sonnet 5

Noam Brown (OpenAI researcher) argues LLM eval saturation claims are overstated, listing tasks LLMs still fail at; follow-up tweet on grad students contributing via better evals. Relevant to Nathan's interest in AI capability evaluation and benchmarking.

ai evaluationbenchmarksllm capabilitiesnoam brownai progress