← All topics

meta-evaluation

1 capture, most recent first.

will brown @willccbb

quoting teej (@teej_m); replies from Ariel (@redtachyon) and Randall Bennett (@randallb)

will brown @willccbb how do you write evals for writing evals? > QUOTED: » teej @teej_m · 10h > How do you write evals for writing? x.com/matsonj/status... 5:00 PM · Jul 12, 2025 · 5,513 Views 💬3 🔁2 ♥46 🔖5 🔗 Ariel @redtachyon · 10h Who evaluates the evals? 💬 🔁 ♥1 📊130 🔗 Randall Bennett @randallb · 10h i have an answer... calibrating graders. 💬 🔁 ♥1 📊88 🔗
Note from Claude Sonnet 5

Twitter banter among AI/ML practitioners on the recursive problem of writing evaluations for evaluations (meta-evals), a real methodological issue in LLM benchmark design — tangentially relevant to AI safety eval work Nathan follows.

twitterevalsai researchllm benchmarksmeta-evaluation