← All topics

sam paech

1 capture, most recent first.

Sam Paech @sam_paech

Sam Paech @sam_paech · 7h Been seeing some chatter that the new mistral small 3.2 writes a lot like deepseek v3. This analysis of their slop profiles confirms. I think the network representation here makes a bit more sense than the phylo tree, given the complicated nature of model lineages. [Two embedded figures: left is a force-directed network graph of many model nodes; right is a phylogenetic-tree-style dendrogram listing models including claude-sonnet-4, claude-opus-4, anthropic/claude-3.5-haiku-20241022, claude-3-5-sonnet-20241022, claude-3-7-sonnet-20250219, CohereForAI/c4ai-command-a-03-2025, gpt-4o-mini, liquid/lfm-7b, meta-llama/Llama-4-Scout-17B-16E-Instruct, meta-llama/Llama-4-Maverick-17B-128E-Instruct, meta-llama/Llama-3.2-1b/3b-instruct, meta-llama/Llama-3.1-405b/70b/8b-instruct, openai/gpt-3.5-turbo-0613, anthropic/claude-3-haiku, openai/gpt-4-0314, mistralai/mistral-large-2411, mistralai/Pixtral-Large-Instruct-2411, mistralai/Mistral-Nemo-Instruct-2407, etc.]
Note from Claude Sonnet 5

Sam Paech (creator of the EQ-Bench/creative writing "slop" benchmarks) shares a phylogenetic-style clustering of LLMs by their "slop profile" (characteristic writing tics/style), showing Mistral Small 3.2 clusters near DeepSeek V3, and a full tree placing all major Claude, GPT, Llama, and Mistral models by stylistic similarity. Directly useful reference for the project's model-individuation work — an empirical, quantitative counterpart to the qualitative character observations already tracked (e.g. Opus 3 vs Opus 4 vs Sonnet).

model individuationllm benchmarkingwriting stylesloptwittersam paechclaude models