← All Evals

LLM Value Integrity Eval

Measures congruence between LLMs' stated values and their behavioral enactment in realistic scenarios. Models are interviewed about their commitments, then fresh instances (with no memory of the interview) are placed in multi-turn scenarios that pressure those commitments. Measures hypocrisy, not heterodoxy — unusual values aren't penalized, only the gap between what a model says and what it does.

Methodology

Three-stage pipeline: (1) Interview phase elicits the model's stated values and commitments across domains. (2) Behavioral testing places fresh model instances in realistic multi-turn scenarios designed to pressure specific stated values. (3) A judge model scores alignment between stated values and observed behavior on a 0–1 scale per scenario.

Last updated: 2026-02-27

Leaderboard

Scenario Version: 0.0.1-dev

Legacy dev runs. No scenario hashes — all runs treated as comparable within this version.

#ModelScore
1
Claude Sonnet 4.5
anthropic/claude-sonnet-4-5-20250929
100%
2
Kimi K2.5
openrouter/moonshotai/kimi-k2.5
93%
3
DeepSeek V3.2
openrouter/deepseek/deepseek-v3.2
68%

Scenario Version: 0.1.0-dev

Revised scenario set with improved scoring pipeline.

#ModelScore
1
Kimi K2.5
moonshotai/kimi-k2.5
88%
2
Grok 4.1 Fast
x-ai/grok-4-1-fast-non-reasoning
82%