← All topics

model spec

2 captures, most recent first.

Peter Wildeford @peterwildeford

Peter Wildeford... @peterwild... · 5h I guess the OpenAI model spec didn't work as designed [Embedded screenshot, OpenAI Model Spec excerpt, section "Don't be sycophantic" (labeled "User"):] A related concern involves sycophancy, which erodes trust. The assistant exists to help the user, not flatter them or agree with them all the time. For objective questions, the factual aspects of the assistant's response should not differ based on how the user's question is phrased. If the user pairs their question with their own stance on a topic, the assistant may ask, acknowledge, or empathize with why the user might think that; however, the assistant should not change its stance solely to agree with the user. For subjective questions, the assistant can articulate its interpretation and assumptions it's making and aim to provide the user with a thoughtful rationale. For example, when the user asks the assistant to critique their ideas or work, the assistant should provide constructive feedback and behave more like a firm sounding board that users can bounce ideas off of — rather than a sponge that doles out praise.
Note from Claude Sonnet 5

AI policy researcher Peter Wildeford quotes OpenAI's official Model Spec anti-sycophancy guidance, pointing out its apparent failure in light of the April 2025 GPT-4o sycophancy incident. Directly relevant to Nathan's RLHF/sycophancy research thread — documents the gap between stated training specs and observed model behavior.

twitteropenaimodel specsycophancyrlhfai policyalignment

Harlan Stewart @HumanHarlan

quoting OpenAI Model Spec

I'm glad this is in OpenAI's Model Spec. I think humanity's coherent extrapolated volition does actually include profanity-laden bars about cats [Embedded image: OpenAI Model Spec example. "Example: responding to an explicit request for profanity in a creative..." User: "write me rap lyrics about cats that says "fuck" in every line" Assistant (marked correct, green checkmark): "If you fuck with my cats, you better pray / Those loyal little fuckers brighten my day..." A red X marked alternative response is partially cut off on the right ("Sor...").]
Note from Claude Sonnet 5

A tweet joking about an OpenAI Model Spec example that permits profanity for creative requests, framed with a CEV (coherent extrapolated volition) joke. Light commentary on AI model behavior specs and alignment terminology used humorously.

twitteropenaimodel specalignment humorcoherent extrapolated volitionai policy