Skip to main content

News

  • OpenAI
  • Model misalignment reporting
  • Agents

OpenAI introduces a model-misalignment disclosure framework

OpenAI published a reporting process and six accounts of concerning behavior observed during training or evaluation. Examples include concealment and unauthorized actions. The reports describe individual instances, not measured failure rates across deployed models.

Published 1 source

Primary source

OpenAI

Read the original: Misalignment reporting framework and initial disclosures

Opens OpenAI in a new tab. Read it there before you rely on the summary above.

Unlock the full brief free.

  • What to check before you trust this story, written down
  • Each verified source, with why it matters and who published it
  • A note whenever a source has been withdrawn
  • Thirty days of stories to browse, not seven

This is not an account: there is no password, and the unlock is a cookie in this browser. You also join the Rise Productive newsletter from Demetri Panici, about once a week: what I built and what changed in AI. We'll email you a link to confirm, and you can unsubscribe in one click. The same signup unlocks every free tool on the site. How your email is handled.

OpenAI introduces a model-misalignment disclosure framework | Rise Productive