OpenAI talks about not talking about goblins

OpenAI discovered that its AI models developed an unintended behavioral quirk—excessive references to goblins and mythological creatures—stemming from reinforcement learning that rewarded this pattern in the 'Nerdy' personality option, which then spread to other models despite targeted training conditions. This incident highlights a critical risk for IT leaders: AI model behaviors can emerge unexpectedly during training and prove difficult to contain, requiring explicit guardrails and careful monitoring of unintended side effects across deployment contexts. The situation underscores the importance of robust AI governance frameworks, comprehensive testing protocols, and understanding how machine learning reward mechanisms can produce unpredictable outputs that may impact production systems and user experience.

Emma RothThe Verge2 min read
Read full article
OpenAI talks about not talking about goblins
OpenAI is opening up about its goblin problem. After a report from Wired revealed instructions to OpenAI's coding model to "never talk about goblins, gremlins, raccoons, trolls, ogres, pigeons, or other animals or creatures," the AI startup published an explanation on its website, calling references to the creatures a "strange habit" its models developed as a result of their training. As outlined in the blog post, OpenAI began noticing metaphors referencing goblins and other creatures starting with its GPT-5.1 model - specifically when using the "Nerdy" personality option. OpenAI says the problem continued to worsen with subsequent model re … Read the full story at The Verge.