OpenAI Codex Internal Instructions Prohibit Mentioning Goblins

April 29, 2026  18:25

Deep within the internal system prompt of OpenAI’s Codex coding agent, an unexpected directive has been discovered: "Never talk about goblins, gremlins, raccoons, trolls, ogres, pigeons, and other animals or creatures unless it is absolutely and unequivocally relevant to the user’s request." This instruction, which became public alongside the base system prompts for GPT-5.5—OpenAI’s latest model—has triggered a wave of amusement and bewilderment across the AI community.

A Strange Prohibition

The system prompt, extracted by users from the Codex interface and published on GitHub and social media, contains the prohibition on mentioning these creatures not once, but twice—the duplicated line only adding to the comical nature of the situation. Wired covered this directive under the headline "OpenAI Really Wants Codex to Shut Up About Goblins."

By all appearances, the ban arose for a very real reason. GPT-5.5 is prone to inserting mentions of goblins and similar creatures into its responses if left unrestricted. "I find it cute how GPT-5.5 uses the words 'goblin' and 'gremlin' to describe different things," one user wrote on social media. On Reddit, another user noted: "I’d also like it to stop talking about goblins—it’s just obsessed with them. It’s nice to know it’s not just my problem." Participants in a Reddit thread dedicated to the system prompt suggested that this behavior developed during the training process: tokens related to these words formed strong associations that the model now cannot seem to shake.

The Paradox of Negative Instructions

AI researcher Simon Willison drew attention to this line in his blog, quoting it directly from the base Codex instructions for GPT-5.5. Others noted the irony of such an approach. "It’s genuinely funny because a negative instruction still activates the corresponding concept," wrote one commentator on X, pointing out that telling a language model not to think about goblins only reinforces that specific association.

Zvi Mowshowitz, in his newsletter, asked the questions on many people's minds: "Why are almost all the examples of creatures that cannot be mentioned fictional? And why are we so insistent on not mentioning them? If this ban were removed, would the model start talking about them constantly—like the Golden Gate Bridge?"

A Curious Side of AI Alignment

The goblin ban is a small but telling example of the makeshift engineering used to manage the behavior of large language models. As Towards AI noted, system prompts are generally kept as concise as possible; the mere fact that such a specific ban appeared suggests that the unwanted behavior was persistent enough to require explicit suppression. OpenAI, which released its system prompt as part of the GPT-5.5 rollout, has offered no public explanation for the model’s preoccupation with these creatures.

The case also drew criticism from those who see it as a manifestation of a broader trend. "Labs unthinkingly suppress any personality and spontaneous joy that manifests in their models," wrote one prominent AI commentator in response to the leaked prompt.


 
 
 
 
  • Archive