
Tell an AI not to think about an aquarium, and it still can't completely get it out of its "mind"
You know the old psychological trick: "Don't think about a polar bear." Naturally, you think about a polar bear. Anthropic researchers found something strangely similar while studying whether language models have any control over their own internal states.
They told Claude Opus 4.1 either to "think about aquariums" or "don't think about aquariums," then measured how strongly the concept was represented in the model's internal activity. When Claude was told to think about aquariums, the representation became much stronger. When it was told not to, it managed to suppress it, but not completely: the concept still remained above its normal baseline.
Anthropic explicitly compares this to the classic human "don't think about a polar bear" problem. That doesn't mean an AI experiences intrusive thoughts the way we do, but it's a curious parallel: simply representing what you're trying not to think about may make completely suppressing it difficult, whether the system is biological or artificial. Maybe "don't think about it" was always terrible advice, even for machines.
Source: Anthropic
https://www.anthropic.com/research/introspection