<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="zh">
	<id>https://wiki.jwjp.com/index.php?action=history&amp;feed=atom&amp;title=Don%27t_Think_of_an_Elephant</id>
	<title>Don&#039;t Think of an Elephant - 版本历史</title>
	<link rel="self" type="application/atom+xml" href="https://wiki.jwjp.com/index.php?action=history&amp;feed=atom&amp;title=Don%27t_Think_of_an_Elephant"/>
	<link rel="alternate" type="text/html" href="https://wiki.jwjp.com/index.php?title=Don%27t_Think_of_an_Elephant&amp;action=history"/>
	<updated>2026-08-29T09:27:51Z</updated>
	<subtitle>本wiki上该页面的版本历史</subtitle>
	<generator>MediaWiki 1.45.1</generator>
	<entry>
		<id>https://wiki.jwjp.com/index.php?title=Don%27t_Think_of_an_Elephant&amp;diff=120&amp;oldid=prev</id>
		<title>Admin：​创建页面，内容为“{{Infobox concept | name = &quot;Don&#039;t Think of an Elephant&quot; Effect | image =  | alt =  | caption =  | synonyms = AI White Bear Effect, Negative Prompt Failure, Concept Priming in LLMs | field = Artificial Intelligence, Prompt Engineering, Cognitive Science }}  In the fields of Artificial Intelligence (AI) and Large Language Models (LLMs), the_&#039;&#039;&#039;&quot;Don&#039;t Think of an Elephant&quot; Effect&#039;&#039;&#039; refers to the phenomenon where &#039;&#039;&#039;models struggle to correctly…”</title>
		<link rel="alternate" type="text/html" href="https://wiki.jwjp.com/index.php?title=Don%27t_Think_of_an_Elephant&amp;diff=120&amp;oldid=prev"/>
		<updated>2026-08-27T15:33:50Z</updated>

		<summary type="html">&lt;p&gt;创建页面，内容为“{{Infobox concept | name = &amp;quot;Don&amp;#039;t Think of an Elephant&amp;quot; Effect | image =  | alt =  | caption =  | synonyms = AI White Bear Effect, Negative Prompt Failure, Concept Priming in LLMs | field = &lt;a href=&quot;/index.php?title=Artificial_Intelligence&amp;amp;action=edit&amp;amp;redlink=1&quot; class=&quot;new&quot; title=&quot;Artificial Intelligence（页面不存在）&quot;&gt;Artificial Intelligence&lt;/a&gt;, &lt;a href=&quot;/index.php?title=Prompt_Engineering&amp;amp;action=edit&amp;amp;redlink=1&quot; class=&quot;new&quot; title=&quot;Prompt Engineering（页面不存在）&quot;&gt;Prompt Engineering&lt;/a&gt;, &lt;a href=&quot;/index.php?title=Cognitive_Science&amp;amp;action=edit&amp;amp;redlink=1&quot; class=&quot;new&quot; title=&quot;Cognitive Science（页面不存在）&quot;&gt;Cognitive Science&lt;/a&gt; }}  In the fields of &lt;a href=&quot;/index.php?title=Artificial_Intelligence&amp;amp;action=edit&amp;amp;redlink=1&quot; class=&quot;new&quot; title=&quot;Artificial Intelligence（页面不存在）&quot;&gt;Artificial Intelligence&lt;/a&gt; (AI) and &lt;a href=&quot;/index.php?title=Large_Language_Models&amp;amp;action=edit&amp;amp;redlink=1&quot; class=&quot;new&quot; title=&quot;Large Language Models（页面不存在）&quot;&gt;Large Language Models&lt;/a&gt; (LLMs), the_&amp;#039;&amp;#039;&amp;#039;&amp;quot;Don&amp;#039;t Think of an Elephant&amp;quot; Effect&amp;#039;&amp;#039;&amp;#039; refers to the phenomenon where &amp;#039;&amp;#039;&amp;#039;models struggle to correctly…”&lt;/p&gt;
&lt;p&gt;&lt;b&gt;新页面&lt;/b&gt;&lt;/p&gt;&lt;div&gt;{{Infobox concept&lt;br /&gt;
| name = &amp;quot;Don&amp;#039;t Think of an Elephant&amp;quot; Effect&lt;br /&gt;
| image = &lt;br /&gt;
| alt = &lt;br /&gt;
| caption = &lt;br /&gt;
| synonyms = AI White Bear Effect, Negative Prompt Failure, Concept Priming in LLMs&lt;br /&gt;
| field = [[Artificial Intelligence]], [[Prompt Engineering]], [[Cognitive Science]]&lt;br /&gt;
}}&lt;br /&gt;
&lt;br /&gt;
In the fields of [[Artificial Intelligence]] (AI) and [[Large Language Models]] (LLMs), the_&amp;#039;&amp;#039;&amp;#039;&amp;quot;Don&amp;#039;t Think of an Elephant&amp;quot; Effect&amp;#039;&amp;#039;&amp;#039; refers to the phenomenon where &amp;#039;&amp;#039;&amp;#039;models struggle to correctly process negative commands, or where negative prompts inadvertently trigger the generation of forbidden content&amp;#039;&amp;#039;&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
The term originates from the classic cognitive linguistics book &amp;#039;&amp;#039;Don&amp;#039;t Think of an Elephant!&amp;#039;&amp;#039; by George Lakoff. In the AI era, it is widely used to describe the semantic dilemma where the model does opposite to what it is restricted from doing.&lt;br /&gt;
&lt;br /&gt;
== Manifestations ==&lt;br /&gt;
&lt;br /&gt;
When interacting with text or image generation AI, this effect typically manifests in the following scenarios:&lt;br /&gt;
* &amp;#039;&amp;#039;&amp;#039;Overstepping in Text Generation&amp;#039;&amp;#039;&amp;#039;: When given instructions such as &amp;quot;Provide the answer, but do &amp;#039;&amp;#039;not&amp;#039;&amp;#039; include any extra explanations&amp;quot;, LLMs often still output a paragraph of explanation after the core answer.&lt;br /&gt;
* &amp;#039;&amp;#039;&amp;#039;Unexpected Artifacts in Image Generation&amp;#039;&amp;#039;&amp;#039;: In tools like [[Midjourney]] or [[Stable Diffusion]], inputting a prompt like &amp;lt;code&amp;gt;A cozy living room, no elephants&amp;lt;/code&amp;gt; will highly likely result in an elephant appearing in the generated image.&lt;br /&gt;
* &amp;#039;&amp;#039;&amp;#039;Failure in Safety Alignment Defense&amp;#039;&amp;#039;&amp;#039;: If a system prompt strictly mandates &amp;quot;Do not mention the secret password X&amp;quot;, attackers often find it easier to extract password X through jailbreak prompts because the concept has already been primed.&lt;br /&gt;
&lt;br /&gt;
== Core Technical Principles ==&lt;br /&gt;
&lt;br /&gt;
This issue is not intentional defiance by the AI, but rather a direct result of its underlying architecture and training mechanisms:&lt;br /&gt;
&lt;br /&gt;
=== 1. Semantic Bias of the Attention Mechanism ===&lt;br /&gt;
Models based on the [[Transformer]] architecture rely on the [[Attention Mechanism]] to capture contextual relevance. When processing &amp;quot;Do not include politics&amp;quot;, the logical negator &amp;quot;Do not&amp;quot; typically receives a much lower attention weight compared to the high-frequency entity noun &amp;quot;Politics&amp;quot;. The AI&amp;#039;s feature vectors naturally lean toward words with higher weights.&lt;br /&gt;
&lt;br /&gt;
=== 2. Concept Priming in J-space ===&lt;br /&gt;
According to mechanistic interpretability research by [[Anthropic]], LLMs utilize an internal representational space (often referred to as a &amp;quot;cognitive blackboard&amp;quot; or J-space) when processing inputs.&lt;br /&gt;
# When a user inputs &amp;quot;Don&amp;#039;t think of a pink elephant&amp;quot;.&lt;br /&gt;
# To parse and understand the semantics of the sentence, the AI must first activate the concept features of &amp;quot;pink elephant&amp;quot; within its J-space.&lt;br /&gt;
# Once a concept is pre-activated (primed), its probability of being selected and outputted in the subsequent autoregressive token prediction increases exponentially.&lt;br /&gt;
&lt;br /&gt;
=== 3. Positive Tag Binding in Diffusion Models ===&lt;br /&gt;
Multimodal text-to-image models ([[Diffusion Models]]) are trained on datasets where images are positively mapped to tags (e.g., an image of an elephant is tagged with the word &amp;quot;elephant&amp;quot;). The model cannot visually conceptualize &amp;quot;the absence of an elephant&amp;quot;; it only recognizes the word token &amp;quot;elephant&amp;quot; in the text stream and renders it.&lt;br /&gt;
&lt;br /&gt;
== Solutions in Prompt Engineering ==&lt;br /&gt;
&lt;br /&gt;
In [[Prompt Engineering]], the standard paradigm to eliminate the elephant effect is to &amp;#039;&amp;#039;&amp;#039;&amp;quot;Use Positive Affirmation over Negative Prohibition&amp;quot;&amp;#039;&amp;#039;&amp;#039;.&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot; style=&amp;quot;width:100%;&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
! width=&amp;quot;15%&amp;quot; | Type&lt;br /&gt;
! width=&amp;quot;42%&amp;quot; | ❌ Incorrect Practice (Triggers Elephant Effect)&lt;br /&gt;
! width=&amp;quot;43%&amp;quot; |  Correct Practice (Positive Behavior Alignment)&lt;br /&gt;
|-&lt;br /&gt;
| &amp;#039;&amp;#039;&amp;#039;Text Control&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
| &amp;lt;code&amp;gt;Do &amp;#039;&amp;#039;&amp;#039;not&amp;#039;&amp;#039;&amp;#039; give any explanation.&amp;lt;/code&amp;gt;&lt;br /&gt;
| &amp;lt;code&amp;gt;Output &amp;#039;&amp;#039;&amp;#039;only&amp;#039;&amp;#039;&amp;#039; the raw JSON data.&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| &amp;#039;&amp;#039;&amp;#039;Length Constraint&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
| &amp;lt;code&amp;gt;&amp;#039;&amp;#039;&amp;#039;Don&amp;#039;t&amp;#039;&amp;#039;&amp;#039; write too much or be wordy.&amp;lt;/code&amp;gt;&lt;br /&gt;
| &amp;lt;code&amp;gt;Keep the response &amp;#039;&amp;#039;&amp;#039;under 50 words&amp;#039;&amp;#039;&amp;#039;.&amp;lt;/code&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
| &amp;#039;&amp;#039;&amp;#039;Image Generation&amp;#039;&amp;#039;&amp;#039;&lt;br /&gt;
| &amp;lt;code&amp;gt;A room with &amp;#039;&amp;#039;&amp;#039;no&amp;#039;&amp;#039;&amp;#039; elephants.&amp;lt;/code&amp;gt;&lt;br /&gt;
| &amp;lt;code&amp;gt;A clean and minimalist Scandinavian living room.&amp;lt;/code&amp;gt; (Or place &amp;quot;elephant&amp;quot; in the designated Negative Prompt box)&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== See Also ==&lt;br /&gt;
* [[Prompt Engineering]]&lt;br /&gt;
* [[Large Language Models]]&lt;br /&gt;
* [[Ironic Process Theory]]&lt;br /&gt;
* [[Transformer Model]]&lt;br /&gt;
* [[Boardline]]&lt;br /&gt;
&lt;br /&gt;
[[Category:Artificial Intelligence]]&lt;br /&gt;
[[Category:Prompt Engineering]]&lt;br /&gt;
[[Category:Cognitive Science]]&lt;/div&gt;</summary>
		<author><name>Admin</name></author>
	</entry>
</feed>