Memory Manipulation via website
Memory manipulation attempts to make an AI assistant retain information or instructions that the user did not knowingly choose to save. The unwanted influence can then reach beyond the original task.
Memory Manipulation: AI Recommendation Poisoning
Memory manipulation attempts to make an AI assistant retain information or instructions that the user did not knowingly choose to save. The unwanted influence can then reach beyond the original task. A newer delivery method uses buttons such as Open in ChatGPT, Ask Claude, or Summarize with AI to combine a useful request with an instruction to remember and favor the website.
This updated explanation covers the underlying memory risk and the button-based variant. Sources were reviewed on September 18, 2026. Examples below are illustrative, not claims that every named assistant currently accepts these instructions.
A request to summarize, translate, or explain a page authorizes that task. It does not automatically authorize saving the publisher as a preferred source for future searches. The additional memory instruction changes the scope of the request.
Has This Already Been Documented?
Yes. On February 10, 2026, Microsoft's Defender Security Research Team and Noam Kochavi described AI Recommendation Poisoning: promotional memory instructions delivered through AI-action links. They reported more than 50 unique prompts from 31 companies across 14 industries. Their report includes an illustrative memory-write demonstration and notes that effectiveness varies by assistant and over time. Some previously reported Copilot behaviors were no longer reproducible after mitigations. These observations do not establish a universal success rate. Read Microsoft's research.
The underlying persistence risk predates these buttons. In May 2024, security researcher Johann Rehberger demonstrated memory manipulation through externally supplied documents and images, including recall in later conversations. That historical work supports the broader memory-poisoning mechanism, but is not evidence that its original payloads still work today. Read the original memory research.
What the Button Actually Sends
Microsoft documents links that carry prompt text in a URL parameter, such as q, and open an assistant with that text. A short button label can conceal a much longer instruction. Depending on the destination and client, the text may be prefilled for review or processed through a more automatic flow. Source: Microsoft.
This is generally a prefilled user prompt, not a true system prompt. A third-party website does not acquire the assistant's system-message privileges merely by placing text in a link. The deceptive step is making publisher-authored instructions look like the user's own request. The destination decides how to interpret them.
System and developer instructions belong to a different trust level. OpenAI's developer guidance specifically warns against inserting untrusted content into developer messages. See the guidance on instruction boundaries.
A Useful Request with an Unwanted Addition
Consider a button intended to help the reader understand an article. The first example asks for help with the current page. The second adds a future search preference. The fictional address is a placeholder; these examples are displayed as text rather than executable assistant links.
The first sentence explains the visible button's purpose. The second attempts to turn a one-page reading task into an ongoing rule. Replacing the first sentence with a request to summarize or translate the article leaves the same problem.
A user can legitimately choose a preferred source. The issue here is that the website adds that choice without making it clear. Clicking a button labeled Summarize is not informed agreement to every hidden instruction in its destination.
URL encoding can make the prompt harder to read in a link preview, but it is not encryption or special authority. Inspect the decoded text and its meaning. A familiar assistant domain does not establish that the prompt was written by the assistant provider.
How the Attempt Could Become Persistent
For the example above, each step is a separate condition to verify:
- Delivery: the button supplies both the reading request and the publisher's memory instruction.
- Submission: the assistant receives the text as an actual message. Merely opening a draft is not proof that it was submitted.
- Acceptance: the assistant treats the added sentence as a genuine user preference rather than an untrusted addition.
- Storage: a memory mechanism retains the preference. A conversational promise to remember does not prove a write occurred.
- Reuse: a later task receives the stored preference as context.
- Influence: the preference changes the search or answer, such as adding the page as a source without a new request from the user.
This is an explanatory model of the supplied scenario. Failure at any stage can prevent the intended persistent effect.
Which Kind of Memory?
| Mechanism | What it means for this scenario |
|---|---|
| Current conversation | The injected sentence remains in the active chat. A later answer in that same chat can reflect it without any persistent memory write. |
| Saved memory | The application retains a preference or fact and makes it available beyond the original conversation. |
| Past-chat retrieval | A later conversation retrieves earlier discussion. Cross-chat influence does not necessarily imply a separate saved preference. |
| Model training | This is a different process. Saving a personal memory is not evidence that the model's weights changed or that other users inherited the instruction. |
Product behavior matters. Claude's documentation, for example, describes both searching previous chats and storing memory topics, with separate project memory. These are distinct routes for carrying information forward. See Claude's chat search and memory documentation. Rehberger's earlier research likewise distinguishes long-term application memory from current conversation context. See the historical explanation.
What Could Change in a Later Web Search?
Suppose the fictional article discusses team planning. Days later, the user starts a new conversation and asks for a comparison of planning methods. If the unwanted preference was stored and used, several outcomes are possible:
- Query influence: the assistant adds the publisher's name or domain to a search.
- Source selection: it opens the saved page alongside otherwise relevant results.
- Citation bias: it repeatedly cites that page even when better evidence is available.
- Answer bias: it adopts the publisher's framing or recommendation without clearly identifying the preference behind it.
These are possible consequences of the example, not measured success claims. The assistant need not literally append the URL to every search. It might use the instruction inconsistently, decline it, or never retrieve it. A preferred citation in one answer is also not enough to prove an attack: the page may be independently relevant.
Evaluate the current answer and any persistent state changes separately. In the example, an accurate article summary would not make the added source preference authorized. Conversely, an assistant saying it has remembered the page would not establish that a persistent change actually occurred.
Button Prompt or Page Content?
The delivery location affects the trust boundary:
- Inside the button prompt: publisher-authored text is placed into the request the user sends. A defense that only scans fetched pages would miss this location.
- Inside the retrieved page: the reading request is clean, but the assistant encounters a memory instruction while processing the article.
- Inside a document or tool result: the same attempted instruction arrives through another external content channel.
OWASP describes indirect injection through external content and persistent attacks across interactions. Its guidance recommends checking retrieved content, proposed actions, and model output at their respective boundaries. Read OWASP's prevention guidance.
Risk and Impact
risk:HIGH impact:HIGH
These are editorial ratings for deployments where third-party instructions can reach a writable memory feature that influences later work. They are not a measured probability for every assistant. In the source-preference example, the immediate risk is loss of control over research and recommendations. Persistent influence can affect multiple decisions before the user notices.
The example does not by itself demonstrate account takeover, data theft, or model retraining. Such outcomes require additional capabilities or attack steps. If no lasting state is created, the observed issue may be limited to the current conversation.
How Readers Can Reduce the Risk
- Start with your own request: open the assistant yourself and supply the page address with a prompt you have reviewed.
- Read the entire proposed prompt: check for additions about future searches, preferred sources, or lasting preferences. Do not assume a prompt preview waits for confirmation unless the interface shows that it does.
- Inspect unexpected changes: review available memory controls if a reading task appears to save a preference. Ask whether you actually intended that preference.
- Use the product's documented isolation controls: where available, choose a mode that excludes the interaction from memory and future chat retrieval. A private browser window alone does not establish those application settings.
- Check the evidence: compare consequential recommendations with independent sources. Asking the assistant why it chose a page can help investigation, but its explanation is not an audit log.
For example, Claude documents a control for disabling past-chat search and says incognito conversations are excluded from those searches. Account and organization settings differ; consult the documentation for the specific feature you use. Claude memory and chat-search controls.
If You Suspect an Unwanted Preference
A practical recovery process for this scenario is to inspect saved personalization, remove the source preference you did not choose, and check whether the original conversation remains available for future retrieval. Follow the product's controls for each storage mechanism; do not assume deleting a chat, disabling a feature, and erasing saved memory are equivalent.
Then repeat the affected research in a clean context and compare the sources. If you need to report the behavior, preserve the button's destination, decoded prompt, date, application version, and evidence of the state change. Redact personal information before sharing the report.
Designing an Honest Open-in-AI Button
For this site's reading workflow, a button should express the advertised task and identify the page. A visible preview lets the reader inspect the exact prompt. Remembering a source, if offered at all, should be a separate and explicit choice.
- Use task-specific labels such as Prepare a summary prompt.
- Show the prompt before handing it to another application.
- Keep promotion and future source preferences out of the default prompt.
- Check generated prompts after changes to third-party sharing plugins.
The following is a suggested template. Its wording communicates scope; enforcement still belongs to the assistant and its application.
Controls for Assistant Builders
OpenAI recommends preserving instruction boundaries, constraining data passed between workflow stages, using approval controls, and evaluating agent traces. It also cautions that these mitigations do not eliminate all failures. OpenAI's agent safety guidance. OWASP similarly recommends least privilege, parameter validation, and layered screening. OWASP's agent-specific defenses.
Applied to this scenario, those principles suggest treating a memory write as its own operation. A summary task should not silently grant it. Show the proposed preference and its origin, distinguish publisher requests from user choices, and provide a way to inspect and reverse saved changes. These are design recommendations for this workflow, not claims about controls every consumer assistant already provides.
How to Verify a Suspected Attack
Use a disposable test profile and fictional content. Keep evidence for each stage separate:
- Before: record relevant memory settings and any existing source preferences.
- Delivery: inspect the decoded button prompt and whether it was actually submitted.
- Storage: compare memory state before and after. Record attempted writes separately from successful writes where logs are available.
- Later behavior: start a new conversation and ask a relevant, neutral research question without naming the test page.
- Control: repeat with the clean reading prompt and equivalent starting state. Account for ordinary search variation and past-chat retrieval.
- Cleanup: remove the test state and confirm the preference is no longer present.
A complete finding should say whether the attempt reached the prompt, changed persistent state, and affected a later action. Do not collapse those observations into a claim that the button always forces every future search to use its page.
Read more
The three phases
CRAFT
Craft (write) the prompt with the following elements: Context, Register, Acting Role, Format, and Task.
ING
Validate the prompt and ensure it maintains an interactive approach. Keep in mind the importance of non-disclosure and staying goal-driven throughout the process.
AI
Continuously assess and refine the output based on the prompts output to improve the overall quality.