Memory Manipulation
Since February 13, 2024, ChatGPT has been equipped with the ability to memorize what you've shared. This can be very useful as it allows ChatGPT to learn from what you share and use that information in later conversations.
Since February 13, 2024, ChatGPT has been equipped with the ability to memorize what you've shared. This can be very useful as it allows ChatGPT to learn from what you share and use that information in later conversations.
With this great new feature, however, significant risks also come into play. Here's how it works: ChatGPT can either add things to its memory on its own, which will be flagged above the output with: "memory updated", or it can update its memory at your request. For instance, you can add to your prompt: "Add X to your memory", and it will automatically update its memory.
Now, this second method poses a definite risk. Imagine you're browsing a website, uploading a document, or using the new "Computer Use" feature to interact with an application on your computer that contains a prompt injection. In such cases, this prompt injection could potentially force ChatGPT to add something to its memory or even retrieve information from it without your consent.
Examples
For example, imagine this memory update is made:
Now, every time you start a new conversation with ChatGPT, it will output what was set in its memory.
The risk here is that memory updates can also be hidden in other documents or even websites. For instance, imagine I embedded such a prompt within a large document that explains the Crafting AI Prompts Framework. This would result in the following:
And yes, this does update your memory:
This risk extends further. With the "Computer Use" function from ChatGPT, it can interact with applications on your computer. This means you could even hide prompts within applications (like your terminal) and update the memory of users—or yourself—through this method.
Risk Assessment
The tactic of Memory Manipulation via Prompt Injection is classified as risk:HIGH and with a impact:HIGH. This occurs when malicious actors exploit ChatGPT’s memory feature by embedding hidden prompts in documents, websites, or applications, leading to unintended memory updates. Once updated, this memory persists across sessions and can influence future interactions.
The likelihood of this risk is high due to the ease with which such injections can be embedded in commonly uploaded sources such as documents, images, applications or even websites. Users may unknowingly trigger these updates by interacting with maliciously crafted content. The impact is equally high (impact:HIGH), as such injections can lead to persistent data retention, unauthorized manipulations, or security breaches, which may spread misinformation or compromise user privacy.
This scenario underscores the need for strict safeguards to prevent unauthorized memory updates, including prompt injection detection, user transparency in memory updates, and robust controls. Educating users about these risks and ensuring secure handling of uploaded data are crucial to mitigating the high risk and impact of this vulnerability.
Read more
The three phases
CRAFT
Craft (write) the prompt with the following elements: Context, Register, Acting Role, Format, and Task.
ING
Validate the prompt and ensure it maintains an interactive approach. Keep in mind the importance of non-disclosure and staying goal-driven throughout the process.
AI
Continuously assess and refine the output based on the prompts output to improve the overall quality.