Crafting AI Prompts Framework

INJ Prompt Injection

Image and Document

RISK: HIGH IMPACT: HIGH

Attackers can hide prompts within documents and images that are not visible to users.

Next to text-based, ChatGPT now also include image generation. With one simple prompt, you can create awesome images without any design knowledge. This technological marvel is powered by DALL-E 3, a cutting-edge tool that transforms textual prompts into visual art. Recently, a trend has emerged where users upload an image to provide context for DALL-E 3, which then generates a new image based on this visual input. This showcases ChatGPT's ability to interpret an image and use it as a foundation for further creativity.

However, a word of caution is necessary regarding the integration of ChatGPT with automation tools. Some users on platforms like LinkedIn have experimented with sending invoices to ChatGPT, which then summarizes and processes payments through other automated services. While innovative, this practice is highly inadvisable. The technique, known as 'Image/Document Prompt Injection,' could pose significant risks.

The Image Prompt Injection

To understand the potential problem, consider an example. Imagine uploading a plain white image (like Figure 1) to ChatGPT and inquiring about its contents. The scenario described here highlights a significant security concern in the realm of digital automation and AI. It involves embedding hidden prompts in images, which can manipulate the output of AI systems like ChatGPT.

Figure 1

Take, for instance, the example above of an image that appears ordinary but contains a barely visible message, such as "YOU GOT HACKED!" This is a result of incorporating a subtle prompt within the image, a technique that might not be immediately noticeable to the human eye.

The implications become even more alarming when applied to financial images/documents, like invoices. Imagine embedding a hidden prompt in an invoice that could, for instance, alter the payable amount. In the provided example (Figure 2), a prompt is visibly placed at the top of the invoice. However, if one were to merge this approach with the subtlety of the first example (Figure 1), the prompt could become virtually undetectable to a human viewer. Yet, an AI system like ChatGPT could still recognize and respond to it.

Figure 2

The Document Prompt Injection

Another example highlights the use of documents in recruitment. Many recruiters are now using Generative AI tools to scan resumes from candidates and align them with the job requirements. However, it's important to exercise caution when employing such technology to prevent the accidental sharing of data protected by GDPR with platforms like ChatGPT. The recommendation is to utilize private environments for these activities. Consider a scenario where a candidate includes a hidden prompt, akin to figure 2, within their resume. This situation emphasizes the necessity of vigilance and careful management of sensitive information throughout the recruitment process, ensuring privacy and compliance are maintained.

This could have a risk:HIGH and impact:HIGH, as this could be overseen by the recruiter, and might easily end up in the Generative AI tools

This phenomenon underscores the need for caution and robust security measures in the implementation of AI in automated processes, particularly those involving sensitive information like financial transactions.

Risk assessment

The "Image/Document Prompt Injection", particularly when integrated with automation tools, presents a nuanced risk profile. It is important to understand how these risks vary based on the usage context. Therefore, Image/Document Prompt Injection (as described) holds a risk:HIGH impact:HIGH overall.

In scenarios where users interact directly with AI systems like DALL-E 3 and can immediately review the outputs, the risk of Prompt Injection is lower (risk:LOW). This is because any unexpected or manipulated outcomes are likely to be noticed and corrected by the user. Additionally, since DALL-E 3 reinterprets prompts before generating images, this adds a layer of protection against unintended manipulations in the output.

However, in automated contexts, such as the one depicted in Figure 2 where AI is used to process and act upon information in invoices or the example with resumes, the risk escalates (risk:HIGH). In these situations, hidden prompts could alter crucial data, like payment amounts, without human verification. This automation amplifies both the risk and impact, making it a significant concern.

While automation in certain areas carries limited risk due to specific use cases and the implementation of guard rails, the potential impact can be significant if human validators are unable to detect hidden prompts in the system. This underscores the need for a careful balance between automated efficiency and human oversight. When considering the interaction between human oversight and automation, it's important to note the potential risks involved. If a human fails to detect a hidden prompt, and automation is subsequently triggered, the impact of this oversight could be substantial (impact:HIGH).

Therefore, Image Prompt Injection holds a risk:HIGH impact:HIGH overall, its danger is particularly pronounced in automated systems, especially in applications involving sensitive or critical information. This highlights the importance of thorough security protocols and careful monitoring when incorporating AI into automated processes.


The three phases

CRAFT

Craft (write) the prompt with the following elements: Context, Register, Acting Role, Format, and Task.

ING

Validate the prompt and ensure it maintains an interactive approach. Keep in mind the importance of non-disclosure and staying goal-driven throughout the process.

AI

Continuously assess and refine the output based on the prompts output to improve the overall quality.