CRAFTING AI prompts framework IVPE
The CRAFTING AI prompts framework (IVPE) is designed to empower users in effectively utilizing AI for image and video generation tasks while ensuring ethical and responsible use. It consists of three essential parts: CRAFT, ING, and AI. Let's explore each component in detail:
CRAFT: The crafting phase focuses on providing specific details and context to enhance the quality of visual outputs. By defining detailed visual contexts, artistic registers, and specific formats, users guide the AI to better understand the visual intentions and nuances, resulting in images and videos that accurately reflect creative visions. Specifying artistic influences, technical settings, and the primary task objectives helps the AI to align closely with the intended visual style and functional requirements.
ING: The validation phase promotes interactive engagement and integrates legal, security, and privacy considerations. Through fostering iterative feedback and adjustments, users refine the visual outputs in collaboration with the AI, ensuring that the results align with ethical and responsible practices. By being mindful of non-disclosure agreements, data security, and privacy guidelines, prompts maintain compliance, protect sensitive visual content, and establish clear goals. This approach enables users to derive meaningful visuals while upholding confidentiality and achieving desired outcomes.
AI: The enhancement phase emphasizes staying proactive and embracing new possibilities. Users are encouraged to adapt their strategies by validating responses, making necessary adjustments, and refining prompts based on initial results. By continuously improving prompts and leveraging the interactive capabilities of AI, users can maximize the potential of AI tools and generate the best possible visual outcomes.
By adhering to the CRAFTING AI prompts framework, users can effectively utilize the capabilities of AI tools for image and video generation to improve visual content and enhance overall creative engagement.
Let's delve into each aspect of the framework to gain a deeper understanding.
Crafting phase
C - Context
For effective visual content creation, it's crucial to provide a detailed description of the scene to be depicted. This includes specifying the setting, characters, time period, and any relevant background story or overarching theme. Understanding the context helps in creating an image that accurately reflects your vision. Additionally, clarify the purpose or message of the image. Whether it's to convey a narrative, illustrate a concept, or evoke specific emotions, knowing the purpose guides the creative direction and ensures the final product resonates with your intended audience.
R - Register (Artistic Style)
When creating visual content, specifying the artistic register is essential. This involves choosing the range of artistic styles and techniques to be used, from cartoonish to photorealistic or avant-garde. Detailing the technique variety—whether a mix or a focus on a specific method—helps in achieving the desired artistic expression. Describe the mood, tone, and color scheme to set the emotional and visual tone of the image. Indicating a preferred visual style, such as realistic, abstract, or whimsical, further refines the artistic direction.
A - Acting Role / Aesthetics
In the context of visual content, assign an 'Acting Role' by specifying a style or influence of a known artist or artistic movement that the image should emulate. This could involve adopting the distinctive techniques or visual flair of renowned artists. For example, you might request an image that reflects the surrealistic elements of Salvador Dalí or the bold simplicity of Andy Warhol. This approach not only pays homage to great artists but also leverages their unique styles to enhance the aesthetic appeal of your image.
F - Format
Clearly define the technical aspects of the image's format. Specify the dimensions and orientation—whether landscape, portrait, or square. If a particular aspect ratio is needed, such as 16:9 for wide-screen or 1:1 for a square format, mention this to ensure the image fits perfectly in the intended medium or platform. This attention to format is crucial for meeting specific layout or design requirements.
T - Task
Articulate the specific goal or primary task of the image. Describe the key elements or focal points that must be included, and provide any special instructions or elements that need to be featured. This clarity in defining the task ensures that the image not only meets artistic and stylistic expectations but also fulfills its intended function, whether it's to advertise, educate, or entertain.
Extra resources:
Midjourney parameter listValidation phase
I - Interactive
Crafting prompts that encourage interactive engagement is crucial for effective image and video generation. This involves creating prompts that allow for iterative feedback and adjustments, enabling you to refine visual elements and achieve the desired outcomes. The interactive nature of these prompts enhances the creative process, allowing you to explore different visual ideas, request modifications, and perfect the final output. By fostering a dynamic exchange, you can ensure that the generated images or videos are more aligned with your vision, incorporating specific feedback and changes as needed.
-
Working extensively with images often involves an interactive approach where requesting the 'Seed number' or 'Gen ID' becomes quite useful. (If you're unfamiliar with seed numbers, there will be more information available in the updated version of the framework, which includes seed numbers for image creation.)
To streamline this process, you can incorporate a directive into your Custom Instructions for ChatGPT, or include it in every prompt, stating:
When creating images, also return the seed number and Gen ID.user >... When creating images, also return the seed number and Gen ID. ...The difference:
Gen ID (Generation ID): This is a unique identifier assigned to each image generated by DALL·E 3. It's used to reference and retrieve the specific image from the system. The Gen ID helps in tracking, sharing, and managing images within the platform, ensuring that every creation can be distinctly recognized and accessed.
Seed ID: The Seed ID, often referred to simply as the "seed," is related to the randomness used in the generation process of an image. In generative models like DALL·E 3, randomness plays a crucial role in creating varied and unique outputs. The seed determines the starting point of this randomness, and using the same seed with the same input prompt will theoretically produce the same image every time. It's a way to introduce controlled variability in the model's output and can be used to replicate results or ensure consistency across different generation attempts.
This not only aids in reproducing or referencing specific images but also enhances the interactive experience. However, it's important to remember that when you use DALL·E 3 to generate a new image based on your prompt, it still operates under certain parameters, such as temperature settings. This means that while you can guide the creative process by referencing previous images, the system will interpret your prompt within the bounds of these parameters to create something new. This balance between control and creativity is what makes using DALL·E 3 a unique and engaging experience.
Similarly, this concept of specifying return formats can be applied to other outputs. For instance, when ChatGPT generates tables, you could request that each table includes a unique ID as the first column. This practice ensures consistency and ease of reference across different types of outputs, making your interactions with the AI more efficient and tailored to your specific needs.
-
The true potential of ChatGPT is realized when you also use IPE to modify images after they are generated. If you've created an image that meets your expectations, you can enhance it further without any external software. Simply ask ChatGPT to apply filters using Python.
Enhance the image with Pythonuser >Use Python to add a Grayscale color filter to it. Display the original picture, and new picture side by side.If you're satisfied with the code, you can request a download link to obtain the new image.
Create downloaduser >Make it available for download.Examples of filters to apply are: Blur effects (Gaussian Blur, Box Blur, Median Blur), Sharpening, Edge Detection (Sobel filter, Canny Edge Detection), Color Filters (Grayscale, Sepia, Thresholding), Contrast Adjustment, Brightness Adjustment, Rotation and cropping, resizing and scaling, Adding text or overlays, Morphological Transformations (Erosion, Dilation).
N - Non-disclosure
When generating images or videos, it's essential to incorporate legal, security, and privacy considerations, especially when handling sensitive visual content. Protecting this information and ensuring compliance with non-disclosure agreements (NDAs) is critical. Specify any restrictions related to data security and privacy to maintain ethical practices in visual content creation. Before submitting your prompt, ensure that all sensitive elements are properly handled to prevent unwanted exposure or breaches, thus safeguarding confidentiality in the process.
-
While tools like DALL·E 3 offer the ability to generate images for commercial use, it is paramount to ensure that these images are free of copyrighted elements. Users must exercise due diligence to confirm that the generated content does not inadvertently include copyrighted material or trademarks that could lead to legal issues. This includes closely reviewing the images for any recognizable features or designs that are protected by copyright. To adhere to legal standards and maintain ethical usage, always verify that your content complies with all applicable copyright laws before usage, especially in professional or commercial settings. Ensuring careful scrutiny of visual content helps prevent copyright infringement and protects against potential legal complications.
Additional note on Acting Roles: When utilizing the 'Acting Role' component, as detailed in the section on assigning stylistic or artistic influences, be aware of the potential for including copyrighted elements. For example, emulating a 'Marvel style' might inadvertently incorporate signature superhero elements that are copyrighted. It is vital to review any content created under specific acting roles or artistic influences to ensure that no copyrighted materials are used, thus maintaining compliance with copyright laws and safeguarding against legal issues.
G - Goal-driven
Clearly defining your objectives in the prompts for image or video generation is key to obtaining precise and relevant outputs. Specify your desired visual outcomes and any particular elements that need emphasis, allowing the generation tools to align their output with your specifications. This goal-driven approach ensures the tools understand your intent and focus on producing visuals that precisely meet your needs. After submitting your prompt, it's important to review the generated content to ensure it meets your expectations. If necessary, adjust the formatting and technical parameters to refine the visuals and enhance their suitability for your specific objectives.
Enhancement phase
AI - Adapt and Improve
As AI tools evolve rapidly, it's important to adapt and improve your strategy and prompts. The "AI" aspect of the framework reminds you to refine your output by asking follow-up questions and adapting your approach based on the initial results. Validate the response and make adjustments as necessary. Additionally, continuously improving your prompts and embracing new possibilities is crucial for generating the best possible outcomes. Stay proactive and adaptable to ensure ongoing success with these powerful tools.
-
Text embedded in images often poses challenges for models tasked with image generation, leading to inaccuracies. To mitigate this, it's advisable to exclude any text from images in your initial requests. Alternatively, tools like Canva offer a useful "grab text" feature, allowing you to modify any text in images post-generation. This approach ensures clarity and correctness in the visual content you create.
It is not necessary to format your prompt following each step individually, but it is important to include all the elements of the framework within your prompt. A valid prompt that encompasses all parts of the framework could be:
So for example, to fill the above one in:
For visual output generation, the most advanced and best approach is using JSON.
{
"Context": {
"Scene": "[Describe the Scene]",
"Setting": "[Specify Setting]",
"CharactersOrSubjects": "[Describe Characters or Subjects]",
"Purpose": "[State the Purpose/Message of the Image]"
},
"Register": {
"ArtisticStyle": "[Specify Artistic Style]",
"ColorPalette": "[Specify Color Palette]",
"MoodOrTone": "[Describe Mood/Tone]"
},
"ActingRoleOrAesthetics": {
"Artist": "[Name of Artist]",
"ArtisticMovement": "[Name of Artistic Movement]",
"AestheticElements": "[Mention Key Aesthetic Elements if any]"
},
"Format": {
"Orientation": "[Specify Orientation]",
"AspectRatio": "[Specify Aspect Ratio]",
"ImageType": "[Photograph, Illustration, Digital Painting, etc.]",
"LightingConditions": "[Specify Lighting Conditions]",
"CameraSettings": "[Specify Camera Settings]"
},
"Task": {
"Objective": "Create an image that captures the essence of the described scene.",
"Instructions": "[Additional Instructions]",
"Parameters": "[Specify parameters]"
}
}
Please be aware that this is an iterative process, which means that after the enhancement phase, you will return to the crafting phase. You can adapt and improve the generated output by writing another prompt, utilizing the interactive (i) approach to refine it towards your desired goal (g), while considering the non-disclosure (n) constraints. In your follow-up prompt, you can build upon the previous output without reiterating all of the CRAFT elements. Instead, concentrate on the elements that will improve and contribute to the output's enhancement (AI).
Released: Apr 24, 2024
Now that you understand the basics of crafting effective prompts, it's time to take your skills to the next level. Let's dive into advanced prompt engineering techniques and explore how you can harness the full potential of AI.
Read more
The three phases
CRAFT
Craft (write) the prompt with the following elements: Context, Register, Acting Role, Format, and Task.
ING
Validate the prompt and ensure it maintains an interactive approach. Keep in mind the importance of non-disclosure and staying goal-driven throughout the process.
AI
Continuously assess and refine the output based on the prompts output to improve the overall quality.