2024 Workflow Thoughts

Embracing Simplicity in a Sea of Complexity
In the expansive world of AI image generation, it's easy to get lost in endless iterations and refinements. I prefer a different route—a single-pass workflow that emphasizes simplicity and exploration over meticulous tweaking. This approach allows me to focus on the learning how AI fundamentally works by poking at its different layers one at a time.
My journey begins with a concept or idea, often crafted through custom GPT tools like Diffusion Fashion Muse or Fantasy Art Prompt Generator. By starting with a well-defined prompt by a language model more knowledgeable and artistic than myself, I aim to minimize the need for subsequent edits or dull images. It's about setting a clear course from the outset, allowing the tools to do what they do best without constant interference.
The Tools That Shape the Journey
Why A1111 with SD 1.5?
I primarily use A1111 with Stable Diffusion 1.5 checkpoints for several reasons:
Speed and Efficiency: SD 1.5 models are lightweight and fast, allowing me to run them on my home setup with two video cards. This speed is crucial for a single-pass workflow, providing quick feedback that aids in refining prompts without long wait times.
Versatility: Despite being less advanced than newer models like SDXL, SD 1.5 offers a balance between performance and resource requirements. It handles a wide range of styles, from realism to watercolor, which suits my exploration of different artistic expressions.
Control: Using SD 1.5 allows me to experiment with various samplers, schedulers, LoRAs, and Textual Inversions. This flexibility is essential for fine-tuning the initial output without relying on multiple editing passes.
Custom Models and Merging Techniques
I often start with my merged artistic model, Recursed Canvas, and switch to Aberrated Perceptron when aiming for realism. Tacking back and forth between these models with A1111 can happen at multiple phases providing basically a slider for artistry versus realism. Both of these models are merged versions that I've fine-tuned over time. By adjusting merge percentages (usually between 5% and 20%) with merge block weighting, I can control which aspects of each model influence the final output of the new model. This technique enhances the initial generation, reducing the need for further adjustments or lackluster creativity.
Crafting the Vision: From Concept to Prompt
The prompt is the cornerstone of my single-pass workflow. I begin with a clear idea—say, an "evil sorceress in a haunted forest wearing couture clothing"—and use my GPT tools to generate a detailed prompt. I then refine this prompt, removing or adjusting elements that might confuse the model due to ambiguous terms or conflicting token weights.
For example, phrases like "bow-shaped lips" might inadvertently introduce literal bows into the image. By identifying these potential pitfalls early, I can adjust the prompt to maintain clarity and focus. This meticulous prompt crafting reduces the likelihood of undesirable elements appearing in the final image and allows the prompt to more likely work across multiple models.
Understanding Denoise Levels: Balancing Artistry and Realism
Denoise levels play a crucial role in defining the style of the generated image. Is my goal to just change skin texture or am is the goal to change the actual face shape? In my ADetailer workflow:
Lower Denoise Levels (0.2–0.3): These settings preserve more of the original structure during upscaling and refining steps. They're ideal for enhancing realism without significantly altering the underlying image.
Higher Denoise Levels (0.35–0.45): These allow for more creative freedom, introducing stronger prompt adherence. I use these when aiming for a more expressive, less literal interpretation of the prompt.
By carefully selecting denoise levels during processes like hires fix and adetailer, I can control the balance between realism and artistic style within a single pass.
Inpainting and Segmented Models: Precision Without Over-Editing
While I strive to minimize post-generation edits, tools like adetailer and segmented models play a vital role in refining specific areas:
Adetailer: This tool allows for targeted enhancements by focusing on particular regions, such as faces or eyes. It operates within the single-pass philosophy by integrating adjustments into the initial generation process.
Segmented Models (e.g., Mediapipe): Using models like mediapipe_face_mesh, I can achieve natural blending when enhancing facial features. These models detect facial landmarks, enabling precise adjustments without affecting the entire image. By using a segmented model, the space of the image being edited does not have the typical tell-tale rectangular blending that makes adetailer's edits less obvious.
For instance, when generating images with watercolor styles inspired by artists like Agnes Cecile, I might use adetailer to start with a face model with a lower denoise level and then transition to a higher denoise level on the eyes to introduce realism.
Challenges and Learning Moments: Navigating Uncharted Waters
Throughout this journey, I've encountered challenges that have shaped my understanding of AI image generation:
Model Limitations: Realizing that Stable Diffusion 1.5 struggles with certain concepts, like accurately rendering "white hair" without defaulting to a platinum blonde, highlighted the importance of model selection and prompt specificity.
Tokenization Nuances: Learning how phrases translate into visual elements taught me to communicate more precisely. Misinterpretations like "cowboy shot" leading to the addition of cowboy hats emphasized the need for careful wording.
Balancing Creativity and Ethics: While exploring styles inspired by specific artists, I grappled with the ethical implications of using their unique styles. This awareness guides me to appreciate and reference their work respectfully.
These experiences reinforce the value of continuous learning and adaptation, essential components of my single-pass workflow.
Evaluating the Outcome: Knowing When to Anchor
Assessment is swift but deliberate. I scan the image for glaring issues-distorted anatomy, off-scale features, or elements that detract from the overall composition. If an image doesn't pass the one sniff smell test, I discard it. This practice maintains quality without getting mired in perfectionism.
I also rely on my background in photography and digital imaging to inform this evaluation. Years of experience with composition, lighting, and aesthetic principles provide an intuitive sense of what works for me and what doesn't.
Sharing and Community: Contributing to Collective Knowledge
Transparency is important to me. I include all generation data when sharing images on platforms like CivitAI. By providing detailed information about prompts, models, steps, and denoise levels, I hope others can learn from my experiments—both successes and failures.
This open approach aligns with the principles of the Creative Commons and fosters a collaborative environment where knowledge is freely exchanged. Almost all of my personal photography was released with the Creative Commons license. I was proud when I reverse image searched some of my favorite shots to find them being used on Wikipedia.
Conclusion: Embracing Simplicity and Exploration
My single-pass workflow is a result of my goal to learn how AI works in creative exploration. By focusing on the foundational elements—the concept, the prompt, and the initial generation—I can navigate the vast possibilities of AI art without becoming overwhelmed by endless parameter adjustments across the workflow.
In the ever-evolving landscape of AI, there's a certain joy in charting your own course. By embracing a single-pass approach, I've found a way to balance learning efficiency with creativity, making the journey as rewarding as the destination.