A couple of weeks ago, out of interest and, perhaps, a hidden desire to generate a ballerina from Atomic Heart, I installed a neural network and began experimenting, first with text queries, then with photo editing, masks, sketches, etc.
I haven’t gone too far, but perhaps the initial tools will be of interest to those who stumble upon this, God forgive me, “material”.
Regarding installation: as they write on the site where I found the instructions, you need a “sufficiently powerful computer”, at least 4GB of video memory and 10GB for installation. I have an Nvidia 3050, 8GB – it may refuse to generate images larger than 1000×1000 pixels, I have to play with filters and change sizes (I’m interested in generating based on photos, so either crop the photo or install filters so that the output image is cropped and formatted). The pictures themselves, depending on the parameters, are generated from 1 to 10 minutes (my standard experiments, I did not try to arrange stress tests).
Instructions here https://remontka.pro/stable-diffusion-install-use/#install
Separately, I want to clarify that the instructions are not complete at all, in my case the component for using nvidia cuda is not installed and does not really want to be installed, this is a very sluggish battle, but maybe someday I will win it, as long as I am guided by the principle “it seems to work”.
If links are not possible (I’m writing here for the first time), I installed the AUTOMATIC1111 version with a web interface, because why not, it’s so convenient for me, don’t judge me. Let me explain, there are versions that work via the API directly via the command line. In my setup both options are possible. You can even create a public link to generate, for example, images while sitting on your phone, using the power of your PC.
After ten days of use, I settled on the Protogen x3 model.4, under 6GB (so-called checkpoints, training models on which convenient scenarios have been worked out). In addition to models, there are additional processors: Hypernetwork, LoRA, Textural Inversion. They weigh less, are connected to the selected model as additional filters in the form of a description. In the case of the web interface – selected – poked – appeared in the description. In the case of API – I don’t know, stop putting pressure on me.
txt2img – the tab https://independentcasinosites.co.uk/ where you first enter
Promt – what you want to see in the resulting picture, from simple phrases to paragraphs found on the Internet, from which you can guess what the image will be like.
Negative prompt – something that should not be in the image
Sampling method – a list of sampling methods, there is a list of options, you can switch and look at the generation result, or you can find the picture you are interested in, see what description and what sampling method was used and select it. I did both, in any case it’s interesting, subjectively I like DPM++ SDE Karras more
Width, Height, that’s clear.
CFG Scale (in fact, there are explanations when you hover over the properties) – how much the output image will correspond to the description, how much creativity the neural network is allowed.
Denoising strength is also actually a scale of creativity, it no longer depends on the description, the higher it is, the more abstract your image is.
Batch count – how many images to generate. Increase time.
Batch size – how many images to generate in 1 run. Will increase the amount of memory consumed
Sampling Steps (UPD) – the greater the number of steps, the more artifacts are removed, in most cases, as it is written in the manuals, after 30 steps little changes, for my part I set it from 50 to 100, in the case of masks it seems that the final drawing is cleaner, self-hypnosis is possible.
img2img – with the same text description you can add a photo and configure it so that the resulting image is at least somewhat similar to it
And Mask is the Inpaint tab. Here on the image you can indicate areas that you want the processor not to interpret or, on the contrary, to interpret only them (in my brain everyone understands what I mean “either transform or ignore”).
There is also a Scetch tab – draw/label a sketch in different colors and throw it into the generator.
A lot depends on the set of samples and how you adjust the sliders, but you can create several pictures in the selected option that you like and there will even be good differences, there will be plenty to choose from.
I both downloaded models for generation and simply searched for popular products. I didn’t get to something more adult, but I decided to create something that is neither an installation instruction, nor a user manual, nor anything useful at all; the denouement is revealed at the beginning of the text, as is customary with Dontsova, King and [insert the name of a popular writer, with a banal plot development]