BlogComfyUI Explained: How Node Workflows Turn Local AI Image Generation into a Pipeline You Control
Web Development10 min read

ComfyUI Explained: How Node Workflows Turn Local AI Image Generation into a Pipeline You Control

A practical guide to ComfyUI: why it uses nodes instead of fixed buttons, what hardware it really needs, how the basic text-to-image graph works, which KSampler settings matter, and how img2img, inpainting, LoRA and ControlNet plug in.

Bahaa Esmail
ISMS & DevOps Lead
Illustration of connected workflow nodes on a navy background with the title ComfyUI: Node workflows for local AI images

Open image models such as Stable Diffusion, Flux and Qwen Image are like an engine without a dashboard. You need software that turns your prompt, settings and reference images into something the model understands, then turns its output back into a picture you can save.

ComfyUI is the most flexible of these tools, and the most intimidating at first sight: instead of a form with a Generate button, you get a canvas of boxes and wires. This guide explains what those boxes do, what hardware you need, and how to build and adjust the standard workflows yourself.

Why nodes instead of fixed buttons?

Form-based interfaces such as Forge or Fooocus wrap the generation pipeline in fixed screens: you fill in fields, press a button, and the steps in between stay hidden. That is fast to learn, but changing the pipeline itself means waiting for the developers to add an option.

ComfyUI shows the pipeline as a graph. Each step (load a model, encode the prompt, sample, decode, save) is a separate node, and you connect them with links. You can see every intermediate result, swap one part without touching the rest, and save the whole graph as a reusable workflow. It is not the only tool with this idea: InvokeAI, for example, also has a node-based workflow editor alongside its canvas.

Scroll horizontally to see all columns

Form-based interfaces vs ComfyUI
AspectForm-based interfacesComfyUI
PipelineFixed by the developersBuilt from nodes you connect and rearrange
VisibilityIntermediate steps hiddenAny output can be previewed or reused
Re-runningUsually reruns the whole jobReruns only nodes whose inputs changed
MediaMostly still imagesImages, video, audio, 3D and more, depending on the model
Learning curveGentleSteeper at first

The re-running row matters: ComfyUI only re-executes the parts of the graph whose inputs changed, so tweaking the final step does not reload the model or redo everything before it.

What hardware do you need?

Three parts of your computer do the work. The graphics card's memory (VRAM) holds the model weights being used and the tensors the sampler works on. System RAM holds models that are loaded but not needed right now, and parts of a model when it does not fit in VRAM. The SSD stores the model files, which are often several gigabytes each.

Diagram: the graphics card memory holds the model weights in use and the sampler data, system RAM holds models on standby and parts that do not fit, and the SSD stores model files; the project says large models can run on 4 GB VRAM plus 8 GB RAM by streaming weights.
How ComfyUI spreads the work across VRAM, system RAM and the SSD.

The project says ComfyUI can run even large models on as little as 4 GB of VRAM and 8 GB of RAM by streaming weights between RAM and the GPU. That makes low-end cards usable, not fast: more VRAM means fewer transfers and quicker results, especially with large image models and video. The official SD1.5 tutorial, for example, describes that roughly 4 GB model as running smoothly on a consumer card with about 6 GB of VRAM.

NVIDIA cards have the broadest support. According to the system requirements, ComfyUI also runs on AMD cards through ROCm (Linux and Windows), Intel Arc, Apple Silicon Macs through PyTorch's MPS backend, and even the CPU, which is much slower. Some community custom nodes are written with NVIDIA in mind, so check before relying on them with other hardware.

There are three ways to install it. The desktop app is the easiest. The Windows portable build bundles its own Python in one folder, which avoids conflicts with other software; there are builds for NVIDIA, AMD and Intel cards. A manual install with Git works on every system and gives you the latest changes first.

Anatomy of a node

Every node follows the same layout. The title bar shows its name and lets you collapse it; right-clicking the node (or using the selection toolbox) changes its mode, for example to bypass. Inside are widgets: number fields, text boxes and drop-down menus for its settings. Inputs sit on the left edge and outputs on the right.

Simplified drawing of a KSampler node: a title bar at the top, input ports model, positive, negative and latent_image on the left, a LATENT output on the right, and widgets for seed, steps, cfg, sampler, scheduler and denoise.
The parts of a node, using KSampler as the example.

Each link carries one data type, and ComfyUI colors ports and links by type. You can only connect ports of the same type, which prevents many mistakes. The colors below are the defaults listed in the official documentation; themes can change them.

Scroll horizontally to see all columns

Common data types in ComfyUI
TypeWhat it carriesDefault color
MODELThe diffusion model (UNet or DiT) that removes noiseLavender
CLIPThe text encoder that turns prompts into vectorsYellow
VAEConverts between pixels and latent spaceRose
CONDITIONINGEncoded prompts and other guidanceOrange
LATENTThe compressed image the sampler works onPink
IMAGENormal RGB pixelsBlue
MASKWhich area of an image to changeGreen

The basic text-to-image workflow

The default workflow has six kinds of nodes, and almost every other workflow is a variation of it.

Graph of the default text-to-image workflow: Load Checkpoint sends CLIP to two CLIP Text Encode nodes, MODEL to KSampler and VAE to VAE Decode; Empty Latent Image feeds KSampler; KSampler feeds VAE Decode, which feeds Save Image.
The default text-to-image workflow and what flows along each link.
  • Load Checkpoint reads a model file and outputs three things: MODEL, CLIP and VAE.
  • Two CLIP Text Encode nodes use CLIP to encode the positive prompt (what you want) and the negative prompt (what you don't). Their CONDITIONING outputs go to the sampler's positive and negative inputs.
  • Empty Latent Image creates the blank canvas in latent space. You set width, height and batch size here.
  • KSampler takes the model, both conditionings and the latent, and removes noise step by step.
  • VAE Decode turns the sampler's latent result into pixels using the VAE.
  • Save Image (or Preview Image) shows the result and saves it to the output folder.

Why a "latent" canvas? Diffusion models do not work on pixels directly. With the VAEs used by Stable Diffusion and Flux, the latent is 8 times smaller than the image on each side, so a 1024 × 1024 image becomes a 128 × 128 latent. That is why generation is feasible on consumer cards, and why the width and height in ComfyUI move in steps of 8.

KSampler settings that matter

KSampler is where the image is actually made, and its settings decide quality, speed and how closely the result follows your prompt.

Cards summarizing KSampler settings: seed, steps (default 20), cfg (default 8, Flux expects 1), sampler_name, scheduler and denoise (1.0 for text-to-image).
The six KSampler settings and their defaults.

Seed sets the starting noise. The same seed with the same model, settings and software gives the same image, and the control after generate option decides whether the seed stays fixed, changes randomly, or goes up or down by one after each run. ComfyUI creates this noise on the CPU, which helps results match across machines, but exact repeats are not guaranteed: a different GPU, driver or ComfyUI version can change the output slightly. Even the --deterministic launch option is documented as not making images deterministic in all cases.

Steps is the number of denoising steps; the default is 20. More steps take longer and usually improve detail only up to a point. Some models are built for very few steps: the official Flux.1 Schnell workflow uses 4.

CFG controls how strongly the image follows the prompt. The default is 8; too high a value hurts quality, often as harsh colors and artifacts, while too low lets the model drift from the prompt. The right range depends on the model: distilled models such as Flux expect a CFG of 1, and ComfyUI's Flux examples say so explicitly.

Sampler is the algorithm that removes the noise. Euler is simple and fast, DPM++ 2M is a common choice that converges well in a moderate number of steps, and the SDE and "ancestral" variants add fresh noise during sampling, which changes the look and makes results vary more.

Scheduler decides how the noise levels are spread across the steps. Karras, for example, spends more of its steps at the low-noise end, which many people find gives finer detail.

Denoise sets how much of the starting latent is replaced. Keep it at 1.0 for text-to-image. Lower values are for image-to-image, covered below.

From latent to PNG, with the workflow inside

VAE Decode runs the VAE's decoder to turn the latent into normal pixels, and Save Image writes the result to ComfyUI's output folder.

ComfyUI also stores the whole workflow, including the seed, inside the PNG's metadata. Drag that PNG back onto the canvas and the exact graph that made it reappears, which makes PNGs a convenient way to share workflows. Two caveats: the --disable-metadata launch option turns this off, and websites or apps that re-save or compress images may strip the metadata.

Image-to-image and inpainting

For image-to-image, replace Empty Latent Image with two nodes: Load Image, then VAE Encode, which turns your picture into a latent. Feed that latent to KSampler and lower denoise.

Image-to-image chain Load Image, VAE Encode, KSampler with denoise below 1, VAE Decode; a denoise scale marking 0.2 to 0.4 as keeping structure and 0.5 to 0.7 as a new style with the same layout; and two inpainting options, VAE Encode (for Inpainting) and Set Latent Noise Mask.
Image-to-image, rough denoise starting points, and two ways to inpaint.

Denoise is now the key setting: the lower it is, the closer the result stays to the original. As rough starting points, low values around 0.2 to 0.4 keep the structure and change details, while 0.5 to 0.7 change the style more while usually keeping the composition. The right value depends on the model and the image. Note that a lower denoise does not mean fewer steps: ComfyUI still runs the steps you set, over the low-noise part of a longer schedule, as the image-to-image guide explains.

Inpainting changes only part of an image. You draw a mask with the built-in mask editor (or use an image's transparency), and the mask tells the sampler where it may work. Core ComfyUI offers several ways to do this (the inpainting guide walks through the first):

  • VAE Encode (for Inpainting) fills the masked area with neutral gray before encoding, so it works best with denoise at 1.0, ideally with a model trained for inpainting. Its grow_mask_by setting widens the mask to soften the seam.
  • Set Latent Noise Mask keeps the original content under the mask, so it suits smaller changes with lower denoise and ordinary models.
  • InpaintModelConditioning prepares the image, mask and prompts together for inpainting models.

LoRA and ControlNet

A LoRA is a small add-on file that adjusts a model toward a style, character or subject without replacing it. LoRAs are usually much smaller than the checkpoint they modify; their size depends on the base model and how they were trained. The Load LoRA node sits between Load Checkpoint and the rest of the graph on both the MODEL and CLIP paths. Its strength_model and strength_clip settings control how strongly it affects the image model and the text encoder, and several Load LoRA nodes can be chained.

Graph showing Load LoRA between Load Checkpoint and the rest of the workflow on the MODEL and CLIP paths, and Apply ControlNet taking the encoded prompts, a ControlNet model and a preprocessed reference map before KSampler.
LoRA modifies the model and text encoder; ControlNet adds guidance to the prompts.

ControlNet guides the composition with a reference map instead of words: edges, depth, or a human pose. A preprocessor turns your reference image into that map. Core ComfyUI includes a Canny edge node, while most other preprocessors (depth, line art, OpenPose and more) come from custom node packs such as comfyui_controlnet_aux.

Load ControlNet loads the model, and Apply ControlNet adds it to both the positive and negative conditioning before they reach KSampler. Strength sets how strongly it constrains the image, and start_percent and end_percent set during which part of the sampling it applies. The ControlNet guide shows a complete example.

Keeping big workflows tidy

Workflows grow quickly. A few habits keep them readable:

  • Groups (Ctrl+G) put a colored frame around related nodes, such as "prompts" or "upscaling".
  • Reroute points move wires out of the way; the canvas has a native reroute feature for new workflows.
  • Bypass (Ctrl+B) skips a node and passes its input straight through, so you can switch off a LoRA without rewiring. Mute (Ctrl+M) stops a node completely, which breaks anything that depends on it.
  • Subgraphs collapse a set of nodes into one reusable node.
  • Primitive nodes let one value, such as width or height, feed several nodes at once.

Conclusion

ComfyUI turns image generation from pressing a button into designing a pipeline. The learning curve is real, but the concepts are few: data types and links, latent versus pixel space, the sampler's settings, and the add-ons that plug into the conditioning and model paths. Once those make sense, you can read any shared workflow, fix it when it breaks, and fit it to your own hardware.

The fastest way to start is the default text-to-image workflow: change one setting at a time, watch what it does, then add image-to-image, a LoRA or a ControlNet one step at a time.

Want this for your product?

Send a short note about your project. We will review it and explain the next useful step.

Contact Our Team