comfy-image-utils

v2026.09.24

ComfyUI non-inference image ops: resize/crop/pad/tile/batch/mask utilities, plus image-to-text captioners (Florence-2, WD14, BLIP, DeepDanbooru). Use when manipulating images or generating captions/tags in a workflow.

GitHub
安装命令
npx skhub add laurigates/comfy-image-utils
Markdown
SKILL.md

ComfyUI image utilities

Image manipulation that doesn't go through a diffusion model. Plus image-to-text inference (Florence-2, WD14, BLIP, DeepDanbooru), which is included here because the input is an image and the node-level setup (model downloads, ONNX dependencies, HF cache paths) is the bulk of the work.

The split:

PackNiche
comfyui-kjnodesBatch ops, resize-v2, crop-by-mask, channel split/merge, Get*SizeAndCount, LoadAndResizeImage (exposes image_path)
comfyui_essentialsResize / Flip / Crop / Tile-Untile / Composite, list↔batch conversion, Mask family (Blur, Flip, FromColor, BoundingBox)
comfyui-easy-useimageCount, imageInsetCrop, imagesCountInDirectory
comfyui-tooling-nodesBase64 load, image cache, ApplyMaskToImage, WebSocket send, Tile Extract/Merge
ComfyUI-CrystoolsCImageGetResolution, CImageLoadWithMetadata, CImageSaveWithExtraMetadata
bjornulf_custom_nodesResizeImage, ResizeImagePercentage, GrayscaleTransform, RemoveTransparency, LoadImageWithTransparency
comfyui_yvann-nodesRepeatImageToCount
comfyui-custom-scripts (pysssss)ConstrainImage (max-dimensions resize with aspect-preserve)
comfyui-variousimage_ops / channel_ops / color_ops / image_sequence / mask_sequence_ops modules
comfyui-florence2Florence-2 vision-language for captioning (PromptGen LoRAs)
comfyui-wd14-taggerWD14 ONNX booru-style tagger
comfyui-art-ventureBLIP captioner, DeepDanbooru anime tagger

When to Use This Skill

Use this skill when...Use instead when...
Manipulating images/masks outside of model inference (resize, crop, tile, batch)Running a diffusion/inference node on an image -> the relevant model-family skill
Generating a caption/tag from an image (Florence-2, WD14, BLIP)Extracting metadata already embedded in an output -> comfy-metadata

Sources of truth

  • custom_nodes/comfyui-kjnodes/nodes/image_nodes.py — batch / resize / channel / size+count
  • custom_nodes/comfyui_essentials/image.py and mask.py — Image*/Mask* family
  • custom_nodes/comfyui-tooling-nodes/ — base64, cache, websocket, tiling
  • custom_nodes/comfyui-florence2/ — Florence-2 loaders + caption nodes
  • custom_nodes/comfyui-wd14-tagger/ — WD14 ONNX wrapper
  • custom_nodes/comfyui-art-venture/modules/interrogate/ — BLIP, DeepDanbooru

Resize decision

You want…Best nodeWhy
Resize to specific width × heightessentials ImageResizeCanonical, scale-mode picker
Resize keeping aspect ratio, target one dimensionkjnodes ImageResizeKJv2keep_proportion flag plus more interpolation options
Resize to a percentage of originalbjornulf ResizeImagePercentageSingle-input scale factor
Constrain to max dimensions (downsize only)pysssss ConstrainImageOne-sided cap, preserves aspect, no upscale
Load + resize in one step + expose source filenamekjnodes LoadAndResizeImageBonus image_path STRING output for filename-templated saves
Crop region by coordinatesessentials ImageCropxywh widgets
Crop to bounds of a maskkjnodes ImageCropByMaskFits to non-zero region of input mask
Inset-crop (chop edges)easy imageInsetCropCrop from edges by pixels or %
Pad on all sideskjnodes ImagePadKJTop/bottom/left/right + fill color

Latent-friendly sizing

Many models require dimensions divisible by 8 (or 16 for some WanVideo configs). ImageResizeKJv2 has a divisible_by widget; essentials ImageResize does not. For latent-aligned resize, prefer kjnodes' v2; otherwise pre-compute the target size via SimpleMath (see comfy-math-strings).

Batch operations

NeedBest node
Concat two batcheskjnodes ImageConcatenate
Concat N batcheskjnodes ImageConcatMulti
Batch N images into one tensor (from separate IMAGE outputs)kjnodes ImageBatchMulti
Extract specific indiceskjnodes GetImagesFromBatchIndexed
Extract a contiguous rangekjnodes GetImageRangeFromBatch
Reverse the batch orderkjnodes ReverseImageBatch
Shuffle (random)kjnodes ShuffleImageBatch
Pick one image by indexessentials ImageFromBatch
Duplicate one image N timesessentials ImageExpandBatch
Repeat batch K timesessentials ImageBatchMultiple
Match a target batch size by repeatingyvann RepeatImageToCount
Batch tensor → list of imagesessentials ImageBatchToList
List of images → batch tensoressentials ImageListToBatch
Count batch sizeessentials BatchCount, easy imageCount

The batch ↔ list distinction matters: LIST is a Python list of single-image tensors processed by INPUT_IS_LIST=True nodes; BATCH is a single 4D tensor (B, H, W, C). Many downstream nodes (samplers, VAE) expect BATCH. Convert with ImageListToBatch before the sampler.

Mask utilities

NeedBest node
Gaussian blur a maskessentials MaskBlur
Horizontal/vertical flipessentials MaskFlip
Combine N masks into a batchessentials MaskBatch
Make a mask from a color range in an imageessentials MaskFromColor
Get the (x, y, w, h) bounding box of maskessentials MaskBoundingBox
Get mask dimensionskjnodes GetMaskSizeAndCount
Apply a mask to an image (alpha composite)tooling-nodes ApplyMaskToImage
Mask → grayscale image(use ApplyMaskToImage on a white image)

For mask shape detection (face mask covers a region; is it empty?), the easy isMaskEmpty probe lives in comfy-conditionals.

Image I/O

NeedNode
Load image from base64 string (API input)tooling-nodes LoadImageBase64
Load mask from base64tooling-nodes LoadMaskBase64
Cache image in memory (reuse across queue runs)tooling-nodes Save Image Cache / Load Image Cache
Send image over WebSockettooling-nodes Send Image WebSocket (for streaming to external tools)
Load image preserving alpha channelbjornulf LoadImageWithTransparency
Convert RGBA → RGB (replace alpha with color)bjornulf RemoveTransparency
Convert image to grayscalebjornulf GrayscaleTransform
Save image with custom PNG metadataCrystools CImageSaveWithExtraMetadata
Load image + emit PNG metadata as JSONCrystools CImageLoadWithMetadata
Get image dimensions (no batch)Crystools CImageGetResolution

For batch / cross-output-directory metadata analysis (figure out which model produced a directory of outputs), use the comfy-metadata skill — it covers PNG tEXt / iTXt, WebP EXIF, MP4 container metadata (kijai WanVideoWrapper's comment blob), and .latent safetensors.

Tiling

NeedNode
Repeat image as a tile patternessentials ImageTile
Undo ImageTileessentials ImageUntile
Extract a specific tile from a grid layouttooling-nodes ExtractImageTile
Reconstruct from extracted tilestooling-nodes MergeImageTile
Extract tile-shaped masktooling-nodes ExtractMaskTile

For tiled diffusion (split a large image into tiles, run each through a sampler, stitch), that's a model-level pattern — see comfyui-tiled-diffusion (separate pack) or the comfyui-inpaint-cropandstitch skill referenced in portrait-outpaint.

Channel ops

NeedNode
Split RGBA into 4 maskskjnodes SplitImageChannels
Merge 4 masks (R, G, B, A) → RGBA imagekjnodes MergeImageChannels
Grayscalebjornulf GrayscaleTransform

The kjnodes split/merge pair is the workhorse for any per-channel image processing.

Image → prompt (image-to-text inference)

These nodes are inference nodes in the sense that they run a neural network, but they're not diffusion — they produce text from an image. They live in this skill because the input is an image and because users coming from comfy-prompting need the node-level setup details.

Decision: which captioner?

SourceOutput styleBest forSetup
Florence-2 PromptGenSD/Flux-friendly natural-language ("portrait of a woman in red, sitting, soft light")Photoreal / general-purposeAuto-downloads ~1.5 GB on first run; PEFT LoRA adapters for prompt style
WD14TaggerBooru-style comma-separated tags ("1girl, solo, red_dress, sitting")Anime / illustration promptsONNX model auto-downloads; needs onnxruntime
BLIP (art-venture)Short generic caption ("a woman in a red dress")Quick descriptions; older / smaller modelAuto-downloads from HF
DeepDanbooru (art-venture)Anime tag classifierSpecific to anime / booru tagsAuto-downloads model

Florence-2 setup

Nodes:

  • DownloadAndLoadFlorence2Model — auto-fetch from HF on first run
  • Florence2ModelLoader — load from a local path (after the first run, point at the cached path)
  • DownloadAndLoadFlorence2Lora — apply a PEFT LoRA adapter (PromptGen variants are LoRAs on top of base Florence-2)

The first run downloads to the HF cache. On a small-root-disk install, set HF_HOME to a larger data disk to keep models off the small root — see ~/.claude/rules/huggingface-downloads.md. Florence-2 weights for the base model are ~1.5 GB; PromptGen LoRA adapters are ~50 MB.

Florence-2 variants and their use:

VariantUse
microsoft/Florence-2-baseBase — generic captioning
microsoft/Florence-2-largeLarger, better quality
MiaoshouAI/Florence-2-large-PromptGen-v2.0Optimized for SD/Flux prompts
gokaygokay/Florence-2-Flux-CaptionerFlux-prompt-style captions
microsoft/Florence-2-large-ftFine-tuned on DocVQA / OCR — useful for screenshots

WD14 setup

Single node: WD14Tagger. ONNX-based, requires onnxruntime (CPU) or onnxruntime-gpu for CUDA acceleration. The ONNX model auto-downloads to the WD14 pack's model directory on first run.

Configurables:

  • model: WD14 model variant (Convnext, ViT, SwinV2, MoAT) — newer variants are slightly more accurate
  • threshold (general tags): 0.35 default; raise to filter weak tags
  • threshold_character: tags for specific known characters
  • exclude_tags: comma-separated tags to skip ("1girl, solo")
  • replace_underscore: convert booru underscores to spaces
  • trailing_comma: append , after the output

BLIP and DeepDanbooru (art-venture)

BlipLoader (or DownloadAndLoadBlip for one-step) + BlipCaption. First-run download ~500 MB. Output is a short caption.

DeepDanbooruCaption for anime — single node, auto-downloads model weights.

Caption chaining

For best-quality prompts, chain:

LoadImage ──► Florence-2 PromptGen ──► STRING (caption)
                                          │
                                          ▼ (optional)
                                  Searge_LLM_Node
                                  (add cinematic detail, camera direction)
                                          │
                                          ▼
                                  CLIPTextEncode (or model-specific encoder)

The Florence-2 → LLM chain produces more nuanced prompts than either alone. See comfy-prompting for the LLM-side details.

Recipes

Sortable per-day output filename with source basename

User wants outputs at output/2026-05-13/143055_michael.png where michael is stripped from the source filename michael.jpg. Project CLAUDE.md covers the native %date:%/%NodeName.widget% syntax; when the native approach fails (e.g. extension stripping), use:

LoadAndResizeImage (kjnodes) ─► image_path (STRING, full path)
                                       │
                                       ▼
                          StringFunction (pysssss, regex)
                          find:    "^.*/|\.(png|jpg|jpeg|webp)$"
                          replace: ""
                                       │
                                       ▼ (bare basename, no path, no ext)
                          JoinStringMulti (kjnodes)
                          in_1: "<bucket>/%date:yyyy-MM-dd%/%date:hhmmss%_%ksampler.sampler_name%_%ksampler.scheduler%_s%ksampler.seed%_"
                          in_2: <bare basename>   # the <descriptor> segment
                                       │
                                       ▼
                          easy imageSave (filename_prefix STRING input)

%date:% tokens pass through verbatim into the SaveImage prefix substitution (resolved at save time). String manipulation chain lives in comfy-math-strings; the save node lives in easy-use. The prefix convention itself is per-install — see "Naming conventions are per-install" in comfy-workflow-json.

Caption a folder of images for batch dataset prep

You have 100 photos in input/dataset/ and want a .txt next to each with a Florence-2 caption:

easy imagesCountInDirectory ──► count
                                  │
                                  ▼ (drives forLoopStart iteration count)
                          easy forLoopStart (total = count)
                                  │
                                  ▼ (per-iteration)
                  LoadImage (path = `input/dataset/{index}.jpg`)
                                  │
                                  ▼
                          Florence-2 PromptGen ──► caption STRING
                                                       │
                                                       ▼
                                                  SaveText (bjornulf)
                                                  path: `input/dataset/{index}.txt`
                                  │
                                  ▼
                          easy forLoopEnd

Pattern hits two skills: this one for Florence-2; comfy-flow-control for the for-loop primitives.

Mask-driven crop + paste workflow

You have a portrait, a face mask, want to upscale only the face:

LoadImage ──► IMAGE
   │
   ▼
GenerateFaceMask (whatever)
   │
   ▼ MASK
   │
   ▼
ImageCropByMask (kjnodes) ──► cropped IMAGE (face region only)
   │                       └─► crop coords (for paste-back)
   ▼
(upscale chain: sampler, etc.)
   │
   ▼ upscaled face
   │
   ▼
ImageComposite (essentials, alpha-paste using mask) ──► final image

ImageCropByMask returns both the cropped image and the bbox; the bbox is needed to paste the result back to the right location.

Tile-based large-image processing

LoadImage (4096×4096) ──► IMAGE
                            │
                            ▼
                  ExtractImageTile (tooling-nodes, 2×2 grid)
                            │
                            ▼ 4 IMAGE tiles
                            │
                            ▼ (process each through a sampler)
                            │
                            ▼ 4 processed tiles
                            │
                            ▼
                  MergeImageTile (tooling-nodes) ──► final 4096×4096

For sampler-aware tiled diffusion (overlap, blending across tile seams), reach for comfyui-tiled-diffusion instead — these tooling-node tile ops are pure image-level.

Gotchas

  • LoadAndResizeImage.image_path is a full path, not just a basename. Strip the dir component via StringFunction regex (see Recipe 1) or use the native %LoadImage.image% substitution if not using LoadAndResizeImage.
  • ImageResize defaults are not latent-aligned. The result may be 519×731, which crashes the VAE on some models. Use ImageResizeKJv2 with divisible_by=8 or pre-compute via SimpleMath.
  • ImageBatchMulti slots are dynamic. Adding/removing inputs rewrites the slot count. After heavy editing, re-add the node to compact unused slots.
  • MaskFromColor is RGB-only. RGBA images need RemoveTransparency first (or split channels and feed RGB).
  • WD14Tagger needs onnxruntime, not torch. The pack ships its own ONNX session; CUDA acceleration requires onnxruntime-gpu (which on this install would need: .venv/bin/python -m pip install onnxruntime-gpu).
  • Florence-2 first run downloads to ~/.cache/huggingface by default. On a small-root-disk install, set HF_HOME to a larger data disk (a systemd unit's Environment= block, or your shell profile). Otherwise the ~1.5 GB weights fill the small root partition. See ~/.claude/rules/huggingface-downloads.md.
  • BLIP captions are short (~10-15 words). For longer prompts, prefer Florence-2 PromptGen or chain BLIP output through a local LLM.
  • DeepDanbooru and WD14Tagger produce overlapping but non-identical tag vocabularies. Stick with one for a given workflow; mixing produces redundant 1girl, 1_girl, solo, person pile-ups.
  • Tooling-nodes' Load Image Cache is in-memory only. The cache doesn't survive a service restart. For persistent caching, save to disk via SaveImage with a known filename and reload.
  • ImageComposite (essentials) requires both images to be the same size. Resize one or use ImagePadKJ to match dimensions first.
  • SplitImageChannels always emits 4 MASKs (R, G, B, A). On an RGB input, the alpha channel comes back as a fully-white mask (not None) — wire only the channels you need.
  • RepeatImageToCount doesn't broadcast smaller dimensions. If the input batch has 1 image and target count is 5, you get 5 copies of that image — useful when matching a batch dimension to drive a per-frame conditioning input.

Cross-refs

  • comfy-prompting — when to reach for image-to-prompt nodes; wildcard / LLM combination with caption output.
  • comfy-conditionals — easy isMaskEmpty to gate mask-driven branches; MaskBoundingBox to compute "is the region large enough?" predicates.
  • comfy-flow-control — index switches over batches; for-loop iteration over image directories.
  • comfy-math-strings — string manipulation for filename templating (the recipe above); resolution math.
  • comfy-debug-preview — Get*SizeAndCount, CImageGetResolution, ImageDetails for inspecting batches and source images.
  • comfy-metadata — offline / cross-file metadata extraction across output directories; the Crystools nodes here are for inline reads, comfy-metadata is for retrospective analysis.
  • photo-restore, portrait-outpaint, video-extend — task-level skills that combine these image utilities with model inference.

Things this skill does NOT cover

  • Model inference on images — KSampler, VAEEncode/Decode, ControlNet apply, IPAdapter. Those are model-family work; see wan / z-image / hidream-o1 / task skills.
  • Tiled diffusion at the sampler level — see comfyui-tiled-diffusion and the inpaint-cropandstitch workflow referenced in portrait-outpaint.
  • Image-format conversion at the file level (JPEG quality, WebP encoding parameters) — that's tools-plugin:imagemagick-conversion territory for external CLI work.
  • 3D / depth / mesh manipulation — depthanythingv2, depthflow-nodes (broken on this install per project CLAUDE.md).
发现
标签

此技能尚未发布标签。

版本
最新版本元数据

版本

v2026.09.24

发布时间

Sep 24, 2026

分类

未分类

许可证

MIT

源路径

comfyui-plugin/skills/comfy-image-utils

默认分支

main

最新提交

1668324

Tree SHA

b2d4cc3