Catalog
Built-in presets
Ready-made forms and pipelines for each kind of creation.
These presets ship in PotionUI’s catalog. Install and assign them through Administration → Presets, or use a matching setup recipe. The tables list the modes each preset exposes; model files download separately, and some modes require different checkpoints within the same family.
Native presets run on the native engine, locally or on a compatible remote worker. Open a preset’s Requirements tab to check missing components. Each linked source includes its configuration and description.
Image generation
| Preset | What it does | Modes |
|---|---|---|
| SDXL | Generate images with SDXL checkpoints or redraw a masked region of an existing image. | Text to image; inpaint |
| Flux1 | Create images from text or transform an input image using a Flux 1 model. | Text to image; image to image |
| Flux2 | Generate and transform images with the Flux2 family, including Klein variants. | Text to image; image to image |
| Krea-2 | Generate images with Krea 2 or refine an already-upscaled image in its experimental Enhance mode. | Text to image; enhance |
| Anima | Create anime and illustration-style characters and scenes from text. | Text to image |
| Z-Image | Generate images with Z-Image. Defaults target Turbo checkpoints; base models use their own sampling settings. | Text to image |
| Qwen-Image | Create images, including compositions with written text, transform inputs, or edit with a compatible editing checkpoint. | Text to image; image to image; edit |
Video generation
Use Video Director to compose supported shots, keyframes, and audio. LTX and MiniMax-H3 integrations are experimental; check the selected preset’s controls and model guide.
| Preset | What it does | Modes |
|---|---|---|
| Wan 2.1 / 2.2 | Generate video from text or a starting image, with shot continuation. Choose matching text-to-video or image-to-video checkpoints. | Video |
| LTX-2 / 2.3 | Direct video with optional synchronized audio and reference conditioning, or upscale a finished clip. Uses the earlier all-in-one LTX checkpoints. | Video; upscale |
| LTX-2.5 | Direct and upscale video using LTX-2.5’s separate model components and Gemma4 text encoder. Audio generation also needs its audio VAE. | Video; upscale |
| MiniMax-H3 | Generate video and its soundtrack together. References mode accepts up to nine images and requires the separate reference-conditioned checkpoint. | Video; references |
| MiniMax-H3 VDN | Experimental Video DeltaNet variant intended for longer clips. Requires extra branch weights and adapters; performance and output quality remain unverified in PotionUI. | Video |
Music generation
| Preset | What it does | Modes |
|---|---|---|
| MiniMax-Music3 | Generate a song from lyrics and a style description, or enable Instrumental for music without vocals. | Song |
| YuE2 | Generate music from style tags and sectioned lyrics. Leave lyrics empty for instrumental output; optional melody or chord planning guides composition. | Song |
3D generation
| Preset | What it does | Modes |
|---|---|---|
| TRELLIS.2 | Turn a reference image into a textured 3D mesh with materials. | Image to mesh |
Media utilities
These tools process existing media. Their forms collect the input and operation settings without requiring a written prompt.
| Preset | What it does | Modes |
|---|---|---|
| SeedVR2 | Restore and upscale images or video using SeedVR2 model files and fixed conditioning. | Image upscale; video upscale |
| Image Tools | Remove or key out backgrounds, crop to a subject, resize and pad a canvas, or prepare a subject for animation. Matting operations need a compatible model. | Prepare; remove background; fit canvas; color key; crop subject |
| Video Tools | Interpolate a clip to 2× or 4× its frame rate using a RIFE checkpoint, preserving duration and audio. | Interpolate |
| Audio Tools | Trim an audio clip by start time and duration. Output is WAV at the source sample rate and channel count; no model is required. | Trim |
Presets supplied by plugins
Enable the providing plugin before installing these presets:
- ComfyUI Backend
- FLUX Klein 9B and QwenImage offer text-to-image and image-to-image modes. Krea 2, SDXL Illustrious, and zImage offer text-to-image. These run on a configured ComfyUI server with its required models and nodes.
- NVIDIA RTX Upscale
- RTX Upscale offers image-upscale and video-upscale modes, including denoise and deblur options. It requires the NVIDIA runtime described in the plugin entry.
For form authors, [Custom] Frontend Field Catalog is a bundled developer preset for inspecting field rendering. Its catalog mode does not generate media.