r/sdforall

Image 1 — If AI Makes Us More Creative, Why Does Everything Look the Same?
(A Painter’s Perspective)
Image 2 — If AI Makes Us More Creative, Why Does Everything Look the Same?
(A Painter’s Perspective)
Image 3 — If AI Makes Us More Creative, Why Does Everything Look the Same?
(A Painter’s Perspective)
Image 4 — If AI Makes Us More Creative, Why Does Everything Look the Same?
(A Painter’s Perspective)
Image 5 — If AI Makes Us More Creative, Why Does Everything Look the Same?
(A Painter’s Perspective)
Image 6 — If AI Makes Us More Creative, Why Does Everything Look the Same?
(A Painter’s Perspective)
Image 7 — If AI Makes Us More Creative, Why Does Everything Look the Same?
(A Painter’s Perspective)
Image 8 — If AI Makes Us More Creative, Why Does Everything Look the Same?
(A Painter’s Perspective)
Image 9 — If AI Makes Us More Creative, Why Does Everything Look the Same?
(A Painter’s Perspective)
Image 10 — If AI Makes Us More Creative, Why Does Everything Look the Same?
(A Painter’s Perspective)
Image 11 — If AI Makes Us More Creative, Why Does Everything Look the Same?
(A Painter’s Perspective)
Image 12 — If AI Makes Us More Creative, Why Does Everything Look the Same?
(A Painter’s Perspective)
Image 13 — If AI Makes Us More Creative, Why Does Everything Look the Same?
(A Painter’s Perspective)
Image 14 — If AI Makes Us More Creative, Why Does Everything Look the Same?
(A Painter’s Perspective)

If AI Makes Us More Creative, Why Does Everything Look the Same? (A Painter’s Perspective)

QUICK NOTE: the question in the title is rhetorical. The carousel explains the nuance and explores several related issues beyond the first slide.

If the design does not work for you, tell me specifically what you would improve. I am still refining the format, so constructive feedback is welcome.

.....

I’m a painter who sometimes writes, and this visual essay started with an odd discovery: I had used the name “Elias Thorne” in a short story, only to realize that AI models often return to that same name, along with motifs like lighthouse keepers, cathedrals, glossy landscapes, and other familiar patterns.

From an artist’s point of view, the question isn’t just whether AI is good or bad, but what happens to authorship and creativity when the tool starts making choices for us.

AI can boost productivity and even enhance individual works, but if we all lean on the same models, it might steer us toward similar ideas, characters, and visual styles.

This carousel looks at visual convergence, originality, transparency, and the role of human intention, with AI-generated images clearly labeled and sources included.

So where’s the line, does AI broaden personal creativity while making our collective output more uniform?

u/MenegattiArt — 1 day ago
▲ 88 r/sdforall+1 crossposts

MiniMax H3 Creator update: presets, and three nodes are now one

Posted this pack here last week. What's happened since:

The sampling knobs I said I'd add if people wanted them are in. Both of H3's flow shifts, since it samples picture and sound on separate schedules, plus a cache pill with FirstBlockCache, TeaCache or core's own EasyCache behind it. Still no custom sigmas, same reason as last time.

Creator and Timeline are one node now. Click under the prompt and the shot becomes a timeline. Delete cards back down to one and it's a shot again. Old workflows load unchanged, Timeline nodes included.

Presets are the new one. Save a setup and put it back in sections, so you can drop a canvas and a step count onto a shot you've already written without touching the prompt. It saves the sampler row as well as the node blob, which matters because the row is where the turbo schedule and the step count live.

The better half of it: you can build a preset from a finished render. The workflow is already embedded in the mp4, so you point at the good one from three prompts ago and get the whole setup back.

Fixed from your reports: the gallery no longer freezes on big libraries, the settings page stopped resetting fields you hadn't touched, and a text-only render no longer loads both VAEs.

Coming next, on a branch and not merged yet, is a faces pill. H3 draws a face worse the smaller the head is in frame, and that's about head size rather than resolution, so it's still there at 768 and upscaling doesn't reach it. So it asks the model the same question again with the face filling the canvas and composites the answer back under a feathered mask, once per pass, re-cropping every frame so a push-in doesn't leave the face small inside a fixed box. Detection is core's SAM3, so there's nothing extra to install. The method is Carasibana's ComfyUI-H3-FaceRefine and zuanfilm's graph on top of it.

Same branch also stops the node randomizing your seed between renders, and puts the last one you actually ran a click away.

https://github.com/roadmaus/ComfyUI-MiniMax-Creator

u/Fine_Rhubarb3786 — 5 days ago
▲ 82 r/sdforall+1 crossposts

ComfyUI Tutorial MiniMax H3 4 Steps Lora + Upscaling + 2X Faster Generation! Best Settings for 2K AI

Hello everyone

Want to get faster MiniMax H3 video generation without sacrificing quality? In this tutorial, I’m testing the new H3 LoRA together with Sage Attention, Sol Attention, and Spectrum nodes to find the best combination for speed and quality. The goal is to push MiniMax H3 as far as possible while cutting generation times by up to , then upscale the results with LTX Upscaler to reach a stunning 2432 × 1344 (2K-class) resolution. By combining both H3 LoRA together with Sage Attention, Sol Attention, and Spectrum nodes I generated video at 0.8 megapixel using "**RTX3060 6GB 16GB RAM "**and I got

 13 minutes vs 41 minutes at 8 steps

 27 minutes vs 52 minutes at 20 steps

LTX 2.3 Upscaler 11 minutes to get 2432 × 1344 resolution

Workflow link

https://civitai.com/articles/34028/comfyui-tutorial-minimax-h3-4-steps-lora-upscaling-2x-faster-generation-best-settings-for-2k-ai

youtu.be
u/cgpixel23 — 4 days ago
▲ 63 r/sdforall+2 crossposts

ComfyUI Tutorial First Test Of LTX 2 5 New Model Better Than Minimax H3

First testing of the LTX 2.5 with a 6GB VRAM I’ve created a low-VRAM workflow that supports both Text-to-Video and Image-to-Video, optimized specifically for GPUs with 6GB of VRAM. The goal is to make LTX 2.5 more accessible to users who don’t have high-end GPUs, while keeping the workflow simple and easy to use. If you’re interested in testing LTX 2.5 on a 6GB GPU, check out the workflow and let me know how it performs on your setup!

Video Resolution : 1344x768 for 7 seconds video
Generated Time : 10 min

The model seems very fast the motions are better, lipsync and sound too, however the quality in minimax is better to me

Workflow link

https://civitai.com/articles/33897/comfyui-tutorial-first-test-of-ltx-2-5-new-model-better-than-minimax-h3

youtu.be
u/cgpixel23 — 7 days ago
▲ 15 r/sdforall+2 crossposts

Noisy outputs in Krea 2 in ComfyUI

I have a problem. Recently my outputs in Krea 2 have blotches/banding/noise. This occurs on smooth surfaces and the edges of objects.

Images without reddit compression: img1, img2, img3

Things I tried today, but that didn't make any difference:

  • Krea2 turbo int8 convrot / fp8 / bf16 / raw fp8
  • Windows ComfyUI 0.31 / 0.30 / 0.26
  • comfy-kitchen 0.2.30 / 0.2.27 / 0.2.10
  • PyTorch 2.12 cu132 / 2.12 cu130 / 2.8 cu128
  • Older NVIDIA driver

anyone else experienced this? any ideas?

test prompt: "a stylized 3D character render of a young woman, waist-up, neutral background, clean materials, and a contemporary high-quality character design presentation, dark environment, gray walls, low light"

u/y3kdhmbdb2ch2fc6vpm2 — 10 days ago
▲ 15 r/sdforall+1 crossposts

Nexfocus Walkthrough – Free & Open-Source Connected AI Workspace with the built-in GIMP plug-in

If an image model is a horse, text prompting is like trying to guide it with verbal commands alone: useful, but too imprecise for fine control. Inpainting, LoRAs, and ControlNets add the bridle, reins, and pieces of the harness, but they still did not feel like a complete system. Nexfocus began with a question: what would it take to build the whole harness around the model?

Answering that question meant following the entire generation process first. We had to understand how each part loads, works, moves, waits, hands its result to the next part, and makes room when its job is done. That expedition became Nexfocus.

AI image models already know how to render- they understand photorealism, lighting, and texture better than any software framework. The real struggle has always been context control: feeding the model the exact spatial and structural context it needs to understand what to render.

This is where human input shines. By adding rough blocks of color and basic shapes to define the composition and color palette, the burden of text prompting is significantly reduced. Instead of struggling to describe a complex 3D scene in words, the artist controls spatial intent and lets the AI do what it already knows best—handling lighting, global occlusion, perspective consistency, and fine rendering.

Through the built-in GIMP plugin and Staging Palette, you can seamlessly composite background plates, sketch structural boundary guides for outpainting, or block out color guidance in GIMP with one click, then pass them back into Nexfocus to drive the diffusion pipeline.

Two development anchors shaped the journey: a GTX 1050 with 3 GB of VRAM and Colab Free's T4 with 12.7 GB of system RAM. We proved that full FP16 SDXL checkpoints and Flux Fill workflows could run in both environments by rethinking how the pipeline uses the hardware available to it.

Two important lessons emerged from the road:

- Keeping the GPU working without interruption became paramount. To do that, we had to find a way to keep feeding it the weights it needed when it needed them.

- Every part of the pipeline must independently account for what it owns, where it belongs, when it can be reused, and when it should make room for something else.

Throughout this journey, my conversations with PyTorch often felt like this:

> PyTorch: "Don't you have a bunch of H100s lying around in your backyard?"

>

> Me: "No. What if every component has to justify exactly where it lives?"

>

> PyTorch: "Get a bigger machine."

Those conversations eventually became the architecture. Nexfocus is both the working application and the field notebook: a record of the constraints, wrong turns, and discoveries that shaped the path forward.

The path is open now. I hope you'll take a walk along the path we built and enjoy the scenery.

GitHub: https://github.com/magekinnarus/Nexfocus

u/magekinnarus — 8 days ago

# ControlNet for FLUX.2

This workflow demonstrates the new ComfyUI custom nodes I developed to implement ControlNet for FLUX.2-dev.

Workflow: JSON | Drag-and-drop PNG

JLC Flux2 ControlNet provides, to the best of my knowledge, the first complete, validated ComfyUI implementation of Alibaba PAI's FLUX.2-dev-Fun-Controlnet-Union-2602.

This implementation is for the FLUX.2-dev ControlNet path built around that Union model. It is not for FLUX.2 Klein or the lightweight Klein-style variants many people currently use; I am currently working on a separate strategy to extend this functionality to those models.

It is also worth making an important distinction: reference images are not ControlNet. There are workflows that feed pose maps, depth maps, edges, or other ControlNet-style hint images into FLUX.2's native reference-image system. Those images can certainly influence composition and structure, and they can often produce a usable approximation, but this is still reference-image conditioning, which is a completely different conditioning mechanism. It does not load a ControlNet model, does not execute a ControlNet branch, and should not be confused with one.

This workflow actually loads and runs Alibaba PAI's FLUX.2 ControlNet model.

The two JLC nodes that enable that path are the FLUX.2 ControlNet Loader and the ControlNet Orchestrator.

The Orchestrator also introduces a non-recursive composition method that lets several control types share a single loaded Union model instead of building a conventional chain of ControlNet applications.

The example shown here uses three controls generated from the same source image:

  • DWPose
  • Depth Anything
  • Color

That is really the point of this workflow: there are very few special pieces required to add actual ControlNet capability to FLUX.2-dev.

Some of the other nodes shown are from my JLC ComfyUI Nodes package and are there mainly for convenience—loading, resizing, preprocessing, LoRAs, and general workflow ergonomics. You can replace those with your preferred ComfyUI nodes.

This is not simply a repackaging of existing ControlNet nodes. The contribution here is making this capability available as a complete ComfyUI implementation of Alibaba PAI's actual FLUX.2 ControlNet model. The Orchestrator also provides practical multi-control composition where a finished implementation was previously missing.

All of the JLC nodes can be installed through the ComfyUI Custom Node Manager, and the repositories contain the documentation and explanation of the implementation.

I hope you find them useful, and I'd be very interested to see what people build with them!

u/jessidollPix — 10 days ago

SenseNova U1.5 vs Nano Banana vs GPT Image 2 — which one actually looks editorial?

Honestly, I expected the closed models to win this pretty easily.

They did on realism—but not necessarily on art direction.

I ran the same fashion-editorial prompt through SenseNova U1.5, Nano Banana, and GPT Image 2. These are the first outputs—no rerolls, edits, or post-processing.

My take:

- Nano Banana wins on background detail. The station feels fuller and more believable.

- GPT Image 2 has the best atmosphere—darker, moodier, and more cinematic.

- SenseNova U1.5 gave me the strongest fashion-editorial look. The styling, composition, and color treatment feel the closest to an actual campaign.

Totally subjective, but for this specific fashion use case, U1.5 feels like it gets about 80% of the way to the closed models overall. And on art direction alone, I actually prefer it.

The remaining gap is mostly in realism: the face and skin still have a slightly plastic-looking AI sheen, while Nano handles the environment better and GPT feels more naturally cinematic.

If the goal is a fashion campaign rather than pure photorealism, I’d pick U1.5. It sells the outfit and art direction best, which is a pretty strong result for an open-source 8B preview model.

Obviously, one prompt isn’t a benchmark. I’ll put the full generation prompt in the comments.

Model links- SenseNova U1.5:

- https://huggingface.co/sensenova/SenseNova-U1.5-8B-MoT-Preview

- https://github.com/OpenSenseNova/SenseNova-U1

Would you trade some realism for stronger art direction, or does the plastic-looking skin already kill the U1.5 result for you?

u/Pristine_Weight_4705 — 8 days ago
▲ 29 r/sdforall+2 crossposts

👋 Welcome to r/MinimaxVideo - Introduce Yourself and Read First!

r/MiniMaxVideo

The community for everything related to MiniMax's AI video models.

Share videos, prompts, workflows, tutorials, news, updates, benchmarks, tips, and discussions about MiniMax video generation—including H3, Hailuo, and future releases.

Whether you're creating cinematic clips, testing prompt techniques, comparing models, or showcasing your best generations, you're welcome here.

Please include your prompt, settings, or workflow whenever possible to help others learn and reproduce your results.

Play with Minimax locally: https://huggingface.co/Comfy-Org/MiniMax-H3

Online: https://hailuoai.com/video

u/Hefty_Scallion_3086 — 14 days ago