New ComfyUI update may change how Minimax H3 interprets the prompt format you use - Re: Tokenizer Fix
https://github.com/Comfy-Org/ComfyUI/pull/15808/changes
https://redd.it/1vvrlxg
@rStableDiffusion
https://github.com/Comfy-Org/ComfyUI/pull/15808/changes
https://redd.it/1vvrlxg
@rStableDiffusion
GitHub
Minimax-H3: Add missing special tokens by kijai · Pull Request #15808 · Comfy-Org/ComfyUI
The Minimax-H3 tokenizer extends standard Qwen2.5/3 with 7 tokens declared only in tokenizer_config.json's additional_special_tokens (not in tokenizer.json/vocab.json): , , <|cutoff|...
Fizgig now trains LoRAs on AMD Radeon - Flux 2 Klein, Krea 2 and MiniMax H3
https://redd.it/1vvqrgp
@rStableDiffusion
https://redd.it/1vvqrgp
@rStableDiffusion
Best opensource image model?
https://preview.redd.it/p90tt3gru1lh1.jpg?width=1232&format=pjpg&auto=webp&s=167847b7b1328ca1abafec942370fdf6f69406f3
opensource AI has been dominating LLMs and video generation but what about image gen? is there any opensource model that can match gpt-image2?
Edit: The reason I am asking this is because lately I haven't been active much on image generation communities. And the leaderboards are a bit confusing and most of them are filled with closed source unlike the llm and video gen leaderboards.
I am very much comfortable with ComfyUI since I've used it in the past for flux.
My use case is for posters and branding. Images with a lot of text.
Edit2: Thanks a lot everyone! I really appreciate the info. Here's the summary:
Krea2 is best overall but gptimage1.5 level.
Ideogram4 for text and branding.
Flux Klein 9b for image editing.
Z-image for realism
Anima and illustrious (by onoma AI) for anime.
Here's the workflow I've decided on:
Krea2/Ideogram4 = Base image generation.
Flux Klein 9B/QwenImage2512 = inpainting.
Wan2.2 low noise = Upscaling.
https://redd.it/1vvwxhp
@rStableDiffusion
https://preview.redd.it/p90tt3gru1lh1.jpg?width=1232&format=pjpg&auto=webp&s=167847b7b1328ca1abafec942370fdf6f69406f3
opensource AI has been dominating LLMs and video generation but what about image gen? is there any opensource model that can match gpt-image2?
Edit: The reason I am asking this is because lately I haven't been active much on image generation communities. And the leaderboards are a bit confusing and most of them are filled with closed source unlike the llm and video gen leaderboards.
I am very much comfortable with ComfyUI since I've used it in the past for flux.
My use case is for posters and branding. Images with a lot of text.
Edit2: Thanks a lot everyone! I really appreciate the info. Here's the summary:
Krea2 is best overall but gptimage1.5 level.
Ideogram4 for text and branding.
Flux Klein 9b for image editing.
Z-image for realism
Anima and illustrious (by onoma AI) for anime.
Here's the workflow I've decided on:
Krea2/Ideogram4 = Base image generation.
Flux Klein 9B/QwenImage2512 = inpainting.
Wan2.2 low noise = Upscaling.
https://redd.it/1vvwxhp
@rStableDiffusion
Sparse Attention, Harder, Better, Faster, Stronger
The nodes in [https://github.com/Zironic/H3-Optimizations](https://github.com/Zironic/H3-Optimizations) have been rewritten to replace the default Sparge Attention backend with a custom Sparse Comfy Kitchen backend.
This comes with some benefits.
* Users no longer have to worry about Sparge being installed properly. All required kernels for supported GPUs are provided directly. Should work on both Windows and Linux.
* Most users should be seeing 5-20% increases in speed for the attention part of compute.
* New backend should use about 500MB less VRAM
* New backend has slightly lower quantization error.
* Apparently in the previous version, the intended chunked kitchen QKV path never properly shipped so the memory optimization node should now actually be slightly speed positive even when used without the Sparse Attention node.
**Caveat:** I've only tested the nodes against the comfy pruned\_int8\_convrot weights. Other versions may work but they're not tested.
As the nodes currently rely on comfy-kitchen 0.2.31 you need ComfyUI v0.33.0 or later.
**IMPORTANT: sparse attention is not free speed. The percentage is effectively a prompt-adherence/quality budget.**
**Density isn't just a speed setting, and its quality effect depends on where you apply it in the diffusion schedule.**
**Early steps:** attention density has a large effect on prompt/action adherence and the overall generation trajectory.
**Middle/later steps:** lowering density tends to show up more as motion/temporal artifacts and lost fine motion detail.
So `10% retained` doesn't simply mean “90% of the quality is gone.” It means you're giving sparse attention very little information to work with, and *what breaks depends heavily on the sampling step.*
PlagueKind's `sparsity_ratio=0.9` means **90% discarded / 10% retained**. My node expresses the inverse quantity, so `Video attention retained=0.10` is the comparable setting. The defaults therefore aren't equivalent.
https://redd.it/1vw1ad0
@rStableDiffusion
The nodes in [https://github.com/Zironic/H3-Optimizations](https://github.com/Zironic/H3-Optimizations) have been rewritten to replace the default Sparge Attention backend with a custom Sparse Comfy Kitchen backend.
This comes with some benefits.
* Users no longer have to worry about Sparge being installed properly. All required kernels for supported GPUs are provided directly. Should work on both Windows and Linux.
* Most users should be seeing 5-20% increases in speed for the attention part of compute.
* New backend should use about 500MB less VRAM
* New backend has slightly lower quantization error.
* Apparently in the previous version, the intended chunked kitchen QKV path never properly shipped so the memory optimization node should now actually be slightly speed positive even when used without the Sparse Attention node.
**Caveat:** I've only tested the nodes against the comfy pruned\_int8\_convrot weights. Other versions may work but they're not tested.
As the nodes currently rely on comfy-kitchen 0.2.31 you need ComfyUI v0.33.0 or later.
**IMPORTANT: sparse attention is not free speed. The percentage is effectively a prompt-adherence/quality budget.**
**Density isn't just a speed setting, and its quality effect depends on where you apply it in the diffusion schedule.**
**Early steps:** attention density has a large effect on prompt/action adherence and the overall generation trajectory.
**Middle/later steps:** lowering density tends to show up more as motion/temporal artifacts and lost fine motion detail.
So `10% retained` doesn't simply mean “90% of the quality is gone.” It means you're giving sparse attention very little information to work with, and *what breaks depends heavily on the sampling step.*
PlagueKind's `sparsity_ratio=0.9` means **90% discarded / 10% retained**. My node expresses the inverse quantity, so `Video attention retained=0.10` is the comparable setting. The defaults therefore aren't equivalent.
https://redd.it/1vw1ad0
@rStableDiffusion
GitHub
GitHub - Zironic/H3-Optimizations
Contribute to Zironic/H3-Optimizations development by creating an account on GitHub.
Minimax H3 Huge Quality Difference between Cloud and Local use
Hi.
I have a decent h3 workflow that I built for a loca use. It use turbo lora etc... If i use the defaut settings in the goal of getting the highest quality possible, meaning res_multistep simple 20 steps or more, I got also good results, but this is not even close to the results you can get on platforms like kie or wavespeed at 768P.
I already convert properly the prompt to the correct H3 digest form, so I'm wondering what's different between local and cloud use of h3? I don't talk about the 2K quality, only 768P, I'm not able to reach the sames results locally, do you guys have maybe workflows, settings, or suggestions to try reaching the same quality level in comfyui ?
https://redd.it/1vw4cgy
@rStableDiffusion
Hi.
I have a decent h3 workflow that I built for a loca use. It use turbo lora etc... If i use the defaut settings in the goal of getting the highest quality possible, meaning res_multistep simple 20 steps or more, I got also good results, but this is not even close to the results you can get on platforms like kie or wavespeed at 768P.
I already convert properly the prompt to the correct H3 digest form, so I'm wondering what's different between local and cloud use of h3? I don't talk about the 2K quality, only 768P, I'm not able to reach the sames results locally, do you guys have maybe workflows, settings, or suggestions to try reaching the same quality level in comfyui ?
https://redd.it/1vw4cgy
@rStableDiffusion
Reddit
From the StableDiffusion community on Reddit
Explore this post and more from the StableDiffusion community
[MiniMax H3] Ultimate SD Upscale can actually fix your bad/low-res generations
https://youtu.be/naUuN5k7_Ew
https://redd.it/1vwgoy2
@rStableDiffusion
https://youtu.be/naUuN5k7_Ew
https://redd.it/1vwgoy2
@rStableDiffusion
YouTube
Fantastic Walk (Ultimate SD Upscale and MiniMax H3, 2560x1440px with 16 GB VRAM locally in ComfyUI)
Ultimate SD Upscale can actually fix your bad/low-res generations.
In this comparison initial clips were made with MiniMax H3 at 1504x832px resolution and then upscaled to 2560x1440px with Ultimate SD Upscale nodes: https://github.com/lisitskyaa/ComfyUI…
In this comparison initial clips were made with MiniMax H3 at 1504x832px resolution and then upscaled to 2560x1440px with Ultimate SD Upscale nodes: https://github.com/lisitskyaa/ComfyUI…
Minimax SEED HUNTER workflow released!
https://www.youtube.com/watch?v=H8JSzhkOmXA
https://redd.it/1vwe2wk
@rStableDiffusion
https://www.youtube.com/watch?v=H8JSzhkOmXA
https://redd.it/1vwe2wk
@rStableDiffusion
YouTube
Minimax SEED HUNTER Workflow Released!
Link to workflow: https://civitai.red/models/2881362/minimax-seed-hunter-workflow-optimized-fast-latent-upscaler-speedups
Please consider supporting me on Patreon: https://www.patreon.com/cw/foxfuressence
It's free to join and I post all of my work there…
Please consider supporting me on Patreon: https://www.patreon.com/cw/foxfuressence
It's free to join and I post all of my work there…
Civitai now has a closed-source models option
I haven't seen a post about this here, and I'm curious what you think about it.
On August 14 Civitai rolled out the option for creators to choose "permanent paid access - selling with no time cap". Previously the only option was temporary "early access".
Let's call this what it is, closed-source. Yes, you can get the weights for a relatively small fee, and yes it's on a very small scale compared to Nano Banana and Midjourney. But a permanent paywall still fits the definition.
---
Personally, I block all creators on Civitai who choose permanent paywall and encourage you to do the same.
---
Here's why:
I'm not opposed to Civitai making money or for all options for model creators to make money. They can do that without permanent paywalls.
IMO, open source AI is a fair trade: models are trained on the hard work of many human artists who aren't compensated, but everyone benefits from the ability to create more art more easily. Closed source is an unfair trade: you have to pay a middle man to access the contributions of others who won't be compensated.
Small scale model creators do some hard work too. But for example, for a lora that reproduces the style of an animated film: the lora creator spent at most a dozen hours of work, while just one of the artists on that film spent thousands of hours of work. If a massive models like Krea2 are free, and if giant "hobby" finetunes like Chroma are free, I can't justify paying any price for a 5,000 step lora except as an optional donation of appreciation.
So far, few creators have chosen the permanent paywall closed-source option. But that could easily change if Civitai made it the default option. They already made an extra 1-buzz fee-to-creator per generation the default, and many models have that.
That's my opinion. If you agree, then the only tool you have to disincentivize that potential is to not pay for these models (disincentive Civitai) and block these creators (disincentive creators).
https://redd.it/1vwilol
@rStableDiffusion
I haven't seen a post about this here, and I'm curious what you think about it.
On August 14 Civitai rolled out the option for creators to choose "permanent paid access - selling with no time cap". Previously the only option was temporary "early access".
Let's call this what it is, closed-source. Yes, you can get the weights for a relatively small fee, and yes it's on a very small scale compared to Nano Banana and Midjourney. But a permanent paywall still fits the definition.
---
Personally, I block all creators on Civitai who choose permanent paywall and encourage you to do the same.
---
Here's why:
I'm not opposed to Civitai making money or for all options for model creators to make money. They can do that without permanent paywalls.
IMO, open source AI is a fair trade: models are trained on the hard work of many human artists who aren't compensated, but everyone benefits from the ability to create more art more easily. Closed source is an unfair trade: you have to pay a middle man to access the contributions of others who won't be compensated.
Small scale model creators do some hard work too. But for example, for a lora that reproduces the style of an animated film: the lora creator spent at most a dozen hours of work, while just one of the artists on that film spent thousands of hours of work. If a massive models like Krea2 are free, and if giant "hobby" finetunes like Chroma are free, I can't justify paying any price for a 5,000 step lora except as an optional donation of appreciation.
So far, few creators have chosen the permanent paywall closed-source option. But that could easily change if Civitai made it the default option. They already made an extra 1-buzz fee-to-creator per generation the default, and many models have that.
That's my opinion. If you agree, then the only tool you have to disincentivize that potential is to not pay for these models (disincentive Civitai) and block these creators (disincentive creators).
https://redd.it/1vwilol
@rStableDiffusion
civitai.red
Civitai Changelog | The latest updates to Civitai
List of the recent features, fixes, and improvements to Civitai.
This media is not supported in your browser
VIEW IN TELEGRAM
Minimax H3 T2VA. You can put 15 different characters or more at the same time on screen.
https://redd.it/1vwr9j5
@rStableDiffusion
https://redd.it/1vwr9j5
@rStableDiffusion
If you’re using MiniMax H3, what prompting tricks have you figured out?
Anyone found useful MiniMax H3 prompting tricks beyond the official guide?
Especially for audio + video prompt structure, camera control, dialogue/audio, consistency, weird tricks that actually work, etc.
Please drop your findings 👇 below so it will help others too.
https://redd.it/1vwtrt1
@rStableDiffusion
Anyone found useful MiniMax H3 prompting tricks beyond the official guide?
Especially for audio + video prompt structure, camera control, dialogue/audio, consistency, weird tricks that actually work, etc.
Please drop your findings 👇 below so it will help others too.
https://redd.it/1vwtrt1
@rStableDiffusion
Reddit
From the StableDiffusion community on Reddit
Explore this post and more from the StableDiffusion community