
Qwen-Image 2.1 replaces FLUX for my covers: 7 times faster on the same 16 GB card, and what I lose
I ran the same prompts and seeds through the FLUX that makes my covers and through Qwen-Image-2.1-Turbo quantized to NF4, on the same 16 GB card. Qwen takes 8.4 s per image against 61 s, writes 13 of 18 texts exactly against 7, and draws the requested scene in 11 of 15 covers against none. I lose the LoRA's atmosphere, and the version that fits 16 GB writes 4 fewer texts than unquantized Qwen.
The covers on this blog come from FLUX.1-dev with a pixel-art LoRA, on a 16 GB RTX 4070 Ti SUPER. Each image takes a minute, text often comes out misspelled and counts are ignored. I tested whether Qwen-Image-2.1-Turbo can replace it on the same card. Everything below uses the same prompt and the same seed on each side.
The cover step, today and with Qwen
Text inside the image
Six prompts with an exact text, three seeds each. The texts are in Spanish on purpose: accents and the decimal comma are where models slip. I checked every image at full size: ✓ if every character came out as asked, ≈ if only a space or a split word is off, ✗ if a character is wrong, extra or missing. A decimal comma turned into a point or a missing accent counts as ✗. A line break between words does not count, and neither does text the prompt did not ask for.
Asked for MEDIDO EN UN PORTÁTIL
Full prompt
A minimalist concert poster on a brick wall, bold sans-serif headline reading "MEDIDO EN UN PORTÁTIL", flat colors, soft daylight
Asked for 18 a 21 % más rápido
Full prompt
A green school chalkboard with neat white chalk handwriting that reads "18 a 21 % más rápido", a piece of chalk on the ledge, classroom light
Asked for efraingaray.com
Full prompt
A pink neon sign glowing on a dark tiled wall at night, the sign spells "efraingaray.com", reflections on wet floor
Asked for Ocho modelos de decisión
Full prompt
A hardcover book lying on a wooden desk, the cover title printed in large serif letters reads "Ocho modelos de decisión", warm lamp light
Asked for AUC 0,966
Full prompt
A glass jar on a kitchen shelf with a white paper label printed in black capitals "AUC 0,966", shallow depth of field
Asked for PostgreSQL 19
Full prompt
A yellow road sign on a mountain road at dusk, black letters reading "PostgreSQL 19", pine trees behind
Unquantized Qwen, on a rented 48 GB A6000, wrote 17 of the 18 texts exactly; the one left halfway was «Post / greSQL», split over two lines. The NF4 version got 13: five got worse and one improved. I don’t isolate the cause, because besides quantizing, the GPU and the torch version change too. The clearest error repeats: on the chalkboard NF4 wrote «recido» all three times. Both Qwens ignored the poster in the first prompt and painted the text straight onto the wall.
This blog’s covers
Five real cover prompts, three seeds each. I checked them against what the prompt asks for: one robot, its action and the key objects. ✓ if everything is there, ≈ if one element is off, ✗ if the action or the main object is missing. These are the unchanged production prompts, and they are long: CLIP, one of FLUX’s two encoders, cuts their end at 77 tokens, right where the scene goes; the other one, T5, gets them whole. I compare the two workflows as they run, not each model’s ability with a prompt made to measure.
p0 Robot with a magnifier at a three-block podium; the tallest one cracked
Full prompt
umempart, pixel art, flat colors, retro 16-bit game sprite style, crisp pixel edges, dark navy background, teal and amber accent lighting, clean composition with empty space, no text, no letters, a single cute robot with a round helmet, one big glowing lens eye and an antenna with an amber spark on top, standing on the left third, holding a large magnifying glass up to a winners podium of three stacked blocks, the tallest middle block of the podium has a big crack running down it, wide horizontal composition, the robot is teal with amber details
p1 Robot pulling a paper scroll that piles up in loops
Full prompt
umempart, pixel art, flat colors, retro 16-bit game sprite style, crisp pixel edges, dark navy background, teal and amber accent lighting, clean composition with empty space, no text, no letters, a single cute robot with a round helmet, one big glowing lens eye and an antenna with an amber spark on top, standing on the left third, pulling an extremely long paper scroll list that unrolls across the whole floor behind it and piles up in loops, the scroll covered in tiny rows of dots, dark navy background, wide horizontal composition, the robot is teal with amber details
p2 Robot holding a giant key in front of a round vault door
Full prompt
umempart, pixel art, flat colors, retro 16-bit game sprite style, crisp pixel edges, dark navy background, teal and amber accent lighting, clean composition with empty space, no text, no letters, a single cute robot with a round helmet and one big glowing lens eye holding a giant brass key in front of a round vault door, wide horizontal composition
p3 Robot judge at a desk sorting film frames into two baskets
Full prompt
umempart, pixel art, flat colors, retro 16-bit game sprite style, crisp pixel edges, dark navy background, teal and amber accent lighting, clean composition with empty space, no text, no letters, a single cute robot judge with a round helmet sitting at a desk sorting a pile of film frames into two baskets, wide horizontal composition
p4 Robot with an umbrella over a city of seven buildings
Full prompt
umempart, pixel art, flat colors, retro 16-bit game sprite style, crisp pixel edges, dark navy background, teal and amber accent lighting, clean composition with empty space, no text, no letters, a single cute robot with an umbrella looking at a cloudy sky over a small city of seven buildings, wide horizontal composition
FLUX with the LoRA has the style: fog, rim light and depth. But the subject comes out small and the action gets lost: the key floats away from the robot, nobody sorts the frames and the city has dozens of buildings. In 6 covers one element is off and in 9 the main thing is missing. Qwen draws the scene in 11 of 15, with a big robot and a clear action, though on a flat background. It does not count right either: I asked for seven buildings and it drew 6, 6 and 5.
Speed
nvidia-smi above the baseline at the start: Qwen NF4 (VAE in tiles) 11.9 GiB, FLUX 13.1 GiB.
PyTorch peak: 36.8 GiB at 1344×768 and 32.6 GiB at 2752×1536.
Medians from timing.jsonl and the FLUX logs, computed by timing.py.
Now in the studio, with the NF4 copy saved on disk so it does not requantize on every start, three covers end to end took 33.6 s, model loading and SSH included. This post’s cover was the first one that came out that way. Then I regenerated all 46 of the blog’s covers with the same engine: two seeds per post, I picked by eye the one that followed the scene and redid the two that did not.
What broke and how I fixed it


To fit everything in 16 GB, the VAE decoded the image in tiles, and that left faint vertical purple streaks on light surfaces. They did not come from NF4: they show up the same in bf16 when decoding in tiles. With 512 px tiles the streak fades but stays, and 768 px tiles no longer fit. What worked was splitting the job in two phases: first the latents of every cover, then I take the transformer and the encoder off the GPU and the VAE decodes each image whole. In the label test, decoding and saving took 0.77 s, with a 7.1 GiB PyTorch peak in that phase.


Verdict
For my covers, with these prompts, Qwen NF4 becomes the engine of the studio’s cover profile. Seven times faster on the same card, with readable text and the scene I ask for, is worth more than the LoRA’s atmosphere. FLUX with the LoRA stays installed as a fallback for scenes where mood matters more than content.
If the text has to be exact, I still check it by eye: unquantized Qwen wrote 17 of 18 on the A6000 I rented, and the 16 GB version, 13.
Frequently asked questions
Does Qwen-Image-2.1-Turbo fit on a 16 GB GPU?
Not in bf16 with everything on the GPU: the pipeline is 32.5 GB on disk and my attempt with offload to 30 GB of RAM failed. It runs with the transformer and the text encoder quantized to NF4 with bitsandbytes. The quantized copy takes 10.6 GiB on disk and, on an RTX 4070 Ti SUPER, nvidia-smi showed 11.9 GiB above what the card held at the start, with the VAE in tiles.
How much faster than FLUX is it?
On the same card at 1344×768, Qwen NF4 took 8.4 s per image (median of 36) and FLUX.1-dev Q6_K with the cover LoRA took 61 s (median of 15). About 7 times. Qwen uses 8 steps and FLUX 30. In the studio, three covers end to end, model loading included, took 33.6 s.
Does it write text better than FLUX?
In these 18 images with a requested text, checked by eye, Qwen NF4 wrote 13 exactly and FLUX 7. Unquantized Qwen, on a 48 GB A6000, wrote 17. The NF4 version wrote «recido» instead of «rápido» all three times.
Why not 8-bit instead of NF4?
Because the transformer in 8 bits (LLM.int8) returned a near-uniform purple field. I isolated it on an A6000 by quantizing one component at a time: with the transformer in NF4 the image comes out right.
Where did the purple streaks come from?
From decoding the VAE in tiles, which was needed to fit everything in 16 GB. They show up in bf16 with tiles too. I removed them by making the latents first, taking the transformer and the encoder off the GPU and decoding the whole image: decoding and saving took 0.77 s, with a 7.1 GiB PyTorch peak in that phase, in the label test.
Can I reproduce it?
Yes. The prompts, the generation scripts, the image-by-image review and the timings are in a gist linked in the sources.
Sources
- Qwen-Image-2.1-Turbo, the model card.
- The Qwen-Image pipeline in diffusers and quantization with bitsandbytes.
- QLoRA, the paper that defines NF4, and LLM.int8(), the 8-bit one.
- FLUX.1-dev and its GGUF version, where the Q6_K comes from.
- gemma3:4b, the automatic control reader.
- The prompts, the scripts, the image-by-image review and the timings.
Measured on October 10, 2026. FLUX.1-dev Q6_K GGUF with the ume_modern_pixelart LoRA (flux_batch.py, the studio’s engine) and Qwen-Image-2.1-Turbo NF4 (torch 2.14.1, diffusers 0.41.0.dev0 at commit 1d5d056, transformers 5.19.0, bitsandbytes 0.50.2) on the same 16 GB RTX 4070 Ti SUPER; Qwen bf16 on a rented 48 GB RTX A6000 (torch 2.8.0). 1344×768, seeds 7, 21 and 88.
Comments
No comments yet. The first one is yours.

































































