MOSAIK Project Page

MOSAIK Multi-Patch Content-Aware Spatial Allocation of Image Tokens for Efficient Generation

Mohammadreza Hami*, Mohammadreza Samadi*, Chao Gao, Negar Hassanpour

Huawei Technologies Canada *Equal contribution

mohammadreza.hami@h-partners.com · {mohammadreza.samadi1, chao.gao4, negar.hassanpour2}@huawei.com

Generations shown over 50 denoising steps. White outlines over each image mark the patch size per region: fine p16 patches settle on the animal, coarse p64 patches cover the background. A bar under each image shows 1,024 of 4,096 tokens in use.
MOSAIK re-plans the patch layout at every denoising step: fine p16 patches where coarsening would do the most damage, coarse p32 and p64 patches elsewhere. Token budget: 1,024 tokens per step (the ≈60% fewer-FLOPs tier). The budget applies to steps 2–50; step 1 runs at full p16, so the mean is 1,085 tokens per step at a 1,024-token budget (74% fewer) and 683 at a 614-token budget (83% fewer).

Abstract

Pixel-space diffusion models avoid the reconstruction ceiling of latent diffusion models by generating directly in image space. However, their substantially higher token count makes generation expensive due to the quadratic complexity of self-attention. Several existing efficiency methods reduce this cost by using larger patches at selected denoising steps, thereby representing the image with fewer tokens. Yet, each step still uses a single patch size uniformly across the entire image, overlooking that different regions suffer different fidelity losses when coarsened.

We introduce MOSAIK, a damage-guided framework that varies patch size across regions and denoising steps. MOSAIK adapts the PixelDiT backbone to generate arbitrary heterogeneous patch layouts, and a lightweight predictor uses intermediate denoising features to estimate the fidelity loss caused by coarsening each region. Given a token budget, our damage-guided layout predictor assigns fine patches to sensitive regions and coarse patches elsewhere. Remarkably, while reducing FLOPs by 70% and token count by 83%, MOSAIK matches the full-compute PixelDiT on GenEval and its DPG-Bench score drops by only 1.0 point. Compared to diverse efficiency paradigms, including temporal patch scheduling and feature caching, our approach delivers highly competitive performance at moderate budgets and consistently outperforms these baselines in highly constrained compute regimes.

One wolf image split down the middle. Left half: every 64 by 64 region at p16, a dense uniform grid of 4,096 tokens. Right half: the MOSAIK layout at a 1,024-token budget, with fine patches around the eyes and nose and coarse patches on the fur and background. Uniform p16 · 4,096 tokens MOSAIK · 1,024 tokens
Same prompt and seed. Left: every region at p16. Right: MOSAIK spends a 1,024-token budget on p16 where detail is densest and p32/p64 elsewhere (layout at step 41 of 50).

Motivation

Bigger patches save compute but can lose detail

Each step from p16 to p32 to p64 cuts the tokens by 4×. Applied to the whole image, larger patches can discard local structure and sharp boundaries: our model run at uniform p64 scores 0.65 GenEval (0.73 at uniform p16; PixelDiT 0.74).

Tiger, hummingbird and wolf, each with one crop generated at uniform p16, p32 and p64. The p64 crops lose stripes, throat feathers and fur detail.
The same crop at uniform p16 (PixelDiT), p32 and p64 (our multi-patch model at a single patch size), same seed.

Regions differ in how much coarsening hurts

In this image, coarsening would damage the bee and barely touch the flat background, yet a uniform patch size pays the same for both.

Honeybee on a daisy with a 16 by 16 region grid. The layout predictor estimates damage of 100, 95 and 56 percent of the worst region for three detailed regions on the bee and the flower, and about zero for three flat background regions. A cumulative curve shows that in this image 7 percent of regions hold half of the predicted damage.
Layout-predictor estimate of per-region damage (how much uniform p64 would deviate from the p16 teacher), mean over steps 10–45 of one run.

Method

One backbone for any layout, one predictor for where detail goes

A 1024² image has 256 regions of 64×64 pixels, each tokenized as p16, p32 or p64 (16, 4 or 1 tokens). Two stages adapt PixelDiT to any layout; a third trains a separate layout predictor.

Stages 1–2 self-distil a LoRA student from the frozen p16 teacher; Stage 3 freezes it and trains the layout predictor.

Inference: spend the budget where the damage is

Step 1 runs at full p16. From step 2 on, the predictor reads the previous step's features and predicts the damage drt of every region r. A greedy rule starts every region at p64 and applies the highest-priority refinement that still fits the token budget: p64→p32 costs 3 tokens with priority drt/3, p32→p16 costs 12 tokens with priority drt/12. Any budget works without retraining.

One step on the tiger prompt with a 1,024-token budget (t = 25).

Layouts in action

The layout is re-planned at every step

Play or scrub the 50 steps. The image is the model's running estimate of the final image (x̂0). Step 1 is a full p16 warm-up. In the last steps p16 drops out: with a 1,024-token budget the final two steps are uniform p32; with 614 tokens the layout settles into a p32/p64 mix.

Prompt

Budget
25
Image area by patch size, per step
  • p16 0%
  • p32 0%
  • p64 0%

One model, any budget

Drag the budget: the same predicted damage map, allocated at any token budget, with no retraining.

Prompt

1,024/ 4,096 tokens per step
Image area by patch size
  • p16 0%
  • p32 0%
  • p64 0%

Results

Full-compute GenEval at ≈70% fewer FLOPs

Compared with temporal patch scheduling (DDiT), which also saves compute with larger patches but uses one patch size for the whole image at each step, at four matched compute tiers.

  • MOSAIK
  • DDiT
  • PixelDiT, full compute
GenEval ↑ 0.62 0.66 0.70 0.74 20% 50% 60% 70% FLOPs saved PixelDiT, full compute: 0.74 DDiT · 20% FLOPs saved · GenEval 0.69 DDiT · 50% FLOPs saved · GenEval 0.65 DDiT · 60% FLOPs saved · GenEval 0.65 DDiT · 70% FLOPs saved · GenEval 0.64 MOSAIK · 20% FLOPs saved · GenEval 0.74 MOSAIK · 50% FLOPs saved · GenEval 0.74 MOSAIK · 60% FLOPs saved · GenEval 0.74 MOSAIK · 70% FLOPs saved · GenEval 0.74 MOSAIK 0.74 DDiT 0.64 DPG-Bench ↑ 80 81 82 83 84 20% 50% 60% 70% FLOPs saved PixelDiT, full compute: 83.5 DDiT · 20% FLOPs saved · DPG 83.1 DDiT · 50% FLOPs saved · DPG 81.6 DDiT · 60% FLOPs saved · DPG 81.5 DDiT · 70% FLOPs saved · DPG 80.8 MOSAIK · 20% FLOPs saved · DPG 83.6 MOSAIK · 50% FLOPs saved · DPG 83.4 MOSAIK · 60% FLOPs saved · DPG 83.1 MOSAIK · 70% FLOPs saved · DPG 82.5 MOSAIK 82.5 DDiT 80.8
DPG-Bench and GenEval at four compute tiers
Method ≈20% FLOPs saved ≈50% ≈60% ≈70%
DPGGenEval DPGGenEval DPGGenEval DPGGenEval
PixelDiT, full compute n = 50 DPG 83.5 · GenEval 0.74
Fewer steps n = 38 / 24 / 19 / 16 83.30.7082.50.6781.50.6580.60.64
TeaCache 83.50.7283.50.7282.80.7182.60.71
TaylorSeer 83.70.7183.60.7182.70.6882.40.68
DDiT 83.10.6981.60.6581.50.6580.80.64
MOSAIK (ours) 83.60.7483.40.7483.10.7482.50.74
MOSAIK tokens saved 25%59%74%83%

Same prompt and seed at ≈60% fewer FLOPs

PixelDiT is the full-compute reference. Bottom row: zoomed crop.

Compare with full compute

fewer FLOPs

Drag the divider. The budget applies to MOSAIK's steps 2–50 (step 1 runs at full p16); PixelDiT uses 4,096 tokens at every step. Same seed.

Placement is the gain

Same model, same token budget, only the layout differs: damage-guided layouts beat random ones at every tier.

GenEval ↑ 0.68 0.70 0.72 0.74 0.76 20% 50% 60% 70% FLOPs saved random · 20% FLOPs saved · GenEval 0.72 random · 50% FLOPs saved · GenEval 0.72 random · 60% FLOPs saved · GenEval 0.71 random · 70% FLOPs saved · GenEval 0.70 damage-guided · 20% FLOPs saved · GenEval 0.74 damage-guided · 50% FLOPs saved · GenEval 0.74 damage-guided · 60% FLOPs saved · GenEval 0.74 damage-guided · 70% FLOPs saved · GenEval 0.74 damage-guided random DPG-Bench ↑ 81 82 83 84 20% 50% 60% 70% FLOPs saved random · 20% FLOPs saved · DPG 83 random · 50% FLOPs saved · DPG 82.1 random · 60% FLOPs saved · DPG 82.1 random · 70% FLOPs saved · DPG 81.6 damage-guided · 20% FLOPs saved · DPG 83.6 damage-guided · 50% FLOPs saved · DPG 83.4 damage-guided · 60% FLOPs saved · DPG 83.1 damage-guided · 70% FLOPs saved · DPG 82.5 damage-guided random

MOSAIK composes with caching

Caching saves compute across time, MOSAIK across space. Together they keep GenEval at 0.73 even at ≈80% fewer FLOPs.

GenEval and DPG-Bench for caching alone and caching plus MOSAIK
≈70% FLOPs saved≈80% FLOPs saved
TeaCache 0.71 → 0.74DPG 82.6 → 82.8 0.67 → 0.73DPG 80.7 → 82.2
TaylorSeer 0.68 → 0.74DPG 82.4 → 82.8 0.43 → 0.73DPG 72.9 → 82.2

GenEval, caching alone → caching + MOSAIK.

The savings hold on hardware

At the ≈70% tier on one NVIDIA GB10, the DiT blocks run 5.40× faster; the fixed pixel decoder and text stream set a 4.4 s floor.