PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation

1NVIDIA2University of Waterloo

Overview

PixelUMM reads and writes raw pixels with a single decoder-only Transformer — no VAE and no vision encoder. Images become 16×16 patches and videos become 4-frame tubes; a Qwen3-8B backbone with separate understanding and generation experts shares one self-attention across text, clean pixels, and noisy pixels.

Architecture overview. The answer text shown for the understanding branch is illustrative.

Empirical findings

The ablations behind PixelUMM are grouped into eight experiment families, F1–F8, with configurations numbered R01, R02, and so on. We start with patch artifacts, the most visible failure mode of a linear pixel head.

F1

Image patch size

16×16 (F1-R01) and 32×32 (F1-R02) patches differ only in the generation input projection and output head. With the same sequence-length budget, 32×32 fits four times as many images per step, yet 16×16 keeps a lower T2I training loss throughout late training.

T2I training loss for 16×16 and 32×32 image patches
T2I pixel-flow MSE: full trajectory (left) and 10K–15K zoom-in (right).

TakeawayStronger spatial compression makes image generation harder to learn.

F2

Video patch size

Four tubelet-to-token mappings: p32/t4, p32/t2, p16/t4 and p32/t1 (spatial patch / frames per token). Less aggressive spatiotemporal compression generally gives lower T2V loss, with p32/t1 the lowest. The final model uses p16/t4 to match the compression of common video VAEs such as Wan2.2.

T2V training loss for four video patch configurations
T2V MSE over 0–68K steps (left) and 40K–68K (right).

TakeawayCompression ranges from 1,024 to 4,096 pixels per video token; stronger compression makes video generation harder to learn.

F4

Pixel-space vs. VAE-space training

A 32×32 pixel-patch run with x-prediction (F4-R01) against a frozen Wan2.2 VAE with 2×2 latent patchification and v-prediction (F4-R02), at the same effective token stride. At 10K steps the VAE-space loss is about 4.3× the pixel-space loss, but because the two losses live in different spaces, this does not show that pixel space learns faster.

Pixel-space and VAE-space loss and gradient-norm curves
Left: T2I velocity MSE in pixel and VAE space. Right: pre-clip global gradient norm.

TakeawayGradient norms are nearly equal; pixel-space training shows occasional loss spikes.

F5

Model size

1.7B (F5-R01) and 8B (F5-R02) models with the same recipe and a global batch size of 256. The 8B model reaches MSE 0.060 at about 14K instead of 40K steps, and CE 0.50 at about 10K instead of 31K steps.

MSE and CE loss curves for the 1.7B and 8B models
Joint Stage 1 MSE (left) and CE (right).

TakeawayThe 8B model reaches comparable generation and text losses in roughly one third of the training steps.

F6

Compute scaling

The same 1.7B model and matched recipe on 8 GPUs (F6-R01) versus 128 GPUs (F6-R02). With more GPUs, CE reaches 1.0 at about 2K instead of 17.5K steps (≈9× fewer), while MSE reaches 0.065 at about 17K instead of 23K steps (≈1.3× fewer). These are steps to a fixed loss, not wall-clock speedups.

MSE and CE loss curves for 8 and 128 GPUs
MSE (left) and CE (right); opaque curves are EMAs of the raw measurements.

TakeawayMore GPUs speed up text (CE) convergence far more than generation (MSE) convergence.

F7

Multimodal context conditioning

Clean conditions (text, reference images, an input video) enter the understanding expert, noisy target tokens enter the generation expert, and both meet in shared self-attention. This extends SenseNova-U1’s encoder-free image conditioning to image-to-video and video editing, with no separate VAE-token stream. After 15K steps of multi-task fine-tuning on eight tasks, image understanding moves both ways (BLINK +2.37, CV-Bench +2.27, SEED-I −5.41), while all four video benchmarks improve by 0.89–3.49 points.

Diagram of interleaved text, image and video conditions routed through the understanding expert
Interleaved visual-context conditioning.
A first-frame condition followed by three frames of the generated video
Image-to-video: the clean first frame, then frames 0, 47 and 95 of the generated clip.

TakeawayClean visual conditions can be routed through the understanding expert without systematically degrading understanding.

F8

Video understanding interfaces

dense_mode samples at 4 FPS and forms 4-frame tubelets; sparse_mode samples at 1 FPS and projects frames independently. Both use the same 448² per-frame budget. At F8-R01 (Und Stage 3, 9K steps):

InterfaceMVBenchVideo-MMELongVideoBenchLVBench
sparse_mode71.1557.8959.3941.38
dense_mode71.2257.6758.7141.38
Δ+0.07−0.22−0.680.00

Takeaway4-FPS tubelet input performs comparably to 1-FPS frame input.

Understanding benchmarks

The released PixelUMM checkpoint, compared with unified models and vision-language models (VLMs). Its overall performance is comparable to the baselines. Since training data differ across models, these results cannot establish which architecture is superior or which model converges faster. Dark and light green mark the best and second-best result in each row.

Image understanding

Unified modelsVLMs
BenchmarkPixelUMM8B MoTBAGEL7B MoTTUNA7B+5BTUNA-27B+5BQwen2.5-VL7BLLaVA-OV-1.58BLLaVA-OV-28BQwen3-VL-Inst.8BNEO-ov8B
MMMU41.6755.3049.8050.7051.3055.40–69.6068.10
MMStar53.99–61.20–62.5067.7064.3070.9067.30
RWQA71.6372.8066.1067.7068.5068.1069.7071.5067.80
SEED-I70.39–74.70–77.5077.30––76.60
AI2D80.1289.2079.3079.6082.6084.2084.3085.7085.40
DocVQA90.42–––94.9095.0095.2096.1091.90
ChartQA82.9678.5085.8085.6084.1086.5085.9089.6086.20
InfoVQA63.67–––81.7078.4074.4083.10–
TextVQA78.76–––84.90–––78.50
OCRBench78.0073.3074.3079.7084.2082.9078.2089.6081.60
MME1809.542388.00––2347.00––––
GQA61.9066.4063.9065.0060.70––––
MMVP80.3385.0070.7077.3078.00––––
SEED2+64.5171.9052.7061.1070.9069.20–––
CV-Bench77.60–––80.0080.70–––
CountBench94.3082.5073.5081.7086.4088.2089.0089.80–
PixMo-Count73.03–––63.3062.2064.0062.40–
V*76.9670.2052.4059.2077.0078.0085.9085.30–
MMMU-Pro27.63–––36.3037.40–––
BLINK53.46–––56.4048.3063.5069.1062.80
MuirBench38.31–––59.60––64.4058.20

PixelUMM is evaluated with the official LMMS-Eval protocol (64,750 generations over 21 tasks).

Video understanding

Unified modelsVLMs
BenchmarkPixelUMM8B MoTShow-o21.5B+0.5BTUNA1.5B+?Lance3B MoTLLaVA-OV-28BQwen3-VL8BKeye-VL-1.58BInternVL-3.58BPLM8BLLaVA-OV-1.58B
MVBench70.5349.8054.4062.0066.2069.0056.9072.1077.1051.20
Video-MME (w/o sub.)57.3348.0049.10–71.9071.4073.0065.9060.5061.10
LongVideoBench59.6149.2049.70–66.9068.0066.0062.4059.6056.20
LVBench40.41–27.40–55.5058.0042.8046.7044.5040.10

PixelUMM is evaluated in sparse_mode with up to 96 sampled frames. Video-MME is evaluated without subtitles. Published models keep their original evaluation protocols.

Limitations

PixelUMM has weaknesses. Generating directly in pixel space removes the VAE, but PixelUMM still shares the typical failure modes of latent video diffusion models.