Papers
arxiv:2609.38597

PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation

Published on Sep 29
Β· Submitted by
Cong Wei
on Oct 1
Authors:
,
,
,
,
,
,
,
,
,

Abstract

Unified Multimodal Models (UMMs) often rely on separate visual representations for understanding and generation, increasing visual context length and complicating integration with established vision-language pretraining pipelines. Recent advances in pixel-space modeling offer an encoder-free alternative, but extending this paradigm from images to videos is non-trivial: video understanding and generation adopt different temporal representations, leaving the design of a unified visual interface an open question. We present PixelUMM, an encoder-free model for unified image and video understanding and generation directly in pixel space. PixelUMM represents images as spatial patches and videos as spatiotemporal tubelets, connecting raw pixels to a shared multimodal backbone through single-layer linear projections. Its Mixture-of-Transformers architecture combines shared attention with task-specific parameters and extends clean-pixel prediction to video generation, jointly supporting autoregressive text prediction and pixel-space flow matching. Experiments show that PixelUMM achieves competitive performance across image and video understanding and generation tasks. We further conduct empirical studies of key design choices, including decoder design and spatial-temporal patch size, providing insights for future pixel-space unified multimodal models.

Community

Paper author Paper submitter
β€’
edited about 6 hours ago

Hi everyone! πŸ‘‹ Author here.
PixelUMM is an encoder-free unified multimodal model for image and video understanding and generation, working directly in pixel space: no VAE, no ViT.

  • 🧩 Built on Qwen3-8B as a Mixture-of-Transformers: separate understanding and generation experts share one self-attention over text, clean pixels, and noisy pixels.
  • 🎞️ Images become 16Γ—16 patches and videos become 4-frame tubes, each fed in by a single linear layer.
  • πŸ“Š Performance is comparable to open-source baselines across image/video understanding and generation.
    🌐 Project page: https://nv-tlabs.github.io/PixelUMM
    πŸ’» Code: https://github.com/nv-tlabs/PixelUMM
    πŸ€— Model: https://huggingface.co/nvidia/PixelUMM
    Code and weights are fully open. Happy to answer any questions!

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.38597
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.38597 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.38597 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.38597 in a Space README.md to link it from this page.

Collections including this paper 2