ReviewedOfficial evidence

NVIDIA Cosmos3 Super Image-to-Video 4-Step

Deployment and Hardware Guide

An evidence-based guide to NVIDIA?s 64B distilled image-to-video model, its four-step inference paths, supported runtimes, and official multi-GPU hardware boundary.

Page typeModel profileEditorial briefing
Evidence basisOfficial documentationLinked upstream
Deployment coverageIncludedRequirements and runtime paths
Independent testingNot performedClearly disclosed
Model briefing
Evidence note

This page organizes first-party documentation and editorial analysis. ModelRun Lab has not independently reproduced the published performance or hardware claims.

What this model is

NVIDIA describes Cosmos3-Super-Image2Video-4Step as a distilled version of Cosmos3-Super-Image2Video. It accepts one or more starting images plus an optional text instruction and generates an MP4 video. The checkpoint uses a fixed four-step sampling schedule without classifier-free guidance.

This page is based on NVIDIA's published model card and linked runtime documentation. ModelRun Lab has not independently run the checkpoint or reproduced NVIDIA's performance claims.

The hardware boundary matters

This is a 64B-parameter model, not a typical consumer-GPU image-to-video checkpoint. NVIDIA states that the model requires either a multi-GPU H100/H200-class node with four to eight GPUs or a single B200-class GPU. The official card says it does not fit on a single smaller GPU.

That makes ?local? deployment possible only in the sense of running the model on infrastructure you control. It should not be interpreted as evidence that a desktop RTX card can run the official checkpoint. Layer-wise offload may reduce device-memory pressure, but NVIDIA does not publish a consumer-GPU configuration for this model.

Officially documented software paths

NVIDIA lists three supported inference routes:

  • PyTorch through the Cosmos framework
  • vLLM-Omni for an OpenAI-compatible serving endpoint
  • Hugging Face Diffusers through the modular pipeline

The documented environment is Linux on NVIDIA Ampere, Hopper, or Blackwell hardware. NVIDIA says only BF16 has been tested; FP16, FP8, and FP4 are not officially supported for this checkpoint.

vLLM-Omni deployment outline

The official multi-GPU example serves the model with HSDP and Ulysses sharding. A four-GPU setup uses four shards, while an eight-GPU setup increases the Ulysses degree and shard count. A single-GPU command is documented for a GPU with enough memory, with B200 given as the example.

The four-step schedule is part of the checkpoint. vLLM-Omni reads it from the fixed sampler configuration, so users should not treat the number of inference steps or flow shift as ordinary tuning controls. NVIDIA specifies a guidance scale of 1.0.

Diffusers deployment outline

The model card says the distilled checkpoint must use Cosmos3DistilledModularPipeline. It is not compatible with the standard Cosmos3 Omni pipeline classes. At the time of the referenced documentation, NVIDIA directs users to install Diffusers from its main branch until a release containing that pipeline is available.

The Diffusers path also requires the Cosmos guardrail component. Access to the gated NVIDIA Cosmos 1.0 Guardrail repository is required for the safety checker.

Input and output defaults

NVIDIA recommends 832 ? 480 output at a 16:9 aspect ratio and uses 189 frames as the default example. The generation result is an MP4 file. The model accepts an initial image and can use a dense natural-language prompt or an optional structured, upsampled prompt.

Prompt upsampling is not required to run the distilled checkpoint. NVIDIA documents it as an optional preparation step that may improve prompt detail, but the example upsampler depends on a separate vision-language-model API.

What the four-step claim does and does not mean

Four-step distillation reduces the number of diffusion-model evaluations. NVIDIA reports up to 25 times fewer evaluations compared with its recommended 50-step base-model setting. The latency table in the model card is explicitly derived from base-model benchmarks rather than direct end-to-end measurements of every four-step configuration.

The published estimates therefore should not be presented as independently measured generation times. Preprocessing, encoding, decoding, communication between GPUs, and output writing do not necessarily scale with the denoising-step reduction.

License and operational cautions

The checkpoint is released under OpenMDW 1.1. NVIDIA's model card describes it as available for commercial and non-commercial use, but deployments still need to follow the full license terms and the rights attached to input images and generated outputs.

NVIDIA also warns that the model is not a physics simulator and may produce temporal, motion, geometry, synchronization, and reasoning errors. Robotics, autonomous systems, and other safety-sensitive uses require separate validation and safeguards.

Primary sources