ReviewedOfficial evidence

Gemma 4 12B IT

Multimodal Capabilities and Official Run Path

A source-based profile of Google DeepMind’s Gemma 4 12B instruction-tuned model, covering its encoder-free multimodal design, 256K context, Transformers run path, and documented evidence boundaries.

Page typeModel profileEditorial briefing
Evidence basisOfficial documentationLinked upstream
Deployment coverageIncludedRequirements and runtime paths
Independent testingNot performedClearly disclosed
Model briefing
Evidence note

This page organizes first-party documentation and editorial analysis. ModelRun Lab has not independently reproduced the published performance or hardware claims.

Model overview

Gemma 4 12B IT is Google DeepMind’s instruction-tuned 12-billion-parameter unified model in the Gemma 4 family. Its official model card describes native handling of text, images, audio, and video inputs with text output.

The 12B model uses an encoder-free design: image patches and audio waveforms are projected directly into the decoder-only transformer through lightweight linear layers instead of separate vision and audio encoders. ModelRun Lab has not independently run the checkpoint or reproduced the provider’s benchmark results.

Published specifications

Google lists the following specifications for the 12B Unified model:

  • 11.95 billion parameters
  • 48 layers
  • 1,024-token sliding window
  • 256K-token context window
  • 262K vocabulary
  • Text, image, audio, and video input support
  • Text generation output
  • Configurable thinking mode
  • Native system-prompt and function-calling support

These are provider-published specifications. They should not be read as independent performance verification.

Official Transformers run path

The model card documents a Hugging Face Transformers workflow using AutoProcessor and AutoModelForMultimodalLM. The listed base dependencies are Transformers, PyTorch, and Accelerate, with torchvision and librosa added for relevant multimodal inputs.

The official examples use device_map="auto" and dtype="auto". Those settings let the runtime choose placement and precision; they do not establish a guaranteed minimum GPU or memory requirement. Use the current model card and current Transformers documentation when implementing the run path because library interfaces can change.

Modality handling

Google’s examples document:

  • Audio input for transcription and speech translation
  • Image input for visual question answering
  • Video input through sequences of frames
  • Mixed multimodal prompts processed with the model’s chat template

The provider recommends placing image content before prompt text and audio content after prompt text. It also documents variable image token budgets of 70, 140, 280, 560, and 1,120, trading additional compute for more visual detail.

Thinking and prompting

The official card describes configurable thinking behavior and standard system, assistant, and user roles. Applications should use the current chat template rather than manually reconstructing control tokens. The provider also recommends temperature 1.0, top-p 0.95, and top-k 64 as general sampling defaults. These are provider recommendations, not universal settings for every workload.

Hardware boundary

The reviewed official model card does not publish one universal consumer-GPU VRAM minimum for Gemma 4 12B IT. Actual memory use depends on precision, quantization, context length, multimodal token count, KV cache, batch size, and runtime overhead.

For that reason, this page does not label a particular GPU as an official minimum and does not claim that ModelRun Lab has tested consumer hardware.

License and access boundary

The Hugging Face metadata labels the repository Apache 2.0 and links to Google’s Gemma 4 license page. This page reports that metadata but does not interpret commercial-use rights or provide legal advice. Users should review the current linked terms, repository files, dependency licenses, and usage restrictions before deployment.

Evidence boundaries

  • Architecture, modalities, context length, prompting controls, and example code are taken from Google’s official model card.
  • Benchmark values are provider-reported and are not reproduced here as independent findings.
  • No universal hardware minimum is asserted.
  • ModelRun Lab has not independently tested the model.

Primary sources