Gemma 3 12B Instruct VRAM and RAM requirements

Mid-sized Google open model. Estimates cover text inference; image processing needs additional memory.

Text inference estimate. KV cache is conservatively allocated at full context for every layer, without sliding-window savings. Vision encoder files and image-processing memory are not included. Runtime optimizations can reduce cache use.

This Google model has approximately 12 billion parameters. The estimates below use 8,192 context tokens, one conversation, FP16 KV cache and all model weights on the GPU.

Estimated memory for Gemma 3 12B Instruct
PrecisionGPU VRAM targetSystem RAM targetEstimated weight size
Q4_K_M12.0 GiB16 GiB7.27 GB
Q8_017.8 GiB24 GiB12.75 GB
FP1630.7 GiB32 GiB24.00 GB

How much context can I use?

This calculator supports up to 131,072 tokens for this model. Longer prompts and generated responses share that budget. KV cache grows with context length; additional conversations need additional cache memory.

Can I run Gemma 3 12B Instruct in Ollama, LM Studio or vLLM?

Q4_K_M and Q8_0 estimates represent GGUF-style quantizations commonly used with Ollama and LM Studio. The calculator supports 16-bit weight estimates for vLLM. Check the runtime version, model format and hardware support before downloading. A memory fit alone does not guarantee compatibility or speed.

What do these estimates include?

Weights are estimated from parameter count and bits per weight, not measured files. We add FP16 attention cache, at least 1 GiB runtime allowance and 10% GPU headroom. System RAM includes an 8 GiB loading allowance. CPU offload, image processing and multi-GPU setups are not modelled.

Official model architecture configuration

Compare hardware and change context in the VRAM calculator

Compare other local LLM models · Privacy and cookies