Choosing 8, 12, 16 or 24 GB of GPU memory

Compare memory capacities using a consistent local AI workload, with worked VRAM and system RAM examples.

Published 16 September 2026 · HostingLinks · Editorial approach

Choose a workload before choosing a card

Write down the model, weight format and context length you intend to use. “An 8B model” is not a complete hardware requirement: 16-bit weights need much more space than compressed weights, and long conversations add attention-cache memory. Start with your actual task, such as short writing assistance or reviewing a long document, then compare memory estimates under the same settings.

Our finder accepts available memory in GiB. Product listings often say GB; 1 GiB is 1,024³ bytes while decimal GB is 1,000³ bytes. Check the capacity reported by your system and leave room for other applications. The hardware recommendation cards conservatively convert advertised GB to GiB; the table below uses the capacity you enter into the finder.

A worked Llama 3.1 8B example

Using our Q4_K_M planning weights, Ollama, one conversation and 4,096 tokens gives a GPU target of 6.7 GiB and a system RAM target of 16 GiB. At 32,768 tokens, the GPU target rises to 10.6 GiB. These are calculated allowances, not measurements.

Same model and format, different available memory
Available VRAM4,096 tokens32,768 tokens
8 GiBFits estimateExceeds capacity
12 GiBFits estimateFits estimate
16 GiBFits estimateFits estimate
24 GiBFits estimateFits estimate

The lesson is not that everyone needs 24 GB. If a smaller capacity fits your real workload with room left, a larger card may add no useful capability for that task. Conversely, a model that fits at a short context can fail when a long conversation increases the cache. Keep the system RAM requirement in view too.

Compare what an upgrade unlocks

Enter your current computer in the upgrade comparison, then try 12, 16 and 24 GiB. The new-model list uses the same runtime, context and precision throughout. Expand the RAM warning to see models whose GPU requirement fits but whose system RAM allowance does not. Increasing only GPU memory cannot resolve that warning.

Memory capacity is only one buying criterion

Two cards with the same VRAM need not generate answers at the same speed. Check runtime support, your power supply, connectors, case clearance, cooling and the specific card’s specifications. A laptop GPU is not interchangeable with a desktop model of a similar name. For used hardware, investigate condition and seller protection before committing.

Our benchmark references illustrate measured performance under explicit test conditions. They are not a forecast for every catalogue model. For Apple Silicon, use the unified-memory option instead of treating total memory as dedicated VRAM. For workloads beyond one device, this calculator does not model multi-GPU splitting or CPU offload.

Try your settings in the calculator