YOUR OWN HARDWARE
NVIDIA GeForce RTX 4090
24 GB VRAM
Calculate LLM VRAM and RAM. Find hardware that fits.
Make your next local AI setup an informed one.
A tiny instruction model for experimenting on memory-limited devices.
4.85 effective bits per weight · 0.30 GB estimated weights
More room for conversation needs more memory.
One conversation, full GPU offload, FP16 KV cache.
Room for your model. Space to think.
Planning estimate, not a measured requirement.
YOUR OWN HARDWARE
24 GB VRAM
UNIFIED MEMORY
64 GB unified memory
ON-DEMAND COMPUTE
80 GB VRAM
Provider links may be affiliate links. We may earn a commission at no extra cost to you. Memory fit does not guarantee runtime compatibility or performance.
Weight-file GB are converted to GiB (2³⁰ bytes). FP16 KV memory is 2 × layers × KV heads × head dimension × context tokens × 2 bytes. Runtime allowance is the larger of 1 GiB or 10% of weights. We then leave 10% of GPU capacity free.
System RAM is weights plus 8 GiB, rounded up to 8 GiB. Unified memory is the greater of GPU target plus 8 GiB or GPU target divided by 0.75, rounded up to 8 GiB. These are conservative loading and OS allowances, not measured peaks. Mac cards use this separate unified-memory target.
Actual memory use depends on your runtime, model file, drivers, batching, and loading strategy. Catalog sizes use approximate parameters and effective bits per weight, except gpt-oss which uses rounded published MXFP4 downloads. Gemma 3 and gpt-oss cache estimates allow full context per layer, without sliding-window savings. Compare the actual file and usable memory before buying. Hardware GB are conservatively treated as decimal bytes. CPU offload, image processing and multi-GPU performance are not modeled. vLLM options support FP16 and native gpt-oss MXFP4 and exclude Macs.
Hosted AI services do the model processing remotely. Gemma and gpt-oss have downloadable weights for local use. Codex is a coding tool, and its memory needs depend on the model behind it. Compare local and hosted AI.