Open the LLM VRAM calculator

Local LLM memory requirements

Explore estimated GPU VRAM and system RAM for each model. Compare supported weight formats at an 8K context, or the model's lower native limit. These are planning estimates for one conversation with full GPU offload.

Looking for ChatGPT, Claude, Gemini or Codex? Read our local vs hosted AI guide.

Longer context, concurrent users and runtime settings can increase memory use. Check the actual model download before buying hardware.