NC LOCAL AI / MEMORY FIT

Can Your PC Run a Local AI Model?

A transparent memory estimate for current local models. Choose a checkpoint, quantization, context length, and graphics card to see whether the weights and working memory fit the selected target.

Selection models released 12 May 2026 to 12 Aug 2026 (frozen cohort) Newest model carried 11 Aug 2026 Sources re-checked 15 Sep 2026 Register data changed 15 Sep 2026

The release selection is a frozen cohort, not a rolling window: models already checked stay in the register when their release window passes, and releases after 12 Aug 2026 are not added automatically.

9 MODELS / 8 PROFILES
Bit-width buckets, not vendor checkpoints: NVFP4, FP8 and native INT4 do not map exactly onto Q4-Q8.
INITIAL ESTIMATE

Gemma 4 12B on GeForce RTX 4060

Insufficient memory for this configuration
Estimated weights6.1 GiB
Context reserve0.4 GiB
Runtime overhead1.5 GiB
Estimated requirement8.0 GiB
Usable target memory7.2 GiB

Google describes this as a local multimodal model and publishes downloadable checkpoints.

METHOD

A memory check, not a speed promise

Methodology: The calculator estimates checkpoint weight memory from total parameter count and the selected bit depth. It adds a 10% packing margin, a model-specific working-memory allowance, and a context reserve. Mixture-of-experts models still need their total checkpoint weights available even when only some parameters are active per token.

The estimate follows the practical distinctions described in Ollama's GPU support documentation and the llama.cpp quantization notes. It does not estimate tokens per second, quality, thermals, driver compatibility, or whether a particular backend supports every quantization.

The weight precision selector uses generic bit-width buckets (Q4-Q8, FP16) for memory estimation. Several vendors publish named formats that do not map exactly to these buckets: Nemotron 3.5 Lightning ships NVFP4, Kimi K3 ships MXFP4 weights with MXFP8 activations, Kimi K2.7 Code ships native INT4, and DeepSeek-V4-Flash ships FP8. A user picking 'Q4' is modelling a bucket, not the vendor's published checkpoint.

LIMITS

What the result cannot tell you

Limitations: A passing estimate is not a compatibility guarantee; runtime, driver, backend, thermals, and context behavior can change the result.

  • A pass means the memory estimate fits the selected target after a 10% reserve. It does not guarantee that the model will load.
  • Long contexts can consume substantially more memory, and the actual KV cache depends on architecture and runtime.
  • GPU, unified-memory, CPU-offload, driver, and backend behavior differ. Check the model card and runtime documentation before downloading.
  • The model register is intentionally dated. When a release or specification changes, this page should be updated and re-verified.
USE THE RESULT

Three checks before you download

  1. Choose a model whose source entry matches the kind of workload you want to run.
  2. Set the quantization and context length you actually plan to use, because both change the memory estimate.
  3. Compare the result with the model card and runtime documentation before treating a pass as a setup guarantee.
DECISION BOUNDARY

Memory is only the first constraint

A model can fit in the estimated memory and still be a poor match for a particular backend, driver, operating system, or workload. The register links each entry to its source and keeps specialized models labelled so they are not mistaken for general chat checkpoints.

If the question is hosted model access rather than local hardware, compare provider plans, model reach, and published limits instead.

RELEASES 12 May 2026 TO 12 Aug 2026

Models

Hardware profiles