All articles (page 3)

Silicon

How much VRAM do you need to run an LLM locally?

How much VRAM does a local LLM need? The ~2 GB per billion parameters rule, the weight of the KV cache, what quantization changes, and what actually fits on your card.

8 min read
  • VRAM
  • Local LLM
  • Quantization
  • KV cache