Silicon
How much VRAM do you need to run an LLM locally?
How much VRAM does a local LLM need? The ~2 GB per billion parameters rule, the weight of the KV cache, what quantization changes, and what actually fits on your card.
Deep dives
Every LeCompute article available in English: silicon, inference runtimes, edge AI, costs and system observability.
27 articles in English , page 3 of 3
Silicon
How much VRAM does a local LLM need? The ~2 GB per billion parameters rule, the weight of the KV cache, what quantization changes, and what actually fits on your card.
Silicon
Windows on Arm or Linux, up to 128 GB of memory: what NVIDIA confirms, what DGX Spark measurements tell us, and what remains unknown about RTX Spark.
Silicon
TorchTPU, PyTorch/XLA, JAX, XLA and vLLM form a stack whose keystone is a compiler. How Google attacks the one asset that keeps developers on NVIDIA: the cost of leaving CUDA.