Articles: Silicon

Silicon

How much VRAM do you need to run an LLM locally?

How much VRAM does a local LLM need? The ~2 GB per billion parameters rule, the weight of the KV cache, what quantization changes, and what actually fits on your card.

7 min read
  • VRAM
  • Local LLM
  • Quantization
  • KV cache

Silicon

RTX Spark vs DGX Spark: two different chips, one 273 GB/s wall

The RTX Spark (N1X) is not a rebadged DGX Spark: it is the Grace Blackwell SoC for Windows on Arm, with 128 GB of unified memory. Why the ~273 GB/s, not the "petaflop", decide what it can actually do.

18 min read
  • RTX Spark
  • DGX Spark
  • Grace Blackwell
  • Unified memory