Articles: Silicon

Silicon

Ryzen AI Halo vs DGX Spark: What 128GB Doesn't Tell You

Ryzen AI Halo and DGX Spark both offer 128 GB, but performance changes with model and protocol. Vendor claims and independent LLM tests, audited.

17 min read
  • Ryzen AI Halo
  • DGX Spark
  • Ryzen AI Max+ 395
  • Local LLM

Silicon

BlueField-4 vs Pensando Salina: the DPU enters the inference data path

BlueField-4 and Pensando Salina compute no tokens. They move networking, storage, security and selected KV-cache transfers away from the host CPU. Here is the data path that can free the GPU, or merely move the bottleneck.

11 min read
  • BlueField-4
  • Pensando Salina
  • DPU
  • KV Cache

Silicon

UALink 2.0 vs NVLink 6: when the network computes all-reduce

UALink 2.0 standardizes compute inside the network; NVLink 6 integrates it into its switches. Here is what the mechanism changes, and why Helios and Vera Rubin still cannot be ranked.

14 min read
  • UALink 2.0
  • UALoE
  • NVLink 6
  • in-network compute

Silicon

RTX PRO 6000 or three RTX 5090s: 96 GB is not one memory pool

An RTX PRO 6000 and three RTX 5090s each add up to 96 GB of GDDR7. For an LLM, that is where the equality ends: one pool, one sharded model and three independent replicas pay different costs.

10 min read
  • RTX PRO 6000
  • RTX 5090
  • Local LLM
  • multi-GPU

Silicon

RTX Spark on Windows: CUDA is native, the LLM stack is not yet

RTX Spark brings up to 128 GB of unified memory and native CUDA to Windows on Arm. As of August 12, CUDA remains a developer preview, TensorRT-RTX is ready, and the LLM serving stack still needs qualification.

11 min read
  • RTX Spark
  • Windows on Arm
  • CUDA 13.4
  • TensorRT-RTX

Silicon

How much VRAM do you need to run an LLM locally?

How much VRAM does a local LLM need? The ~2 GB per billion parameters rule, the weight of the KV cache, what quantization changes, and what actually fits on your card.

8 min read
  • VRAM
  • Local LLM
  • Quantization
  • KV cache