Silicon
Threadripper Halo Station vs DGX Station: what the memory specs mean
Your model has outgrown one GPU. AMD and NVIDIA offer different ways to make room, but total memory alone will not tell you which system fits your work.
GPUs & Silicon
GPUs, HBM memory, NVLink interconnects and microarchitecture: what really sets the performance and cost of AI compute, from the die to the datacenter.
12 articles
Silicon
Your model has outgrown one GPU. AMD and NVIDIA offer different ways to make room, but total memory alone will not tell you which system fits your work.
Silicon
Ryzen AI Halo and DGX Spark both offer 128 GB, but performance changes with model and protocol. Vendor claims and independent LLM tests, audited.
Silicon
BlueField-4 and Pensando Salina compute no tokens. They move networking, storage, security and selected KV-cache transfers away from the host CPU. Here is the data path that can free the GPU, or merely move the bottleneck.
Silicon
UALink 2.0 standardizes compute inside the network; NVLink 6 integrates it into its switches. Here is what the mechanism changes, and why Helios and Vera Rubin still cannot be ranked.
Silicon
Apple launched the M3 Ultra Mac Studio with 256 GB and 512 GB options. Only 96 GB remains. The chip did not get slower, but an entire class of local models has left the store.
Silicon
An RTX PRO 6000 and three RTX 5090s each add up to 96 GB of GDDR7. For an LLM, that is where the equality ends: one pool, one sharded model and three independent replicas pay different costs.
Silicon
RTX Spark brings up to 128 GB of unified memory and native CUDA to Windows on Arm. As of August 12, CUDA remains a developer preview, TensorRT-RTX is ready, and the LLM serving stack still needs qualification.
Silicon
The MI455X carries 432 GB of HBM4, versus 288 GB on Rubin. AMD still has to prove that UALoE can turn that capacity advantage into rack-scale performance.
Silicon
The DGX Spark and the RTX 5090 answer two opposite questions: hosting a very large model, or serving a mid-size one fast. The duel in numbers, measurements and mechanisms included.
Silicon
How much VRAM does a local LLM need? The ~2 GB per billion parameters rule, the weight of the KV cache, what quantization changes, and what actually fits on your card.
Silicon
Windows on Arm or Linux, up to 128 GB of memory: what NVIDIA confirms, what DGX Spark measurements tell us, and what remains unknown about RTX Spark.
Silicon
TorchTPU, PyTorch/XLA, JAX, XLA and vLLM form a stack whose keystone is a compiler. How Google attacks the one asset that keeps developers on NVIDIA: the cost of leaving CUDA.