Silicon
AMD MI455X vs NVIDIA Rubin: can more memory beat NVLink?
The MI455X carries 432 GB of HBM4, versus 288 GB on Rubin. AMD still has to prove that UALoE can turn that capacity advantage into rack-scale performance.
Silicon
The MI455X carries 432 GB of HBM4, versus 288 GB on Rubin. AMD still has to prove that UALoE can turn that capacity advantage into rack-scale performance.
Silicon
The DGX Spark and the RTX 5090 answer two opposite questions: hosting a very large model, or serving a mid-size one fast. The duel in numbers, measurements and mechanisms included.
Runtimes
NVIDIA Dynamo 1.0, SGLang 0.5.12, TensorRT-LLM 1.3, Tenstorrent Blackhole: prefill/decode disaggregation reaches production. Implementations, benchmarks, and counter-arguments.
Runtimes
The KV cache is no longer a VRAM byproduct you just absorb: it has its own quantization (TurboQuant, ~3 bits), its transfer protocol, its tiered storage (KVBM G1→G4) and its own routing. Anatomy of that shift.
Silicon
How much VRAM does a local LLM need? The ~2 GB per billion parameters rule, the weight of the KV cache, what quantization changes, and what actually fits on your card.
Silicon
The RTX Spark (N1X) is not a rebadged DGX Spark: it is the Grace Blackwell SoC for Windows on Arm, with 128 GB of unified memory. Why the ~273 GB/s, not the "petaflop", decide what it can actually do.
Silicon
TorchTPU, PyTorch/XLA, JAX, XLA and vLLM form a stack whose keystone is a compiler. How Google attacks the one asset that keeps developers on NVIDIA: the cost of leaving CUDA.