Search
Search the deep dives
Search every LeCompute deep dive available in English: silicon, runtimes and edge AI.
11 articles available in English.
- Silicon
AMD MI455X vs NVIDIA Rubin: can more memory beat NVLink?
The MI455X carries 432 GB of HBM4, versus 288 GB on Rubin. AMD still has to prove that UALoE can turn that capacity advantage into rack-scale performance.
- Runtimes
GGUF vs GPTQ vs AWQ: a format, two algorithms, one problem
GGUF is a file format, GPTQ and AWQ are algorithms. At 4 bits, each attacks the outlier problem differently. Anatomy of three constructions, from GPTQ's Hessian to GGUF's k-quants, and of what the choice of format decides for you.
- Costs
Stargate, Spud and subscriptions: the compute lever behind OpenAI's comeback
10 gigawatts 'secured years ahead of schedule', 0.3 gigawatt plugged in at Abilene. Behind OpenAI's comeback in AI coding, a compute bet that must be read in three states: the secured, the built, the plugged-in. And a subscription economics that agents cracked open.
- Runtimes
Nobody knows how to measure coding agents anymore
On July 8, 2026, OpenAI disavowed SWE-Bench Pro: ~30% broken tasks. It is the third reference instrument declared dead in three generations. Anatomy of a metrology crisis: noise ceiling, contamination, and a harness effect worth a whole model generation.
- Costs
GPT-5.6 and the Codex merger: anatomy of OpenAI's comeback
On July 8, 2026, OpenAI disavowed the reference benchmark for agentic coding. On the 9th, it merged Codex into the ChatGPT app and launched GPT-5.6. Behind the sequence, a comeback that turns less on intelligence than on prices, subscriptions and the harness.
- Silicon
DGX Spark vs RTX 5090: capacity or speed, the choice that decides your local LLM
The DGX Spark and the RTX 5090 answer two opposite questions: hosting a very large model, or serving a mid-size one fast. The duel in numbers, measurements and mechanisms included.
- Runtimes
Prefill/decode disaggregation reaches production: state of play, May 2026
NVIDIA Dynamo 1.0, SGLang 0.5.12, TensorRT-LLM 1.3, Tenstorrent Blackhole: prefill/decode disaggregation reaches production. Implementations, benchmarks, and counter-arguments.
- Runtimes
The KV cache is no longer a side effect: it is the center of LLM serving in 2026
The KV cache is no longer a VRAM byproduct you just absorb: it has its own quantization (TurboQuant, ~3 bits), its transfer protocol, its tiered storage (KVBM G1→G4) and its own routing. Anatomy of that shift.
- Silicon
How much VRAM do you need to run an LLM locally?
How much VRAM does a local LLM need? The ~2 GB per billion parameters rule, the weight of the KV cache, what quantization changes, and what actually fits on your card.
- Silicon
RTX Spark vs DGX Spark: two different chips, one 273 GB/s wall
The RTX Spark (N1X) is not a rebadged DGX Spark: it is the Grace Blackwell SoC for Windows on Arm, with 128 GB of unified memory. Why the ~273 GB/s, not the "petaflop", decide what it can actually do.
- Silicon
TorchTPU, XLA, JAX: how Google attacks NVIDIA's software lock-in
TorchTPU, PyTorch/XLA, JAX, XLA and vLLM form a stack whose keystone is a compiler. How Google attacks the one asset that keeps developers on NVIDIA: the cost of leaving CUDA.
No English article matches this search. Try another term, browse all deep dives, or search the full French catalogue.