About
The French publication on AI infrastructure
LeCompute is the French publication covering AI infrastructure: GPUs, kernels, runtimes, edge AI and the cost of inference, explained from the silicon to the code.
Our angle
AI compute is commented on abundantly, but rarely explained. Vendor "peak" numbers travel faster than sustained measurements; comparisons confuse raw throughput with real cost; and the software layer (runtimes, kernels, quantization) stays a blind spot of the tech press.
LeCompute takes the problem from the bottom up, from the silicon to the code. We dissect what actually drives the performance and the cost of an inference service: memory architecture, KV cache, interconnects, numeric formats, scheduling, with reproducible measurements and no hype.
What we cover
Five topic clusters, built as a reference library rather than a news feed:
- Silicon, GPUs, HBM memory, NVLink interconnects and microarchitecture: what really sets the performance and cost of AI compute, from the die to the datacenter.
- Runtimes, vLLM, llama.cpp, TensorRT-LLM, KV cache and quantization: how to serve language models efficiently, from the datacenter to a local machine.
- Edge AI, Jetson, NPUs, embedded accelerators: deploying models under thermal, memory and power constraints, far from the cloud.
- Costs, Cost per token, GPU cloud pricing, H100/H200/B200 rental, the cloud / local / edge trade-off: the real economics of AI inference, with the numbers.
- Kernel & Perf, eBPF, perf, traces, NUMA, scheduling: diagnosing and optimizing an inference stack at the system level, where the abstractions stop.
Our editorial principles
- Measurements before marketing. We report sustained throughput, not theoretical peaks.
- Source what can be sourced. Every article cites its method and its references.
- Explain the why. A benchmark without a mental model teaches nothing.
- No hollow content. If a subject does not deserve a solid page, it gets no page.
Who writes
The analyses are signed:
- Killian Pluenet, C/C++ developer, embedded systems and low-level programming. Reads AI compute where it is actually decided: memory, silicon, performance. Projects and profile at killianpluenet.com.
- Christophe Gerardin, Networking and infrastructure for AI compute: architecting a platform, serving models, keeping the whole thing solid at scale. A background in industrial-systems cybersecurity, applied here to networks and infra.
A French publication
LeCompute publishes in French. A selection of the deep dives is adapted into English, and that selection is what you find in this section. The full catalogue, along with the glossary, lives on the French site.
Contact
A correction, a dataset to share, a subject to suggest? Write to redaction@lecompute.fr. The RSS feed tracks new English publications.