Engineering
Kernels, memory, and the parts of the engine that decide the number.
- Sep 11, 2026
Atlas is the inference engine from desktop to hyperscaler
Speed. Security. Governance. Atlas Inference from desktop to hyperscaler. Owners of dedicated systems and renters who deliver by API both need more inference per watt.
2 min · AD - Sep 1, 2026
DFLASH-2: the fastest single-machine numbers Atlas has produced
66.6 tokens per second on a stock build, one DGX Spark, one stream, and every figure reproducible from a commit.
4 min · RS - Aug 31, 2026
Seven Tenets Powering Atlas Inference Accelerated Workloads
Atlas Inference is a free and open source LLM inference engine written from scratch in Rust. These are the seven philosophical tenets we started it on, and why we left the Python vLLM stack to do it.
7 min · TB