Skip to content
Atlas Blog
LatestEngineeringBenchmarksDesign
atlasinference.io
blog.atlasinference.io

Notes from the inference layer

Kernel work, measured benchmarks, and what it takes to run frontier models on hardware you own. Everything we publish is reproducible from a commit.

All EngineeringBenchmarksReleasesDesign
Design Featured

Atlas is the inference engine from desktop to hyperscaler

Speed. Security. Governance. Atlas Inference from desktop to hyperscaler. Owners of dedicated systems and renters who deliver by API both need more inference per watt.

Alexi Derkatsch Sep 11, 2026 2 min read
  • Sep 1, 2026

    DFLASH-2: the fastest single-machine numbers Atlas has produced

    66.6 tokens per second on a stock build, one DGX Spark, one stream, and every figure reproducible from a commit.

    Engineering 4 min · RS
  • Aug 31, 2026

    Seven Tenets Powering Atlas Inference Accelerated Workloads

    Atlas Inference is a free and open source LLM inference engine written from scratch in Rust. These are the seven philosophical tenets we started it on, and why we left the Python vLLM stack to do it.

    Engineering 7 min · TB
Atlas Inference Engine

Zero-trust inference on hardware you own. Pure Rust and CUDA, built in North Carolina.

Blog

  • Latest
  • Engineering
  • Benchmarks
  • RSS feed

Atlas

  • atlasinference.io
  • Documentation
  • Benchmarks
  • Download

Community

  • GitHub
  • Discord
  • X
© 2026 Atlas Inference · Community Edition AGPLv3
blog.atlasinference.io