Native inference runtime for generative AI

Every image model. Faster. On the GPUs you already have.

engine reproduces each model exactly the way its authors built it, then runs it faster than the software it ships with. No retraining, no new hardware, no quality traded away.

A bioluminescent whale between rain-soaked towers, generated by engine on Qwen-Image 2.1
Generated by engine · Qwen-Image 2.1
435 / 436head-to-head comparisons won on current image models. One tie. Zero slower.
3.86×median speedup from a cold start: process launch to first image.
1.70×median speedup on a warm start, caches already on disk.
7 · 111current model families and measured routes (text-to-image, editing, control, LoRA, quantised…).

Measured on one NVIDIA RTX PRO 6000 Blackwell against the best completed stack on the same machine (e.g. Diffusers, ComfyUI, stable-diffusion.cpp, vendor reference code). A speed result only counts when the image passes the same quality gate as the reference.

The problem

The AI bill is inference.
Most of it is software waste.

01 · FRAGMENTED

Every model ships its own stack

Each release arrives with its own Python environment, loaders and quirks. Teams run a different stack per model, and each one is tuned by nobody.

02 · SLOW TO START

Cold starts burn GPU time

Loading and compiling a modern image model can take a minute before the first pixel. On autoscaling fleets, that is paid GPU time doing nothing.

03 · SILENTLY DIFFERENT

Same weights, different pictures

We measured two official ways to run FLUX.2 Klein 9B on the same GPU, weights, prompt and noise: as low as 25 dB apart. Most stacks never notice.

Proof, not promises

Faster at every point in a model's life.

Median speedup over the best completed same-machine competitor, across 111 routes on 7 current image-model families and 436 lifecycle comparisons.

3.86×

Cold start

New process, nothing cached: launch to first finished image.

1.70×

Warm start

New process with the model files already in the OS cache.

1.20×

Resident

Model loaded, one request at a time: pure per-image latency.

1.20×

Steady state

Back-to-back requests: sustained images per GPU-hour.

6.96×

FLUX.2 Klein 4B, cold start

From launch to the first image, against the fastest completed competitor on the same GPU.

engine9.35 s
best competitor65.1 s
FLUX.2 KleinFLUX.1 (incl. Kontext, Fill, Redux, Canny, Depth)Qwen-Image (incl. 2.1 & Edit)Z-ImageKrea 2SDXLMarigold

Correct first, then fast

We match the authors' own code. Then we beat it.

Speed is easy if you are allowed to change the picture. engine is not. Every route is first matched against the model authors' reference implementation, tensor by tensor, at matched noise. Only then does it count.

"Parity first: matched noise, reference-exact mode. Timings: the fastest lossless path." — our engineering rule

  • Bit-exact with the authors' stack where the reference is deterministic: e.g. Qwen-Image 2.1 text-to-image, FLUX.2 Klein, SDXL's SDE sampler with its seeded Brownian tree.
  • Speed results are refused unless both sides pass the same per-route quality gate (PSNR/SSIM floors set before measuring).
  • Every number is reproducible: one benchmark harness, one result file per workload, every image kept.
  • One native binary (C/CUDA) for all families, served over a simple HTTP API with warm models resident across requests.

Straight out of engine

The latest models. At engine speed.

Every image below came out of engine on one RTX PRO 6000, untouched. The time under each is the real render time per image with the model loaded, measured as it was made.

Z-Image Turbo

1024×1280 · 9 steps~2.0 s / image
An elderly man playing chess with pigeons in a Lisbon square, generated by engine on Z-Image Turbo
A rainy Hong Kong night market under a neon sign reading 夜市 NIGHT MARKET, generated by engine on Z-Image Turbo
A snow leopard on a rocky Himalayan ridge, generated by engine on Z-Image Turbo

FLUX.2 klein 9B

1152×1440 · 6 steps~2.8 s / image
A glass perfume bottle sculpted like a breaking wave, labelled ENGINE, generated by engine on FLUX.2 klein 9B
Aerial view of misty rice terraces at dawn with a farmer in red, generated by engine on FLUX.2 klein 9B
A mechanical hummingbird beside a scarlet flower, generated by engine on FLUX.2 klein 9B

Krea 2 Turbo

1024×1280 · 8 steps~3.5 s / image
Portrait of a freckled young woman in a wheat field at golden hour, generated by engine on Krea 2 Turbo
Still life of ripe figs in a cracked bowl in window light, generated by engine on Krea 2 Turbo
A brutalist concrete house cantilevered over a misty fjord, generated by engine on Krea 2 Turbo

Qwen-Image 2.1

1344×1680 · 40 steps~21 s / image
A concert poster reading ENGINE LIVE with crisp lettering, generated by engine on Qwen-Image 2.1
A rainy Tokyo bookshop with a sign reading 夜の本屋 Night Books, generated by engine on Qwen-Image 2.1
A bioluminescent whale between skyscrapers, generated by engine on Qwen-Image 2.1

Who it's for

Anyone paying for GPUs to make images.

GPU CLOUDS & INFERENCE PLATFORMS

More output per GPU you own

Faster cold starts mean denser autoscaling; faster steady state means more images per GPU-hour, on hardware already paid for.

PRODUCT TEAMS

Ship new models the week they launch

One runtime, one API for every family, so adopting the next FLUX or Qwen release is a model file, not a new stack.

ENTERPRISE & ON-PREM

Exact, auditable, self-hosted

A single native binary on your own GPUs, reproducing the authors' outputs, with every result traceable.

Investors · design partners

Build the future of inference with us.

The runtime works and the numbers are real. We're talking to investors and to design partners who want their models to run exactly as built, faster, on the GPUs they already have. Let's talk.