Another week, another release. We’re happy to announce the release of zml/llmd 20261009.0, which brings considerable performance improvements across the board.

Faster Compilation

In this release, we have spent a considerable amount of effort into improving compilation performance. Notably, when in multiple GPU configurations, auto-tunning is now done in parallel across GPUs. This has a significant impact on cold-start performance, as shown below.

This changes impact NVIDIA, AMD, Intel and Apple Metal backends.

Cold Start, deepseek-ai/DeepSeek-V4.1-Flash, 4xGB300, less is better
Cold Start, deepseek-ai/DeepSeek-V4.1-Flash, 4xGB300, less is better zml/llmd 20261009.0: 3m8s. zml/llmd 20260929.0: 5m9s. SGLang: 35m17s. vLLM: 44m57s. Bars start at zero. zml/llmd 20261009.0 3m8s zml/llmd 20260929.0 5m9s SGLang 35m17s vLLM 44m57s

And here is a visual representation of the compilation time improvements (BEFORE / AFTER):

Faster Qwen3.8-27B

In this release, we have decided to concentrate some fundamental performance work and apply them to Qwen/Qwen3.8-27B as a test ground.

We chose to benchmark it with the ridiculous bs=1024 pure decode path because we found it tends to highlight ineficiencies better.

Token Generation, Qwen/Qwen3.8-27B, w16a16, bs=1024, 4xGB300, more is better
Token Generation, Qwen/Qwen3.8-27B, w16a16, bs=1024, 4xGB300, more is better zml/llmd 20261009.0: 21500. zml/llmd 20260929.0: 19500. SGLang (high throughput, --enable-torch-compile): 13500. Bars start at zero. zml/llmd 20261009.0 21500 zml/llmd 20260929.0 19500 SGLang (high throughput, --enable-torch-compile) 13500
Apple M3 Max, nvidia/Qwen3.8-27B-NVFP4, pp1024/tg128, bs=1, more is better (tg tok/s)
Apple M3 Max, nvidia/Qwen3.8-27B-NVFP4, pp1024/tg128, bs=1, more is better zml/llmd 20261009.0: 17.7 tg tok/s. zml/llmd 20260929.0: 16.2 tg tok/s. oMLX 0.7.0: 13.4 tg tok/s. Bars start at zero. zml/llmd 20261009.0 17.7 tg tok/s zml/llmd 20260929.0 16.2 tg tok/s oMLX 0.7.0 13.4 tg tok/s
Command lines

zml/llmd

llmd --model=/var/models/Qwen/Qwen3.8-27B --batch-size=1024

SGLang

uv run sglang serve \
    --trust-remote-code \
    --model-path /var/models/Qwen/Qwen3.8-27B \
    --mem-fraction-static 0.85 \
    --reasoning-parser qwen3 \
    --tool-call-parser qwen3_coder \
    --mamba-full-memory-ratio 3.67 \
    --host 0.0.0.0 \
    --port 30000 \
    --mamba-radix-cache-strategy extra_buffer_lazy \
    --mamba-ssm-dtype float32 \
    --max-running-requests 1024 \
    --tp 4 \
    --cuda-graph-max-bs-decode 1024 \
    --enable-torch-compile

Faster DeepSeek-V4.1-Flash

As stated in the previous blog post, we chose to release DeepSeek-V4.1-Flash early, even though we were not yet at the performance level we can be. This release brings incremental performance improvements on DeepSeek-V4.1-Flash, as shown below.

Token Generation, deepseek-ai/DeepSeek-V4.1-Flash, bs=1, 4xGB300 (tok/s)
Token Generation, deepseek-ai/DeepSeek-V4.1-Flash, bs=1, 4xGB300 zml/llmd 20261009.0: 150 tok/s. zml/llmd 20260929.0: 130 tok/s. vLLM: 163 tok/s. SGLang: 200 tok/s. Bars start at zero. zml/llmd 20261009.0 150 tok/s zml/llmd 20260929.0 130 tok/s vLLM 163 tok/s SGLang 200 tok/s

Speculative decoding support on Metal

We had missed enabling speculative decoding on Metal in the previous release. This release brings it back, and it works well. DFlash, DFlash2 and DSpark are all supported. If the model natively supports it, you can also pass the --speculative-drafter-model=native flag to enable it.

ZIO is now the default IO backend

This release now enables ZIO has the default Zig std.Io backend.

ZIO is an asynchronous runtime for Zig, in the same spirit as Go’s runtime or Tokio: it schedules lightweight coroutines onto a pool of OS threads, and gives you blocking-looking network, file, and process I/O that’s actually backed by non-blocking, event-driven OS APIs under the hood.

Thanks to it, zml/llmd runs on top of kqueue on macOS, epoll or io_uring on Linux (detected at runtime).

ZIO is exactly what we wanted to build ourselves (and had started to prior to std.Io in older zml/zml versions). So huge shout out to Lukáš Lalinský. Great job.

The old std.Io.Threaded backend is still available and can be enabled with the --io-impl=threaded flag.

Try it today

NVIDIA

docker run -p 8000:8000 --shm-size=256GB --gpus=all -e HF_TOKEN -it zmlai/llmd:cuda \
    --model=hf://deepseek-ai/DeepSeek-V4.1-Flash

AMD

docker run -p 8000:8000 --shm-size=256GB --device=/dev/kfd --device=/dev/dri -e HF_TOKEN -it zmlai/llmd:rocm \
    --model=hf://deepseek-ai/DeepSeek-V4.1-Flash

Intel

docker run -p 8000:8000 --device=/dev/dri -e HF_TOKEN -it zmlai/llmd:oneapi \
    --model=hf://Qwen/Qwen3.8-27B

Google TPU

docker run --net=host --privileged -e HF_TOKEN -it zmlai/llmd:tpu \
    --model=hf://Qwen/Qwen3.8-27B

Apple Metal

brew install zml/zml/llmd
llmd --model=hf://nvidia/Qwen3.8-27B-NVFP4 --batch-size=1