Another week, another release. We’re happy to announce the release of zml/llmd 20261009.0, which brings considerable
performance improvements across the board.
Faster Compilation
In this release, we have spent a considerable amount of effort into improving compilation performance. Notably, when in multiple GPU configurations, auto-tunning is now done in parallel across GPUs. This has a significant impact on cold-start performance, as shown below.
This changes impact NVIDIA, AMD, Intel and Apple Metal backends.
And here is a visual representation of the compilation time improvements (BEFORE / AFTER):
Faster Qwen3.8-27B
In this release, we have decided to concentrate some fundamental performance work and apply them to Qwen/Qwen3.8-27B as a test ground.
We chose to benchmark it with the ridiculous bs=1024 pure decode path because we found it tends to highlight
ineficiencies better.
Command lines
zml/llmd
llmd --model=/var/models/Qwen/Qwen3.8-27B --batch-size=1024
SGLang
uv run sglang serve \
--trust-remote-code \
--model-path /var/models/Qwen/Qwen3.8-27B \
--mem-fraction-static 0.85 \
--reasoning-parser qwen3 \
--tool-call-parser qwen3_coder \
--mamba-full-memory-ratio 3.67 \
--host 0.0.0.0 \
--port 30000 \
--mamba-radix-cache-strategy extra_buffer_lazy \
--mamba-ssm-dtype float32 \
--max-running-requests 1024 \
--tp 4 \
--cuda-graph-max-bs-decode 1024 \
--enable-torch-compile
Faster DeepSeek-V4.1-Flash
As stated in the previous blog post, we chose to release DeepSeek-V4.1-Flash early, even though we were not yet at the performance level we can be. This release brings incremental performance improvements on DeepSeek-V4.1-Flash, as shown below.
Speculative decoding support on Metal
We had missed enabling speculative decoding on Metal in the previous release. This release brings it back, and it works
well. DFlash, DFlash2 and DSpark are all supported. If the model natively supports it, you can also pass the
--speculative-drafter-model=native flag to enable it.
ZIO is now the default IO backend
This release now enables ZIO has the default Zig std.Io backend.
ZIO is an asynchronous runtime for Zig, in the same spirit as Go’s runtime or Tokio: it schedules lightweight coroutines onto a pool of OS threads, and gives you blocking-looking network, file, and process I/O that’s actually backed by non-blocking, event-driven OS APIs under the hood.
Thanks to it, zml/llmd runs on top of kqueue on macOS, epoll or io_uring on Linux (detected at runtime).
ZIO is exactly what we wanted to build ourselves (and had started to prior to std.Io in older zml/zml versions).
So huge shout out to Lukáš Lalinský. Great job.
The old std.Io.Threaded backend is still available and can be enabled with the --io-impl=threaded flag.
Try it today
NVIDIA
docker run -p 8000:8000 --shm-size=256GB --gpus=all -e HF_TOKEN -it zmlai/llmd:cuda \
--model=hf://deepseek-ai/DeepSeek-V4.1-Flash
AMD
docker run -p 8000:8000 --shm-size=256GB --device=/dev/kfd --device=/dev/dri -e HF_TOKEN -it zmlai/llmd:rocm \
--model=hf://deepseek-ai/DeepSeek-V4.1-Flash
Intel
docker run -p 8000:8000 --device=/dev/dri -e HF_TOKEN -it zmlai/llmd:oneapi \
--model=hf://Qwen/Qwen3.8-27B
Google TPU
docker run --net=host --privileged -e HF_TOKEN -it zmlai/llmd:tpu \
--model=hf://Qwen/Qwen3.8-27B
Apple Metal
brew install zml/zml/llmd
llmd --model=hf://nvidia/Qwen3.8-27B-NVFP4 --batch-size=1