We’re happy to announce the release of zml/llmd 20260929.0, which notably brings deepseek-ai/DeepSeek-V4.1-Flash support on NVIDIA and AMD GPUs.

Models

This release adds 2 model families:

As stated, DeepSeek is only available on platforms with enough VRAM to host it natively. We do not plan to support custom quants in the near future.

DeepSeek-V4.1-Flash

Obviously the highlight, this is our first release supporting a frontier model. ZML is now mature enough to support complex frontier models, and we wanted to bring it to you. Bear in mind this is a 763B model, weighing 510GB, so for now we only support it on platforms LLMD supports with enough VRAM, namely NVIDIA and AMD GPUs.

That said, this is a cool release because it pushes the limits and shows how powerful ZML can be: MXFP4 MoEs, block-FP8, DSpark speculative decoding, DSA, HCA, and more. Nothing was spared. And it all works.

Sadly, we couldn’t bring it to TPUs because Google told us there were no TPUv7s available “for at least a few weeks”. LLMD doesn’t support Trainium yet due to an AWS Neuron bug that has since been fixed. We’ll revisit that.

On performance, we’re about 20% slower than vLLM and 30% slower than SGLang on 4xGB300. We could’ve waited to close the gap, but we wanted to bring it to you now. The roadmap is clear and we’ll get there, so why wait?

Token Generation, deepseek-ai/DeepSeek-V4.1-Flash, bs=1, 4xGB300 (tok/s)
Token Generation, deepseek-ai/DeepSeek-V4.1-Flash, bs=1, 4xGB300 zml/llmd: 130 tok/s. vLLM: 163 tok/s. SGLang: 200 tok/s. Bars start at zero. zml/llmd 130 tok/s vLLM 163 tok/s SGLang 200 tok/s

Speculative decoding is also supported on DeepSeek-V4.1-Flash, and it works well. It can be enabled with the --speculative-drafter-model=native flag.

Cold start

We’re pretty happy about our cold-start performance. We already have improvements that are not in this release that we can’t wait to share with you soon.

Cold Start, deepseek-ai/DeepSeek-V4.1-Flash, 4xGB300
Cold Start, deepseek-ai/DeepSeek-V4.1-Flash, 4xGB300 zml/llmd: 5m9s. SGLang: 35m17s. vLLM: 44m57s. Bars start at zero. zml/llmd 5m9s SGLang 35m17s vLLM 44m57s

Bringing frontier IR support to ZML

It was always possible in zml/zml to emit custom kernels as MLIR bytecode using the built-in codegen builder. This allows to leverage ZML’s memory allocation, stream management, concurrency and sharding while going off the trail if needed. We used that feature to emit Triton and Mosaic (Pallas) kernels.

Since the last release, we have added support for CuTe IR, CUDA Tile IR, and Fly IR.

Why does this matter? Well, because Triton cannot get the full performance of the underlying hardware. That’s why there’s Gluon now (which we’ll add support for soon).

On NVIDIA, CUDA Tile is faster than Triton, and CuTe is as fast as raw C++ CUDA (while still being a JIT). On AMD, Fly IR is faster than AMD’s preferred Triton path.

Using these new backends, we replaced the few custom kernels we had (MoE and PagedAttention) using their fastest respective IR. For instance, on NVIDIA, CuTe IR is as fast as C++ CUDA. So here goes.

As always with us, this was achieved with zero Python. No compromise.

Quantization support on Metal

Metal gains MXFP8, MXFP4 and NVFP4 block scaled dot products and block-FP8 MoE paths.

That means you can run NVFP4 models on Apple Silicon:

$ llmd --model=hf://nvidia/Qwen3.8-27B-NVFP4 --batch-size=1

The performance is nice, too:

Apple M3 Max, nvidia/Qwen3.8-27B-NVFP4, 1024/128, bs=1 (tg tok/s)
Apple M3 Max, nvidia/Qwen3.8-27B-NVFP4, 1024/128, bs=1 zml/llmd: 16.2 tg tok/s. oMLX: 13.9 tg tok/s. Bars start at zero. zml/llmd 16.2 tg tok/s oMLX 13.9 tg tok/s

Various improvements

AMD

This release ships with ROCm 10.

Intel

This release brings fixes for four GPU configurations.

Serving diagnostics

Prometheus metrics (/metrics) now cover scheduling, request outcomes, prompt-token counts, and speculative-decoding efficiency.

ignore_eos is also now supported for benchmarking.

Try it today

NVIDIA

docker run -p 8000:8000 --shm-size=256GB --gpus=all -e HF_TOKEN -it zmlai/llmd:cuda \
    --model=hf://deepseek-ai/DeepSeek-V4.1-Flash

AMD

docker run -p 8000:8000 --shm-size=256GB --device=/dev/kfd --device=/dev/dri -e HF_TOKEN -it zmlai/llmd:rocm \
    --model=hf://deepseek-ai/DeepSeek-V4.1-Flash

Intel

docker run -p 8000:8000 --device=/dev/dri -e HF_TOKEN -it zmlai/llmd:oneapi \
    --model=hf://Qwen/Qwen3.8-27B

Google TPU

docker run --net=host --privileged -e HF_TOKEN -it zmlai/llmd:tpu \
    --model=hf://Qwen/Qwen3.8-27B

Apple Metal

brew install zml/zml/llmd
llmd --model=hf://nvidia/Qwen3.8-27B-NVFP4 --batch-size=1