We’re happy to announce the release of zml/llmd 20260929.0, which notably brings
deepseek-ai/DeepSeek-V4.1-Flash support on NVIDIA and AMD GPUs.
Models
This release adds 2 model families:
- deepseek-ai/DeepSeek-V4.1-Flash (with native DSpark support)
- meta-models/Muse-Glimmer-30B
As stated, DeepSeek is only available on platforms with enough VRAM to host it natively. We do not plan to support custom quants in the near future.
DeepSeek-V4.1-Flash
Obviously the highlight, this is our first release supporting a frontier model. ZML is now mature enough to support complex frontier models, and we wanted to bring it to you. Bear in mind this is a 763B model, weighing 510GB, so for now we only support it on platforms LLMD supports with enough VRAM, namely NVIDIA and AMD GPUs.
That said, this is a cool release because it pushes the limits and shows how powerful ZML can be: MXFP4 MoEs, block-FP8, DSpark speculative decoding, DSA, HCA, and more. Nothing was spared. And it all works.
Sadly, we couldn’t bring it to TPUs because Google told us there were no TPUv7s available “for at least a few weeks”. LLMD doesn’t support Trainium yet due to an AWS Neuron bug that has since been fixed. We’ll revisit that.
On performance, we’re about 20% slower than vLLM and 30% slower than SGLang on 4xGB300. We could’ve waited to close the gap, but we wanted to bring it to you now. The roadmap is clear and we’ll get there, so why wait?
Speculative decoding is also supported on DeepSeek-V4.1-Flash, and it works well.
It can be enabled with the --speculative-drafter-model=native flag.
Cold start
We’re pretty happy about our cold-start performance. We already have improvements that are not in this release that we can’t wait to share with you soon.
Bringing frontier IR support to ZML
It was always possible in zml/zml to emit custom kernels as MLIR bytecode using the built-in codegen builder. This allows to leverage ZML’s memory allocation, stream management, concurrency and sharding while going off the trail if needed. We used that feature to emit Triton and Mosaic (Pallas) kernels.
Since the last release, we have added support for CuTe IR, CUDA Tile IR, and Fly IR.
Why does this matter? Well, because Triton cannot get the full performance of the underlying hardware. That’s why there’s Gluon now (which we’ll add support for soon).
On NVIDIA, CUDA Tile is faster than Triton, and CuTe is as fast as raw C++ CUDA (while still being a JIT). On AMD, Fly IR is faster than AMD’s preferred Triton path.
Using these new backends, we replaced the few custom kernels we had (MoE and PagedAttention) using their fastest respective IR. For instance, on NVIDIA, CuTe IR is as fast as C++ CUDA. So here goes.
As always with us, this was achieved with zero Python. No compromise.
Quantization support on Metal
Metal gains MXFP8, MXFP4 and NVFP4 block scaled dot products and block-FP8 MoE paths.
That means you can run NVFP4 models on Apple Silicon:
$ llmd --model=hf://nvidia/Qwen3.8-27B-NVFP4 --batch-size=1
The performance is nice, too:
Various improvements
AMD
This release ships with ROCm 10.
Intel
This release brings fixes for four GPU configurations.
Serving diagnostics
Prometheus metrics (/metrics) now cover scheduling, request outcomes, prompt-token counts, and speculative-decoding efficiency.
ignore_eos is also now supported for benchmarking.
Try it today
NVIDIA
docker run -p 8000:8000 --shm-size=256GB --gpus=all -e HF_TOKEN -it zmlai/llmd:cuda \
--model=hf://deepseek-ai/DeepSeek-V4.1-Flash
AMD
docker run -p 8000:8000 --shm-size=256GB --device=/dev/kfd --device=/dev/dri -e HF_TOKEN -it zmlai/llmd:rocm \
--model=hf://deepseek-ai/DeepSeek-V4.1-Flash
Intel
docker run -p 8000:8000 --device=/dev/dri -e HF_TOKEN -it zmlai/llmd:oneapi \
--model=hf://Qwen/Qwen3.8-27B
Google TPU
docker run --net=host --privileged -e HF_TOKEN -it zmlai/llmd:tpu \
--model=hf://Qwen/Qwen3.8-27B
Apple Metal
brew install zml/zml/llmd
llmd --model=hf://nvidia/Qwen3.8-27B-NVFP4 --batch-size=1