Archive

Performance

ZML is a production inference stack, purpose-built to decouple AI workloads from proprietary hardware.

01

engineering

Laguna S 2.1

Poolside's open-weight model for agentic coding and long-horizon work.

Jul 21, 2026 inference / performance
02

engineering

ZML/LLMD alpha

The universal LLM server.

Jul 8, 2026 inference / performance
03

engineering

10x faster tokenization

Integrating with the IREE tokenizer for a 10x uplift in tokenization performance.

Apr 7, 2026 monitoring / iree
04

engineering

Introducing zml-smi

A universal diagnostic and monitoring tool for GPUs, TPUs and NPUs.

Mar 30, 2026 monitoring / diagnostics
05

engineering

Introducing ZML/v2

Frontier performance through composability.

Mar 24, 2026 compilers / inference