ZML is an AI engineering lab building the fastest, easiest and most universal inference software on the planet.
Starting with Python-free inference
technology that decouples AI models from the underlying
hardware, at peak performance or faster.
This is made possible by embracing the semantic pareto:
abstract the 80% that should be, verticalize on the
20% that shouldn't.
Python runtimes? Non, merci.
Hidden state? Non, merci.
Abstractions overhead? Non, merci.
Explicit over implicit.
Composability over systems.
Predicatibility over magic.
We build inference from model to the metal because this is where performance lies.