ZML launches free inference server for multi-chip LLM deployment

ZML, a French AI startup backed by Turing Award winner Yann LeCun, has released ZML/LLMD, a free inference-performance server that enables open-source large language models to run across diverse chips including Nvidia, AMD, Google TPU, Apple Metal, and Intel Arc.

The goal is to break vendor lock-in and give enterprises and clouds the flexibility to use a mix of chips that may be less costly or more energy-efficient.

ZML founder Steeve Morin told TechCrunch that optimizing inference has become more important than training as AI integrates into daily life, but software and architecture barriers still cause fragmentation.

Unlike ZML‘s earlier open-source ML framework released in 2024, ZML/LLMD is not open source but is launching as a free product to measure usage before deciding on a monetization model.

Morin emphasized the lean team of 20 people and $20 million raised from investors including 20VC, >commit, AALVC, Drysdale Ventures, Kima Ventures, Kindred Capital, LocalGlobe, and Puzzle Ventures, as well as backing from founders like Solomon Hykes, Clément Delangue, and Julien Chaumond.

The software could particularly help emerging European AI chipmakers like Axelera, Fractile, Kalray, OLIX, Q.ANT, SiPearl, SpiNNcloud, and VSORA.

ZML competes in the inference space with Baseten ($13B valuation), Inferact (from creators of vLLM), and RadixArk (behind SGLang), but Morin claims a broader scope that includes co-designing silicon.

He credits Paris as the ideal location for building ZML.

Hot French startup ZML releases free product to speed inference across lots of AI chips | TechCrunch

View Original