Key Takeaways
- ZML has released ZML/LLMD, a free inference server designed to help open source large language models run on a range of chips.
- The software is meant to reduce hardware silos and improve inference performance across systems from Nvidia, AMD, Google TPU, Apple Metal and Intel Arc.
- The launch puts ZML in a growing inference software market where cost, efficiency and chip flexibility are becoming major concerns.
What happened
ZML, a Paris-based AI startup, has launched ZML/LLMD, a new inference server that the company says can help a variety of open source large language models run on several different chip families. According to the company, that list includes Nvidia GPUs, AMD hardware, Google’s TPU, Apple Metal and Intel Arc.
Founder Steeve Morin described the product to TechCrunch as part of a broader effort to break down the silos that often shape AI infrastructure. In his view, the current stack is still constrained by software and architectural barriers that make it harder to move workloads between vendors or get the best possible performance from a given system.
The focus here is inference rather than training. Inference is the stage where a model processes prompts and generates outputs, and Morin said that this layer has become increasingly important as AI becomes more embedded in everyday products and enterprise workflows. The challenge, he argued, is that the systems behind inference are often fragmented and expensive to operate.
ZML/LLMD is not the company’s first public project. The startup previously released an inference-focused machine learning framework in 2024 and updated it in March. But ZML says the new server is different in both purpose and distribution: it is launching as a free product, though it is not open source.
Morin said the decision to make it free is meant to help the team learn how people use it before deciding how to monetize it later. He framed that as a practical approach to growth, rather than trying to force revenue too early.
The company’s pitch is that the software could help enterprises and cloud operators mix and match chips, including hardware that may be cheaper or more energy efficient for certain jobs. ZML is presenting that flexibility as a way to improve efficiency while giving customers more control over their own systems.
Why it matters
The launch lands at a moment when AI infrastructure is increasingly shaped by the economics of inference. While model training often gets the attention, the ongoing cost of serving models to users can quickly become a major operational issue. ZML is positioning its software around that problem.
That matters because the market for inference tooling is becoming crowded, but also more strategically important. TechCrunch notes competition from companies such as Baseten, Inferact and RadixArk, all of which are working in or around the same part of the stack. ZML’s wager is that there is room for a tool that spans a broader range of hardware while still pushing for performance.

The chip angle is especially notable. Nvidia remains dominant, and Morin was careful not to frame ZML as hostile to the company. He told TechCrunch that ZML has a good relationship with Nvidia and sees the chip maker as part of the inference shift rather than simply an obstacle. Still, ZML is clearly trying to make more room for alternatives.
That includes newer chipmakers, several of which are based in Europe. Morin cited companies such as Axelera, Fractile, Kalray, OLIX, Q.ANT, SiPearl, SpiNNcloud and VSORA as examples of hardware vendors that could benefit from software able to work across architectures. The larger point, he said, is not just regional pride but the chance to do things that have not been done before across this mix of chips.
There is also a broader business signal here. ZML has raised $20 million from investors including 20VC, >commit, AALVC, Drysdale Ventures, Kima Ventures, Kindred Capital, LocalGlobe and Puzzle Ventures. That funding, combined with a 20-person team, gives the startup enough runway to release products without needing to rush directly into a paid model.
What to watch
The biggest question is whether ZML/LLMD can prove that it really delivers on cross-chip performance in real deployments. ZML’s claims are ambitious because they involve not just compatibility, but doing so at or near maximum available speed on a range of hardware.
Adoption will be another test. Since the software is free but not open source, users may try it without necessarily committing to a long-term commercial relationship. That could still be valuable for ZML if it helps the company learn which customers care most about its approach and what features matter most.
It is also unclear when the product could become paid, or what pricing would look like if and when that happens. For now, the company appears more interested in expanding usage and gathering data than in locking in a business model.
More broadly, ZML’s release reflects a larger shift in AI infrastructure: the market is no longer just about building models, but about making them cheaper and more flexible to run. If ZML can make that case convincingly, it may help define a slice of the inference stack that is becoming more valuable by the month.
The company’s backing and its network of well-known founders and researchers suggest it is attracting attention, but the real proof will come from whether developers and enterprises actually adopt the software. For now, ZML is making a clear bet that performance, flexibility and cost control can be packaged together — across many chips, not just one dominant platform.



