mudpie

Company profile · 3 min read

Hoid: An agentic compiler for model-specific inference performance

Hoid profiles and optimizes AI models for target hardware, helping inference teams navigate bottlenecks, kernels and hardware-specific performance tradeoffs.

Published · Updated

Hoid is an agentic AI compiler that profiles and optimizes models for the hardware they run on. It fits an inference team that has a real model, a target chip and a performance gap that is too expensive to close through manual kernel and compiler work.

What it does

Hoid’s current site describes a software layer between a model and its silicon. The system profiles a workload, identifies bottlenecks, benchmarks alternatives and navigates an optimization space specific to the model and hardware. The company contrasts that with fixed compiler rules and hand-tuned kernels that usually work for one stack at a time.

The product is best understood as an optimization pilot, not a universal “make AI faster” button. Hoid’s FAQ says a customer chooses a model and target hardware and agrees on the performance targets that matter. It also says the system can work with custom model optimizations and does not require an open-source model. That gives a technical buyer a concrete starting point: one workload, one hardware target, one latency or throughput target.

The homepage cites external examples of how long it can take inference providers to optimize a new model and shows a 100x improvement example for a DeepSeek model after a month of hand tuning. Those references describe the problem and comparison context; they are not a Hoid benchmark. Hoid does not publish a customer performance result on the retained page.

Founder context and tradeoffs

The Speedrun profile identifies Momčilo Mrkaić, Pavle Padjin and Vladimir Zeljkovic as founders, with backgrounds at Databricks, Tenstorrent, Cambridge and machine-learning research. That is directly relevant to compiler and hardware optimization.

Pricing is not public. The buyer should ask whether Hoid needs access to weights, traces or proprietary kernels, how optimizations are validated against model quality, what hardware is supported and how a generated implementation is maintained when the model or runtime changes.

Editorial take

I would shortlist Hoid for an inference team where utilization, latency or cost is already a board-level constraint. I would not bring it in for a model that has no stable workload or target hardware. The proof is a bounded pilot that beats the team’s current implementation on the metric that actually drives its bill or product experience.

Quick facts

Field Sourced detail
Product Agentic model compiler and hardware-specific inference optimization
Buyers Inference providers, AI product teams and accelerator users
Pilot shape One model, target hardware and agreed performance targets
Pricing Pilot/demo-led; no numeric price observed
Main gate Model access, hardware coverage, quality preservation and maintenance

Sources checked

Source Checked
Speedrun profile 2026-09-19
Hoid homepage 2026-09-19
Hoid blog 2026-09-19

About the author

I cofound Lazyweb and publish Mudpie. This is an owner-written publication, not an independent testing organization. Research notes distinguish observations, sourced reporting and editorial judgment.

First1000 ↗ · X ↗