mudpie

Company profile · 4 min read

Conifer: local-first model routing for lower AI spend

Conifer routes AI requests across local hardware, cloud providers, models and caches through one gateway with cost receipts and local-only options.

Published · Updated

Conifer is a model-routing and local-inference layer for teams that want to reduce token spend without rebuilding every coding-agent or API integration. It puts local hardware, cloud models, providers, caches and a cost receipt behind one gateway, with a local-only path for sensitive work.

What it does

The current Conifer site says the runtime, router and local inference are free, and that its gateway exposes more than 200 models with no markup over the model’s own rate. The docs describe one API key and endpoint, named-model routing or an auto choice, provider failover, caching and a receipt naming the serving model and cost. It lists integrations for Claude Desktop, Codex, Cursor, VS Code and other agent surfaces.

The YC launch says Conifer begins with the customer’s own hardware, then falls back to efficient cloud or frontier models, and reports up to 80% lower paid token volume. It also says its Rust inference engine reaches up to 60% faster decode than llama.cpp on Apple Silicon. Those are company claims, not an independent cost or speed benchmark.

Why I’d look closer

The advantage is a concrete cost and privacy lever. Simple requests can stay local while harder ones use cloud models, and a team can see which model served each request. Founder context fits the systems problem: the YC profile describes Charles Muehlberger’s work on multimodal inference and Michael Jeffords’ experience with ML and clinical systems.

The tradeoff is routing quality and operational trust. The cheapest model may be wrong for a task; failover can change behavior; local models can be slower or weaker; and one gateway key becomes a sensitive control point. Local-only mode is a useful boundary, but it does not make cloud routing safe by default.

What I’d ask

How does auto choose a model, and can teams pin or deny providers? What is included in a cost receipt when caching or failover occurs? How are API keys, prompts and logs handled? Can a team compare local and cloud output on its own regression set before allowing traffic to shift?

My editorial take

Shortlist Conifer if inference cost, provider sprawl or local-data requirements are already painful. Start with named models and receipts, then trial routing on reversible coding or batch workloads before letting it steer production traffic. The appeal is not “AI is cheaper”; it is a measurable control plane for deciding where each request runs.

Quick facts

Field Sourced detail
Product Local inference, model gateway, least-cost routing, failover and caching
Buyer High-token AI teams, coding-agent users and privacy-sensitive operators
Pricing Runtime/router free; gateway says tokens bill at catalog model rate with no upcharge
Company claims Up to 70–80% lower paid token volume; up to 60% faster local decode on Apple Silicon
Main question Can routing save money without hiding quality, provider or data-boundary changes?

Sources checked

Source Checked
YC company profile 2026-09-19
Conifer homepage 2026-09-19
Conifer docs 2026-09-19
Conifer YC launch 2026-09-19

Cohort context

Conifer is listed in Summer 2026. In our 2026-09-18 directory snapshot, 119 of 232 listed companies in that cohort have YC’s primary industry label B2B (51.3%). This is a current-directory comparison, not an original intake count or a performance ranking. Nine-cohort dataset.

Public website snapshot

Observed 2026-09-19T16:18:41.405Z in raw homepage HTML. This records visible metadata and advertised links, not agent execution or product quality.

Signal Homepage observation
Product description metadata Observed
Canonical link Observed
H1 or H2 heading Observed
Typed structured data Observed
Docs/developer link Observed
Pricing link Not observed in this response
llms.txt link Not observed in this response
Markdown alternate Not observed in this response

Public observations · Collection method. Missing links here do not establish that a capability or file is absent elsewhere.

About the author

I cofound Lazyweb and publish Mudpie. This is an owner-written publication, not an independent testing organization. Research notes distinguish observations, sourced reporting and editorial judgment.

First1000 ↗ · X ↗