Company profile · 4 min read
Conifer: local-first model routing for lower AI spend
Conifer routes AI requests across local hardware, cloud providers, models and caches through one gateway with cost receipts and local-only options.
Published · Updated
Conifer is a model-routing and local-inference layer for teams that want to reduce token spend without rebuilding every coding-agent or API integration. It puts local hardware, cloud models, providers, caches and a cost receipt behind one gateway, with a local-only path for sensitive work.
What it does
The current Conifer site says the runtime, router and local inference are free, and that its gateway exposes more than 200 models with no markup over the model’s own rate. The docs describe one API key and endpoint, named-model routing or an auto choice, provider failover, caching and a receipt naming the serving model and cost. It lists integrations for Claude Desktop, Codex, Cursor, VS Code and other agent surfaces.
The YC launch says Conifer begins with the customer’s own hardware, then falls back to efficient cloud or frontier models, and reports up to 80% lower paid token volume. It also says its Rust inference engine reaches up to 60% faster decode than llama.cpp on Apple Silicon. Those are company claims, not an independent cost or speed benchmark.
Why I’d look closer
The advantage is a concrete cost and privacy lever. Simple requests can stay local while harder ones use cloud models, and a team can see which model served each request. Founder context fits the systems problem: the YC profile describes Charles Muehlberger’s work on multimodal inference and Michael Jeffords’ experience with ML and clinical systems.
The tradeoff is routing quality and operational trust. The cheapest model may be wrong for a task; failover can change behavior; local models can be slower or weaker; and one gateway key becomes a sensitive control point. Local-only mode is a useful boundary, but it does not make cloud routing safe by default.
What I’d ask
How does auto choose a model, and can teams pin or deny providers? What is included in a cost receipt when caching or failover occurs? How are API keys, prompts and logs handled? Can a team compare local and cloud output on its own regression set before allowing traffic to shift?
My editorial take
Shortlist Conifer if inference cost, provider sprawl or local-data requirements are already painful. Start with named models and receipts, then trial routing on reversible coding or batch workloads before letting it steer production traffic. The appeal is not “AI is cheaper”; it is a measurable control plane for deciding where each request runs.
Quick facts
| Field | Sourced detail |
|---|---|
| Product | Local inference, model gateway, least-cost routing, failover and caching |
| Buyer | High-token AI teams, coding-agent users and privacy-sensitive operators |
| Pricing | Runtime/router free; gateway says tokens bill at catalog model rate with no upcharge |
| Company claims | Up to 70–80% lower paid token volume; up to 60% faster local decode on Apple Silicon |
| Main question | Can routing save money without hiding quality, provider or data-boundary changes? |
Sources checked
| Source | Checked |
|---|---|
| YC company profile | 2026-09-19 |
| Conifer homepage | 2026-09-19 |
| Conifer docs | 2026-09-19 |
| Conifer YC launch | 2026-09-19 |
Cohort context
Conifer is listed in Summer 2026. In our 2026-09-18 directory snapshot, 119 of 232 listed companies in that cohort have YC’s primary industry label B2B (51.3%). This is a current-directory comparison, not an original intake count or a performance ranking. Nine-cohort dataset.
Public website snapshot
Observed 2026-09-19T16:18:41.405Z in raw homepage HTML. This records visible metadata and advertised links, not agent execution or product quality.
| Signal | Homepage observation |
|---|---|
| Product description metadata | Observed |
| Canonical link | Observed |
| H1 or H2 heading | Observed |
| Typed structured data | Observed |
| Docs/developer link | Observed |
| Pricing link | Not observed in this response |
| llms.txt link | Not observed in this response |
| Markdown alternate | Not observed in this response |
Public observations · Collection method. Missing links here do not establish that a capability or file is absent elsewhere.
