mudpie

Company profile · 3 min read

Ashr: custom models and evals for production AI

Ashr trains, benchmarks, hosts and monitors purpose-built open-weight models against a company’s own data and written quality bar.

Published · Updated

What it does

Ashr builds Manifold, a platform for training, hosting, evaluating and continuously improving purpose-built open-weight models. Its current homepage describes a loop that captures an organization’s documents, transcripts, logs and corrections, trains against a written rubric, benchmarks against a current frontier API and keeps retraining in the customer’s cloud. The YC profile also describes the earlier Ashr evals product for mimicking real user behavior and testing agents.

The fit is a company with enough production data and a repeatable AI workload to justify a specialist model, but not enough research infrastructure to build the training/evaluation loop alone. Ashr’s key promise is ownership: encode the organization’s judgment in weights and run them in an environment the customer controls.

Why I’d look closer

The public workflow is unusually disciplined. Ashr says teams define correctness, freeze a dataset, run a head-to-head benchmark and only promote a model when it clears a written bar. The docs expose a Python SDK with offline evals, server-side grading and optional production observability. That gives an evaluator a concrete first step before fine-tuning a model.

The homepage displays customer references and one company-reported example of a custom model tripling browser-agent performance on EHR workflows at more than 20 clinics. Those are company/customer claims, not an independent benchmark. The YC founder biographies describe Rohan Kulkarni as a Berkeley-trained co-founder of Ask Geri and Shreyas Kaps as building Manifold for purpose-built models.

What I’d ask

Which workloads are repeatable enough for a specialist model, how are labels and graders validated, and what leaves the customer cloud? I’d run a frozen holdout against the current frontier model, inspect latency/cost and failure slices, and compare a specialist model to a better prompt before committing to retraining.

My editorial take

Ashr is a strong fit for teams that have real production traces and a measurable model-cost or quality problem. The written benchmark gate is the best part of the pitch. The buyer should demand evidence on its own workload, not generalize from the EHR claim.

Quick facts

Field Sourced detail
Buyer fit Teams with repeatable AI workloads and proprietary production data
Product Evals, post-training, hosting, observability and continual learning
Deployment claim Customer-cloud runs and open-weight model control; company-stated
Public pricing Not exposed in the sources checked

Sources checked

Checked 2026-09-19.

Source Used for
YC company profile Product history, founders and launch context
Ashr homepage Manifold workflow, customer claims and deployment posture
Ashr Python SDK docs Eval/observability implementation surface
Ashr blog Public product and research context

Cohort context

Ashr is listed in Winter 2026. In our 2026-09-18 directory snapshot, 126 of 199 listed companies in that cohort have YC’s primary industry label B2B (63.3%). This is a current-directory comparison, not an original intake count or a performance ranking. Nine-cohort dataset.

Public website snapshot

Observed 2026-09-19T16:20:00.978Z in raw homepage HTML. This records visible metadata and advertised links, not agent execution or product quality.

Signal Homepage observation
Product description metadata Observed
Canonical link Observed
H1 or H2 heading Observed
Typed structured data Observed
Docs/developer link Observed
Pricing link Observed
llms.txt link Not observed in this response
Markdown alternate Not observed in this response

Public observations · Collection method. Missing links here do not establish that a capability or file is absent elsewhere.

About the author

I cofound Lazyweb and publish Mudpie. This is an owner-written publication, not an independent testing organization. Research notes distinguish observations, sourced reporting and editorial judgment.

First1000 ↗ · X ↗