Company profile · 3 min read
Ashr: custom models and evals for production AI
Ashr trains, benchmarks, hosts and monitors purpose-built open-weight models against a company’s own data and written quality bar.
Published · Updated
What it does
Ashr builds Manifold, a platform for training, hosting, evaluating and continuously improving purpose-built open-weight models. Its current homepage describes a loop that captures an organization’s documents, transcripts, logs and corrections, trains against a written rubric, benchmarks against a current frontier API and keeps retraining in the customer’s cloud. The YC profile also describes the earlier Ashr evals product for mimicking real user behavior and testing agents.
The fit is a company with enough production data and a repeatable AI workload to justify a specialist model, but not enough research infrastructure to build the training/evaluation loop alone. Ashr’s key promise is ownership: encode the organization’s judgment in weights and run them in an environment the customer controls.
Why I’d look closer
The public workflow is unusually disciplined. Ashr says teams define correctness, freeze a dataset, run a head-to-head benchmark and only promote a model when it clears a written bar. The docs expose a Python SDK with offline evals, server-side grading and optional production observability. That gives an evaluator a concrete first step before fine-tuning a model.
The homepage displays customer references and one company-reported example of a custom model tripling browser-agent performance on EHR workflows at more than 20 clinics. Those are company/customer claims, not an independent benchmark. The YC founder biographies describe Rohan Kulkarni as a Berkeley-trained co-founder of Ask Geri and Shreyas Kaps as building Manifold for purpose-built models.
What I’d ask
Which workloads are repeatable enough for a specialist model, how are labels and graders validated, and what leaves the customer cloud? I’d run a frozen holdout against the current frontier model, inspect latency/cost and failure slices, and compare a specialist model to a better prompt before committing to retraining.
My editorial take
Ashr is a strong fit for teams that have real production traces and a measurable model-cost or quality problem. The written benchmark gate is the best part of the pitch. The buyer should demand evidence on its own workload, not generalize from the EHR claim.
Quick facts
| Field | Sourced detail |
|---|---|
| Buyer fit | Teams with repeatable AI workloads and proprietary production data |
| Product | Evals, post-training, hosting, observability and continual learning |
| Deployment claim | Customer-cloud runs and open-weight model control; company-stated |
| Public pricing | Not exposed in the sources checked |
Sources checked
Checked 2026-09-19.
| Source | Used for |
|---|---|
| YC company profile | Product history, founders and launch context |
| Ashr homepage | Manifold workflow, customer claims and deployment posture |
| Ashr Python SDK docs | Eval/observability implementation surface |
| Ashr blog | Public product and research context |
Cohort context
Ashr is listed in Winter 2026. In our 2026-09-18 directory snapshot, 126 of 199 listed companies in that cohort have YC’s primary industry label B2B (63.3%). This is a current-directory comparison, not an original intake count or a performance ranking. Nine-cohort dataset.
Public website snapshot
Observed 2026-09-19T16:20:00.978Z in raw homepage HTML. This records visible metadata and advertised links, not agent execution or product quality.
| Signal | Homepage observation |
|---|---|
| Product description metadata | Observed |
| Canonical link | Observed |
| H1 or H2 heading | Observed |
| Typed structured data | Observed |
| Docs/developer link | Observed |
| Pricing link | Observed |
| llms.txt link | Not observed in this response |
| Markdown alternate | Not observed in this response |
Public observations · Collection method. Missing links here do not establish that a capability or file is absent elsewhere.
