Company profile · 4 min read
Belvedir: a private-model improvement loop for AI teams
Belvedir collects traces, trains custom models, benchmarks them, routes traffic and deploys private AI systems for teams with recurring agent workloads.
Published · Updated
Belvedir is building a private-model factory for teams that want their own AI behavior, data boundary and improvement loop without staffing a full training and inference platform. The product starts with traces from existing agents, turns them into training data and environments, then trains, benchmarks, deploys and improves a model the customer owns.
What it does
The current Belvedir site describes collection through a trace SDK, app MCP and computer-use agents, followed by custom models, harnesses and routers. Its public workflow is collect, understand, train, deploy and improve. The homepage says data stays in the customer’s account, inference can run in its VPC or on dedicated hardware, and the system can sit alongside the existing AI stack while a new model earns traffic on customer benchmarks.
The YC launch says Belvedir uses LoRA and can host multiple adapters on one cluster, and reports up to 10x lower inference costs and 2x better benchmark performance for ten companies. Those are company-reported outcomes, not an independent cost or quality benchmark. The current FAQ says Belvedir is in closed alpha with a small group of teams; public pricing is not listed.
Why I’d look closer
The advantage is the whole loop. A fine-tuning API usually stops at a checkpoint; Belvedir is promising trace collection, curation, evals, deployment, routing and recursive improvement. Founder context is relevant: YC identifies Zachary Yu as the founder of Traverse, an RL-environment company, and a Waterloo computer-science graduate.
The tradeoff is governance. Production traces can contain customer data, secrets, personal information and hidden policy assumptions. “Private” needs to be tested across collection, training, hosted inference, logs, adapters, support access and deletion. A router that gradually shifts traffic also needs rollback and a benchmark that captures regressions rather than just average score.
What I’d ask
What exactly is collected by each SDK, MCP and CUA path? Can we keep raw traces in our account and inspect every training example? Which model weights, adapters and logs are exportable? How do we approve a deployment, cap traffic, roll back a regression and prove that one customer’s data cannot train another customer’s model?
My editorial take
Shortlist Belvedir if your agent has enough repeatable traffic to generate useful traces and your team cares about private improvement rather than a one-off fine-tune. Start with a non-sensitive workflow and a customer-owned benchmark. The product is compelling when the loop compounds; the diligence burden is making ownership and isolation real at every step.
Quick facts
| Field | Sourced detail |
|---|---|
| Product | Trace collection, custom model training, harnesses, routing and private deployment |
| Buyer | Startups and SMBs with recurring agent traffic, data and inference costs |
| Availability | Closed alpha according to the current FAQ |
| Company-reported result | Up to 10x lower inference cost and 2x benchmark improvement for ten companies |
| Main question | Can the team own, isolate and roll back the entire improvement loop? |
Sources checked
| Source | Checked |
|---|---|
| YC company profile | 2026-09-19 |
| Belvedir homepage | 2026-09-19 |
| Belvedir YC launch | 2026-09-19 |
Cohort context
Belvedir is listed in Summer 2026. In our 2026-09-18 directory snapshot, 119 of 232 listed companies in that cohort have YC’s primary industry label B2B (51.3%). This is a current-directory comparison, not an original intake count or a performance ranking. Nine-cohort dataset.
Public website snapshot
Observed 2026-09-19T16:18:37.138Z in raw homepage HTML. This records visible metadata and advertised links, not agent execution or product quality.
| Signal | Homepage observation |
|---|---|
| Product description metadata | Observed |
| Canonical link | Not observed in this response |
| H1 or H2 heading | Observed |
| Typed structured data | Not observed in this response |
| Docs/developer link | Not observed in this response |
| Pricing link | Not observed in this response |
| llms.txt link | Not observed in this response |
| Markdown alternate | Not observed in this response |
Public observations · Collection method. Missing links here do not establish that a capability or file is absent elsewhere.
