# Aemon: an AI R&D engineer built around measurable evals

Canonical: https://mudpie.ai/companies/aemon/
Breadcrumb: [Home](https://mudpie.ai/) / [Companies](https://mudpie.ai/companies/) / [Aemon: an AI R&D engineer built around measurable evals](https://mudpie.ai/companies/aemon/)
Author: Ali Abouelatta (https://mudpie.ai/authors/ali-abouelatta/)
Published: 2026-09-19
Updated: 2026-09-19
Research type: Company profile
Method: Company and accelerator sources checked 2026-09-19. Product claims are attributed to their sources; this is research, not a hands-on product trial.

Aemon is an AI R&D engineer for teams that know how to measure success but do not know which approach will reach it. The product’s useful wedge is a locked evaluation: give Aemon a problem, a codebase or research context, and a metric, then let it generate and test many candidate approaches while experts steer the search.

## What it does

The [current Aemon site](https://aemon.ai/) describes a four-step engagement: define the goal and failure cases, harden the evaluation, connect Aemon to the codebase and research context, then run continuous research loops that return validated improvements. The [YC launch](https://www.ycombinator.com/launches/PXs-aemon-the-ai-r-d-engineer) says Aemon synthesizes state-of-the-art papers, evolves thousands of approaches against an eval and lets experts constrain the search.

The strongest public demonstration is the circle-packing result. Aemon says it beat Google DeepMind’s 2025 AlphaEvolve result with less than $10 of compute and links to Google’s public verifier. The homepage labels the replay as a visualization rather than the product UI or a limit on what it can solve. That is a useful evidence trail, but I have not independently run the verifier or reproduced the result here.

## Why I’d look closer

The buyer is an engineering or research team with a real objective function: ranking quality, latency, cost, material performance, route efficiency or another metric that can be automated. The founder background is unusually relevant. YC identifies Richard Zhou as a Waterloo CS dropout and math/robotics medalist, and Ray Xu as a UIUC CS dropout with publications at ICLR and EMNLP before turning 20.

The advantage is search breadth with an evaluation loop, not just a generated suggestion. The tradeoff is that the eval becomes the product’s boundary. If the metric misses reliability, maintainability, safety, data leakage or real-world constraints, Aemon can optimize the wrong thing at machine speed. The homepage says a first engagement can reach a technical breakthrough in two weeks; that is a company process claim, not a guarantee.

## What I’d ask

Who owns the eval and the generated code? Can the team inspect every candidate, failed branch and benchmark comparison? How are compute budgets, codebase access, secrets, regressions and human approvals handled? What happens when the best measured solution is too brittle or expensive to ship?

## My editorial take

Shortlist Aemon when an R&D problem has a trusted, executable eval and the upside of a better solution is meaningful. Start with a sandboxed benchmark and a human review gate before allowing changes into production. Aemon’s compelling promise is disciplined search; without a disciplined objective, it is simply faster experimentation.

## Quick facts

| Field | Sourced detail |
| --- | --- |
| Product | Forward-deployed AI research engineer for technical optimization problems |
| Buyer | CTOs, heads of R&D, AI/ML and computational research teams |
| Public demonstration | Circle-packing result with a linked official verifier; not independently reproduced here |
| Pricing | Not published in the checked pages |
| Main question | Is the customer’s eval strong enough to make autonomous search useful? |

## Sources checked

| Source | Checked |
| --- | --- |
| [YC company profile](https://www.ycombinator.com/companies/aemon) | 2026-09-19 |
| [Aemon homepage](https://aemon.ai/) | 2026-09-19 |
| [Aemon YC launch](https://www.ycombinator.com/launches/PXs-aemon-the-ai-r-d-engineer) | 2026-09-19 |
| [Linked public circle-packing verifier](https://colab.research.google.com/github/google-deepmind/alphaevolve_results/blob/master/mathematical_results.ipynb) | 2026-09-19 |

## Cohort context

Aemon is listed in Winter 2026. In our 2026-09-18 directory snapshot, 126 of 199 listed companies in that cohort have YC’s primary industry label B2B (63.3%). This is a current-directory comparison, not an original intake count or a performance ranking. [Nine-cohort dataset](https://mudpie.ai/research/yc-cohorts-2026-09-19.json).

## Public website snapshot

Observed 2026-09-19T16:20:00.344Z in raw homepage HTML. This records visible metadata and advertised links, not agent execution or product quality.

| Signal | Homepage observation |
| --- | --- |
| Product description metadata | Observed |
| Canonical link | Observed |
| H1 or H2 heading | Observed |
| Typed structured data | Observed |
| Docs/developer link | Not observed in this response |
| Pricing link | Not observed in this response |
| llms.txt link | Not observed in this response |
| Markdown alternate | Not observed in this response |

[Public observations](https://mudpie.ai/research/yc-homepage-links-2026-09-19.json) · [Collection method](https://mudpie.ai/research/yc-homepage-methods/README.md). Missing links here do not establish that a capability or file is absent elsewhere.


## Author disclosure

I cofound Lazyweb and publish Mudpie. This is an owner-written publication, not an independent testing organization. Research notes distinguish observations, sourced reporting and editorial judgment.
