# Cartpole: reinforcement-learning environments with limited public product detail

Canonical: https://mudpie.ai/companies/cartpole/
Breadcrumb: [Home](https://mudpie.ai/) / [Companies](https://mudpie.ai/companies/) / [Cartpole: reinforcement-learning environments with limited public product detail](https://mudpie.ai/companies/cartpole/)
Author: Ali Abouelatta (https://mudpie.ai/authors/ali-abouelatta/)
Published: 2026-09-19
Updated: 2026-09-19
Research type: Company profile
Method: Company and accelerator sources checked 2026-09-19. Product claims are attributed to their sources; this is research, not a hands-on product trial.

Cartpole’s current public identity is reinforcement-learning environments for frontier models, but the public evidence is thin. The official profile is current and specific about the category; the homepage is a one-page contact surface, while the most detailed public launch describes an earlier bug-finding product called Jazzberry.

## What it does

The [YC profile](https://www.ycombinator.com/companies/cartpole) says Cartpole is building reinforcement-learning environments for training frontier models. Its founder, Mateo Perez, is described as a CU Boulder PhD and the company is categorized under AI, reinforcement learning, data labeling and ML. The [current homepage](https://cartpole.com/) only says “cartpole RL environments” and invites contact.

The linked [Jazzberry launch](https://www.ycombinator.com/launches/NQN-jazzberry-ai-bug-finding-with-real-code-execution) is useful history but not evidence that Cartpole currently sells that product. Jazzberry was described as a GitHub pull-request agent that cloned code into a microVM, executed tests and returned a markdown bug report. I am keeping that identity separate from the current Cartpole profile.

## Why I’d look closer

The buyer fit is a frontier-model lab or enterprise AI team that needs environments with realistic tools, long-horizon tasks and measurable rewards. The reason to watch Cartpole is the category itself: a model can look strong on static data while failing when it has to explore, recover and complete a multi-step task.

The tradeoff is evidence and packaging. Public pages do not show a catalog, environment examples, pricing, supported simulators or current customer results. That makes a diligence conversation more useful than a confident product comparison. A buyer should also separate a safe sandbox for training from any environment that can execute customer code or reach real systems.

## What I’d ask

Which environments exist today, and which are custom builds? What tasks, tools and reward signals do they model? Can a customer inspect the environment code, reset state and reproduce a trajectory? What isolation, data-handling and stop controls apply if the environment executes code or interacts with external systems?

## My editorial take

Keep Cartpole on the shortlist for teams building long-horizon model evaluations, but ask for a concrete environment and a customer-owned success metric before treating it as a vendor decision. The company’s current category is clear; the current product proof is not yet public enough to recommend beyond a scoped conversation.

## Quick facts

| Field | Sourced detail |
| --- | --- |
| Current product direction | Reinforcement-learning environments for frontier models |
| Buyer | Frontier AI labs and enterprise teams training/evaluating agents |
| Public availability | Homepage is a contact surface; catalog and pricing not published |
| Historical launch | Jazzberry bug-finding agent; not treated as current Cartpole product evidence |
| Main question | What environment can a buyer inspect, run and measure today? |

## Sources checked

| Source | Checked |
| --- | --- |
| [YC company profile](https://www.ycombinator.com/companies/cartpole) | 2026-09-19 |
| [Cartpole homepage](https://cartpole.com/) | 2026-09-19 |
| [Jazzberry launch linked from YC](https://www.ycombinator.com/launches/NQN-jazzberry-ai-bug-finding-with-real-code-execution) | 2026-09-19 |

## Cohort context

Cartpole is listed in Spring 2025. In our 2026-09-18 directory snapshot, 97 of 143 listed companies in that cohort have YC’s primary industry label B2B (67.8%). This is a current-directory comparison, not an original intake count or a performance ranking. [Nine-cohort dataset](https://mudpie.ai/research/yc-cohorts-2026-09-19.json).

## Public website snapshot

Observed 2026-09-19T16:15:33.304Z in raw homepage HTML. This records visible metadata and advertised links, not agent execution or product quality.

| Signal | Homepage observation |
| --- | --- |
| Product description metadata | Observed |
| Canonical link | Not observed in this response |
| H1 or H2 heading | Observed |
| Typed structured data | Not observed in this response |
| Docs/developer link | Not observed in this response |
| Pricing link | Not observed in this response |
| llms.txt link | Not observed in this response |
| Markdown alternate | Not observed in this response |

[Public observations](https://mudpie.ai/research/yc-homepage-links-2026-09-19.json) · [Collection method](https://mudpie.ai/research/yc-homepage-methods/README.md). Missing links here do not establish that a capability or file is absent elsewhere.


## Author disclosure

I cofound Lazyweb and publish Mudpie. This is an owner-written publication, not an independent testing organization. Research notes distinguish observations, sourced reporting and editorial judgment.
