mudpie

Company profile · 4 min read

Halluminate: resettable computer-use environments for agent training

Halluminate provides managed sandboxes, datasets and evaluations for teams training and testing browser and computer-use agents.

Published · Updated

Halluminate builds training data, evaluations and resettable environments for teams teaching AI to use browsers and business software. Its current public wedge is financial services, but the underlying buyer is any model lab or enterprise that needs computer-use behavior tested without letting an agent touch a live customer system.

What it does

The current Halluminate site describes reinforcement-learning environments for investment banking, private equity and consulting workflows. The YC profile describes a broader suite: managed sandboxes modeled after systems such as Salesforce, Slack and ticketing software, proprietary datasets and expert-annotated evaluation services. The launch explains why the sandbox matters: real sites are noisy, dynamic and hard to reset, with authentication, ads and other side effects mixed into the test.

The buyer is a frontier-model lab or enterprise AI team building browser or computer-use agents. Halluminate says paying customers include leading computer-use model labs and the two largest browser-agent companies in the space; that is company-reported customer context. Its offer is less “give us a benchmark score” and more “give us a reproducible place to find failure modes, train against them and compare the next version.”

Why I’d look closer

The advantage is experimental control. A sandbox can reset state, run in parallel and expose a clean reward or failure signal without sending an agent to a real CRM, ticket queue or financial system. Founder context fits the problem: Jerry Wu previously led product and research at Capital One Labs and co-authored patents; Wyatt Marshall is described as a data and software engineer across early startups.

The tradeoff is realism. A simulated Salesforce or financial workflow is only useful if its permissions, data dependencies, error messages and incentives resemble the customer’s real environment. A model can learn to win the sandbox while failing on the live system it was meant to operate.

What I’d ask

Which financial workflows and sandboxes are available today? Can a customer define its own roles, records, tool contracts and reset conditions? How are annotations audited? Does a benchmark measure recovery from partial failure and permission denial, or only task completion? Can the same trajectory be replayed after the environment changes?

My editorial take

Shortlist Halluminate if computer-use reliability is blocked by unsafe or irreproducible testing. Start with one owned workflow and compare sandbox failures against a controlled fixture, keeping real credentials and customer data out of the environment. The product is strongest when it makes the failure surface legible; its risk is optimizing for a clean simulation nobody actually uses.

Quick facts

Field Sourced detail
Product Managed RL environments, datasets and expert evaluation for computer-use agents
Buyer Model labs and enterprise AI teams, with current public focus on financial services
Environments named Salesforce, Slack, ticketing and financial-services workflows
Pricing Not published in the checked pages
Main question Does the sandbox preserve the failure modes that matter in the live workflow?

Sources checked

Source Checked
YC company profile 2026-09-19
Halluminate homepage 2026-09-19
Halluminate YC launch 2026-09-19

Cohort context

Halluminate is listed in Summer 2025. In our 2026-09-18 directory snapshot, 112 of 166 listed companies in that cohort have YC’s primary industry label B2B (67.5%). This is a current-directory comparison, not an original intake count or a performance ranking. Nine-cohort dataset.

Public website snapshot

Observed 2026-09-19T16:18:06.434Z in raw homepage HTML. This records visible metadata and advertised links, not agent execution or product quality.

Signal Homepage observation
Product description metadata Observed
Canonical link Not observed in this response
H1 or H2 heading Observed
Typed structured data Observed
Docs/developer link Not observed in this response
Pricing link Not observed in this response
llms.txt link Not observed in this response
Markdown alternate Not observed in this response

Public observations · Collection method. Missing links here do not establish that a capability or file is absent elsewhere.

About the author

I cofound Lazyweb and publish Mudpie. This is an owner-written publication, not an independent testing organization. Research notes distinguish observations, sourced reporting and editorial judgment.

First1000 ↗ · X ↗