mudpie

Company profile · 3 min read

Kashikoi: Simulation-based evaluation for AI agents

Kashikoi simulates customized multi-turn interactions to benchmark AI agents, expose behavioral failures and keep evaluations aligned with a team’s goals.

Published · Updated

Kashikoi is a simulation engine for benchmarking AI agents. It fits a team that wants to test an agent against customized, multi-turn scenarios and behavioral goals instead of trusting public benchmarks, hand-written prompts or a few happy-path traces.

What it does

Kashikoi’s current site lets a team connect an agent, describe what it should do and run simulated interactions that produce metrics and actionable recommendations. The examples show security, data-analysis and code agents evaluated on accuracy, context understanding, severity, false positives and task-specific success.

The useful decision is whether the benchmark reflects the product’s actual values and users. Kashikoi’s YC launch describes CPU-friendly world models that interview agents, generate diverse data and detect stale regression suites. The team says a customer can compare its agent with a competitor without writing every prompt by hand. That is an attractive promise, but the buyer still needs to inspect how scenarios are sampled, how judges are calibrated and how a score turns into an engineering change.

Kashikoi’s site displays example scores for threat detection, a data analyst and code exploration. These are product demonstrations, not independent customer results. Pricing is not public; the current path is a founder demo or waitlist-style conversation.

Founder context and tradeoffs

The YC profile identifies Aaksha Meghawat as founder and CTO and describes simulation/evaluation work at Moveworks, transformer research at CMU and edge speech models at Apple. The launch also identifies Tim Michaud as co-founder and publishes security-research context. Their background maps to the evaluation problem, but a buyer still needs to validate its own domains and failure modes.

Editorial take

I would shortlist Kashikoi for an agent team whose eval suite is stale or too expensive to maintain by hand. Start with one task and a known failure distribution. If the team cannot state what “good” means beyond a composite score, more simulation will create measurement noise rather than confidence.

Quick facts

Field Sourced detail
Product Simulation, benchmarking and behavioral evaluation for AI agents
Buyers Agent developers, QA, safety and product teams
Public workflow Custom world models, multi-turn interviews, metrics and recommendations
Pricing Not publicly listed
Main gate Scenario validity, judge calibration, regression freshness and actionability

Sources checked

Source Checked
YC profile 2026-09-19
Kashikoi homepage 2026-09-19
Kashikoi launch 2026-09-19
Kashikoi research blog 2026-09-19

Cohort context

Kashikoi is listed in Spring 2025. In our 2026-09-18 directory snapshot, 97 of 143 listed companies in that cohort have YC’s primary industry label B2B (67.8%). This is a current-directory comparison, not an original intake count or a performance ranking. Nine-cohort dataset.

Public website snapshot

Observed 2026-09-19T16:15:41.975Z in raw homepage HTML. This records visible metadata and advertised links, not agent execution or product quality.

Signal Homepage observation
Product description metadata Observed
Canonical link Observed
H1 or H2 heading Observed
Typed structured data Observed
Docs/developer link Not observed in this response
Pricing link Not observed in this response
llms.txt link Not observed in this response
Markdown alternate Not observed in this response

Public observations · Collection method. Missing links here do not establish that a capability or file is absent elsewhere.

About the author

I cofound Lazyweb and publish Mudpie. This is an owner-written publication, not an independent testing organization. Research notes distinguish observations, sourced reporting and editorial judgment.

First1000 ↗ · X ↗