Company profile · 3 min read
Kashikoi: Simulation-based evaluation for AI agents
Kashikoi simulates customized multi-turn interactions to benchmark AI agents, expose behavioral failures and keep evaluations aligned with a team’s goals.
Published · Updated
Kashikoi is a simulation engine for benchmarking AI agents. It fits a team that wants to test an agent against customized, multi-turn scenarios and behavioral goals instead of trusting public benchmarks, hand-written prompts or a few happy-path traces.
What it does
Kashikoi’s current site lets a team connect an agent, describe what it should do and run simulated interactions that produce metrics and actionable recommendations. The examples show security, data-analysis and code agents evaluated on accuracy, context understanding, severity, false positives and task-specific success.
The useful decision is whether the benchmark reflects the product’s actual values and users. Kashikoi’s YC launch describes CPU-friendly world models that interview agents, generate diverse data and detect stale regression suites. The team says a customer can compare its agent with a competitor without writing every prompt by hand. That is an attractive promise, but the buyer still needs to inspect how scenarios are sampled, how judges are calibrated and how a score turns into an engineering change.
Kashikoi’s site displays example scores for threat detection, a data analyst and code exploration. These are product demonstrations, not independent customer results. Pricing is not public; the current path is a founder demo or waitlist-style conversation.
Founder context and tradeoffs
The YC profile identifies Aaksha Meghawat as founder and CTO and describes simulation/evaluation work at Moveworks, transformer research at CMU and edge speech models at Apple. The launch also identifies Tim Michaud as co-founder and publishes security-research context. Their background maps to the evaluation problem, but a buyer still needs to validate its own domains and failure modes.
Editorial take
I would shortlist Kashikoi for an agent team whose eval suite is stale or too expensive to maintain by hand. Start with one task and a known failure distribution. If the team cannot state what “good” means beyond a composite score, more simulation will create measurement noise rather than confidence.
Quick facts
| Field | Sourced detail |
|---|---|
| Product | Simulation, benchmarking and behavioral evaluation for AI agents |
| Buyers | Agent developers, QA, safety and product teams |
| Public workflow | Custom world models, multi-turn interviews, metrics and recommendations |
| Pricing | Not publicly listed |
| Main gate | Scenario validity, judge calibration, regression freshness and actionability |
Sources checked
| Source | Checked |
|---|---|
| YC profile | 2026-09-19 |
| Kashikoi homepage | 2026-09-19 |
| Kashikoi launch | 2026-09-19 |
| Kashikoi research blog | 2026-09-19 |
Cohort context
Kashikoi is listed in Spring 2025. In our 2026-09-18 directory snapshot, 97 of 143 listed companies in that cohort have YC’s primary industry label B2B (67.8%). This is a current-directory comparison, not an original intake count or a performance ranking. Nine-cohort dataset.
Public website snapshot
Observed 2026-09-19T16:15:41.975Z in raw homepage HTML. This records visible metadata and advertised links, not agent execution or product quality.
| Signal | Homepage observation |
|---|---|
| Product description metadata | Observed |
| Canonical link | Observed |
| H1 or H2 heading | Observed |
| Typed structured data | Observed |
| Docs/developer link | Not observed in this response |
| Pricing link | Not observed in this response |
| llms.txt link | Not observed in this response |
| Markdown alternate | Not observed in this response |
Public observations · Collection method. Missing links here do not establish that a capability or file is absent elsewhere.
