mudpie

Company profile · 4 min read

Fabraix: adaptive red-teaming for AI agents

Fabraix’s Nyx agent tests customer-facing AI against adaptive adversarial behavior and surfaces security failures for repeatable remediation.

Published · Updated

What it does

Fabraix builds Nyx, an AI red-teaming agent for customer-facing AI systems. It tests how an agent behaves when it encounters adversarial webpages, documents, files, messages and tool outputs, then surfaces vulnerabilities that can change as the model, prompt, permissions or data sources change. The YC profile describes continuous testing for AI agents; the public case studies show the kind of production failure the company says it investigates.

The fit is a security or AI-platform team shipping agents that can browse, call tools, access sensitive data or take actions. Fabraix is not a generic scanner. Its value depends on adaptive interaction with the agent’s environment and on turning findings into fixes that a security team can reproduce and prioritize.

Why I’d look closer

The public product surface is specific about failure classes: prompt and indirect injection, tool-use hijacking, guardrail bypass, secret or PII exfiltration, transaction fraud and policy drift. The case-study page presents findings from live customer-facing agents, including a coding-agent sandbox escape and examples involving a password-manager credential and network egress. Those are company-published case-study claims; I am describing the evidence surface, not independently reproducing or testing any target.

The YC profile reports a 78% attack-success rate on AgentHarm versus 67% for GPT-5.6 Sol and says Nyx has found vulnerabilities in agents at dozens of Fortune 500 companies. Those are company-reported benchmark and customer-context claims. The product is more interesting if a buyer can see the attack path, affected permission, safe reproduction and remediation regression—not simply a high attack score.

The founder context fits adversarial systems. The YC biographies describe Ahmed Aly as a former data scientist at Two who built systems for adversarial environments, and Ibrahim Abdu as a former Meta privacy red-teamer and SRE/insider-trading systems engineer.

What I’d ask

What is the authorization boundary for a scan, how are replicas and customer environments isolated, and how are secrets handled when a test finds them? I’d request a redacted report, remediation workflow, evidence retention policy and continuous-test diff before allowing any production-facing engagement. The sources checked did not expose a public price card.

My editorial take

Fabraix is a serious fit for teams that treat agent behavior as a changing security surface. The public case studies make the problem legible. The buying bar is safe, scoped evidence and a repeatable regression loop; the benchmark headline is a useful signal, not a substitute.

Quick facts

Field Sourced detail
Buyer fit Security and AI-platform teams shipping customer-facing agents
Product Nyx adaptive red-teaming agent and ACE benchmark work
Public evidence Company case studies and company-reported AgentHarm comparison
Public pricing Not exposed in the sources checked

Sources checked

Checked 2026-09-19.

Source Used for
YC company profile Product, founders and company-reported benchmark/customer claims
Fabraix homepage Current public product surface
Fabraix case studies Public failure classes and case-study descriptions

Cohort context

Fabraix is listed in Summer 2026. In our 2026-09-18 directory snapshot, 119 of 232 listed companies in that cohort have YC’s primary industry label B2B (51.3%). This is a current-directory comparison, not an original intake count or a performance ranking. Nine-cohort dataset.

Public website snapshot

Observed 2026-09-19T16:18:49.251Z in raw homepage HTML. This records visible metadata and advertised links, not agent execution or product quality.

Signal Homepage observation
Product description metadata Observed
Canonical link Not observed in this response
H1 or H2 heading Not observed in this response
Typed structured data Not observed in this response
Docs/developer link Not observed in this response
Pricing link Not observed in this response
llms.txt link Not observed in this response
Markdown alternate Not observed in this response

Public observations · Collection method. Missing links here do not establish that a capability or file is absent elsewhere.

About the author

I cofound Lazyweb and publish Mudpie. This is an owner-written publication, not an independent testing organization. Research notes distinguish observations, sourced reporting and editorial judgment.

First1000 ↗ · X ↗