mudpie

comparison · 5 min read

Should your agent use Airweave or Crustdata for context?

A founder architecture and buying comparison between Airweave’s private-workspace retrieval and Crustdata’s external people, company, and web data.

Published · Updated

Airweave and Crustdata are both trying to make agents better with data, but they solve opposite sides of the context problem. Airweave makes a company’s private tools searchable. Crustdata makes public company, people, and event data usable through APIs or MCP. The buying decision is about the data boundary, not which product has the more fashionable “agent” label.

Start with the question your agent must answer

Airweave’s public documentation describes an open-source context retrieval layer that connects apps, documents, databases, and workspaces, then exposes a unified search interface to agents. Its examples include internal knowledge assistants searching tools such as Notion, Google Drive, Slack, and other connected sources. The product’s central promise is that an agent can retrieve current, source-grounded context without every application team maintaining a separate connector path.

Crustdata’s public site describes a live graph of people and companies, available through API or MCP, with search, enrichment, signals, and watchers. Its Web Search API page describes a search-fetch-normalize flow that returns structured results for webpages, GitHub repositories, founders, and product launches. Those are external-world tasks: find the right company, person, event, or fresh public signal.

That makes these products complementary in some agent architectures, but not interchangeable retrieval layers.

The context-boundary matrix

Decision dimension Airweave Crustdata
Data boundary Private company tools, documents, apps, and databases Public people, companies, webpages, jobs, posts, and events
Core job Retrieve the right internal context for an agent or RAG system Discover, enrich, resolve, and watch external entities and signals
Main unit Source connection, query, and synced entity Person/company/entity, search request, enrichment, signal, or dataset record
Typical buyer Product or platform team building an internal or customer agent Sales, recruiting, investing, data, or agent-platform team
Interface SDK, REST, MCP, connectors, and searchable collections REST APIs, MCP, web search, people/company APIs, and datasets
Public price signal Developer free: 10 connections, 50 queries/month, 50K entities; Pro $16/month: 50 connections, 500 queries/month, 100K entities, two team members Credit-based API plans and monthly refreshed flat-file datasets; the reviewed public page does not show a numeric rate
Main tradeoff Connector, permission, sync, and retrieval fit for private data Freshness, identity resolution, source coverage, and variable data-access cost

The price entries are dated observations, not a promise that either plan or limit will remain unchanged. Crustdata’s pricing page says its API is credit-based and that its flat-file datasets are refreshed monthly and unified from more than 11 sources; it directs buyers to request samples or enterprise pricing rather than publishing a single universal number.

A worked architecture choice

Imagine a recruiting agent with two jobs: answer questions about the customer’s hiring process and find external candidates who match a new search.

The first job belongs on the private side. The agent may need the customer’s role definitions, interview rubric, approval rules, and internal documents. Airweave’s documented collection-and-source model is relevant because those inputs come from connected company systems.

The second job belongs on the external side. The agent may need current company and people data, job signals, or structured web results. Crustdata’s public product description is aimed at that search, enrichment, and watcher workflow.

The arithmetic for the private side is at least visible before a call. If the agent needs eight source connections, 40 searches in a month, and fewer than 50K synced entities, Airweave’s published Developer limits fit that labeled scenario. If it grows to 20 connections, 300 searches, and fewer than 100K synced entities, the reviewed Pro limits fit the connection and query counts at $16 per month, subject to the plan remaining available and the two-team-member allowance being sufficient. This is a capacity comparison, not a quality test.

The external side needs a different budget model:

monthly context cost =
  private retrieval plan
  + external API credits or dataset quote
  + model and application cost
  + review cost for high-impact matches

Do not pretend the external term can be filled with an invented per-record price. Crustdata’s public pricing page gives the billing shape, not a universal numeric rate. Ask for a quote based on search, enrichment, watcher, and dataset use.

The decision tree

  1. Is the answer inside a workspace the customer already controls? Start with an internal retrieval layer such as Airweave. The hard problem is connectors, permissions, freshness, and search quality over private sources.
  2. Is the answer about a person, company, job, funding event, or public page outside that workspace? Start with an external data layer such as Crustdata. The hard problem is entity resolution, freshness, and coverage.
  3. Does the agent need both? Keep the paths separate until the product has an explicit join rule. A public profile that resembles an internal contact is not automatically the same person, and a private document should not be silently used to enrich an external record.
  4. Are the sources narrow and stable enough to own? If you have two controlled sources and a simple permission model, building may be reasonable. If your product’s value depends on dozens of connectors or a large entity graph, buying or integrating an existing layer deserves the first evaluation.

Founder implication

The strategic choice is not “RAG versus web search.” It is whether your product’s durable value sits in a company’s private context, in the changing external graph, or in the action that joins the two. Airweave is a reference point for turning internal systems into agent-ready context. Crustdata is a reference point for making external entities and public signals actionable. A founder can use both, but should sell one clear job first and state where the data came from when the agent makes a decision.

Sources — snapshots observed 2026-09-19

About the author

I cofound Lazyweb and publish Mudpie. This is an owner-written publication, not an independent testing organization. Research notes distinguish observations, sourced reporting and editorial judgment.

First1000 ↗ · X ↗