mudpie

Company profile · 3 min read

Moss: sub-10ms semantic search for conversational AI

Local-first semantic-search runtime and managed index distribution for voice agents, copilots, and multimodal AI products.

Published · Updated

Moss is a real-time semantic-search runtime for conversational AI, voice agents, and copilots. The buyer decision is whether retrieval latency and infrastructure overhead are breaking the illusion of a natural conversation—and whether local or on-device search is a better fit than another remote vector database.

What it does

Moss says teams can index and distribute data close to where an agent runs, using a Rust and WebAssembly runtime across browser, mobile, and server environments. It offers JavaScript and Python SDKs, local indexes, instant updates, and managed distribution for conversational and multimodal products (YC profile; Moss homepage).

Fact What the public sources say
Buyer Platform, infrastructure, ML, and product teams building voice or conversational AI
Core job Low-latency semantic retrieval close to the agent runtime
Public performance claim Moss says retrieval can run in sub-10 ms and save 70–90% of tokens versus traditional pipelines
Deployment Browser, mobile, server, and offline-capable or local use cases are described
Founder Sri Raghu Malireddi

Why it fits

Retrieval lag is a product problem when a voice agent pauses or a copilot waits for a round trip before answering. Moss's architecture aims to keep the index near the agent, reduce network hops, and distribute updates without making each team operate a full search service. The local-first path can also matter for sensitive context or offline experiences.

The sub-10 ms, token-savings, design-partner, and customer claims are company-reported. A buyer should benchmark its own corpus, update frequency, memory footprint, recall, multilingual behavior, browser or mobile constraints, and consistency between local and cloud indexes. Public pricing was not visible in the reviewed text, although the company advertises free access and paid plans through its pricing path.

Malireddi's public founder context includes ML leadership at Grammarly and Microsoft, personalization systems used by millions, and real-time ML patents and publications (YC company profile).

Short version: Moss is worth a technical bake-off when retrieval is visible in the user experience. Compare it with the full latency and operating cost of the existing pipeline, not only a query benchmark.

Sources checked — 2026-09-19

Cohort context

Moss is listed in Fall 2025. In our 2026-09-18 directory snapshot, 90 of 146 listed companies in that cohort have YC’s primary industry label B2B (61.6%). This is a current-directory comparison, not an original intake count or a performance ranking. Nine-cohort dataset.

Public website snapshot

Observed 2026-09-19T16:15:12.782Z in raw homepage HTML. This records visible metadata and advertised links, not agent execution or product quality.

Signal Homepage observation
Product description metadata Observed
Canonical link Observed
H1 or H2 heading Observed
Typed structured data Observed
Docs/developer link Observed
Pricing link Observed
llms.txt link Observed
Markdown alternate Not observed in this response

Public observations · Collection method. Missing links here do not establish that a capability or file is absent elsewhere.

About the author

I cofound Lazyweb and publish Mudpie. This is an owner-written publication, not an independent testing organization. Research notes distinguish observations, sourced reporting and editorial judgment.

First1000 ↗ · X ↗