# David AI: audio data for speech models

Canonical: https://mudpie.ai/companies/david-ai/
Breadcrumb: [Home](https://mudpie.ai/) / [Companies](https://mudpie.ai/companies/) / [David AI: audio data for speech models](https://mudpie.ai/companies/david-ai/)
Author: Ali Abouelatta (https://mudpie.ai/authors/ali-abouelatta/)
Published: 2026-09-19
Updated: 2026-09-19
Research type: Company profile
Method: Company and accelerator sources checked 2026-09-19. Product claims are attributed to their sources; this is research, not a hands-on product trial.

## What it does

David AI develops audio datasets for speech recognition, translation, synthesis and conversational AI. Its current [homepage](https://www.withdavid.ai/) describes a research process that starts with a model capability, designs the data shape, runs targeted collection, evaluates quality, scales to thousands of hours and releases the dataset. The [YC profile](https://www.ycombinator.com/companies/david-ai) calls the company data for audio AI.

The fit is a speech-model or voice-product team that needs more than scraped audio: diverse conversations, language/dialect metadata, speaker separation and a dataset that is licensed for the intended use. David’s value is in the research loop around data design, not only in the number of hours.

## Why I’d look closer

The public dataset suite is concrete. Converse focuses on natural two-speaker English conversations, Atlas spans 15+ languages with dialect and accent metadata, Chorus covers multi-speaker separation, and Dialog focuses on expert conversations. The site says customers can request samples, enter a use-specific license and receive off-the-shelf access within one to two days. Those terms make the first evaluation straightforward, while the buyer still needs to inspect consent, rights and annotation quality.

The founder context is relevant. The [YC biographies](https://www.ycombinator.com/companies/david-ai) describe Tomer Cohen as a former Chief of Staff at Scale AI and McKinsey consultant, and Ben Wiley as the former engineering lead for Scale’s Public Sector GenAI platform and a Microsoft engineer. The company says its datasets are used by Fortune 100 companies and research labs; that is company-reported customer context.

## What I’d ask

What are the collection and consent terms for each language and speaker type, and how are diarization, accent and noise errors measured? I’d request samples with metadata and a license summary, run them through the target model, and compare signal quality to a smaller internally collected set before purchasing scale.

## My editorial take

David AI is a strong fit for model teams that treat data design as research. The hypothesis-to-release process is more useful than a generic “audio dataset” catalog. The key purchase decision is provenance and measured task lift, not corpus size.

## Quick facts

| Field | Sourced detail |
|---|---|
| Buyer fit | Speech recognition, translation, synthesis and conversational-AI teams |
| Dataset examples | Two-speaker, multilingual, multi-speaker and expert-conversation audio |
| Delivery | Samples, licensing, API/contact-led access; company-described |
| Public pricing | Not exposed in the sources checked |

## Sources checked

Checked 2026-09-19.

| Source | Used for |
|---|---|
| [YC company profile](https://www.ycombinator.com/companies/david-ai) | Product, founders and company-reported customer context |
| [David AI homepage](https://www.withdavid.ai/) | Dataset suite and research/production process |
| [David AI news](https://www.withdavid.ai/news) | Public company update surface |

## Cohort context

David AI is listed in Summer 2024. In our 2026-09-18 directory snapshot, 161 of 248 listed companies in that cohort have YC’s primary industry label B2B (64.9%). This is a current-directory comparison, not an original intake count or a performance ranking. [Nine-cohort dataset](https://mudpie.ai/research/yc-cohorts-2026-09-19.json).

## Public website snapshot

Observed 2026-09-19T16:17:00.072Z in raw homepage HTML. This records visible metadata and advertised links, not agent execution or product quality.

| Signal | Homepage observation |
| --- | --- |
| Product description metadata | Observed |
| Canonical link | Not observed in this response |
| H1 or H2 heading | Observed |
| Typed structured data | Not observed in this response |
| Docs/developer link | Not observed in this response |
| Pricing link | Not observed in this response |
| llms.txt link | Not observed in this response |
| Markdown alternate | Not observed in this response |

[Public observations](https://mudpie.ai/research/yc-homepage-links-2026-09-19.json) · [Collection method](https://mudpie.ai/research/yc-homepage-methods/README.md). Missing links here do not establish that a capability or file is absent elsewhere.


## Author disclosure

I cofound Lazyweb and publish Mudpie. This is an owner-written publication, not an independent testing organization. Research notes distinguish observations, sourced reporting and editorial judgment.
