Company profile · 3 min read
David AI: audio data for speech models
David AI designs, collects and licenses speech datasets for recognition, translation, synthesis and conversational AI teams.
Published · Updated
What it does
David AI develops audio datasets for speech recognition, translation, synthesis and conversational AI. Its current homepage describes a research process that starts with a model capability, designs the data shape, runs targeted collection, evaluates quality, scales to thousands of hours and releases the dataset. The YC profile calls the company data for audio AI.
The fit is a speech-model or voice-product team that needs more than scraped audio: diverse conversations, language/dialect metadata, speaker separation and a dataset that is licensed for the intended use. David’s value is in the research loop around data design, not only in the number of hours.
Why I’d look closer
The public dataset suite is concrete. Converse focuses on natural two-speaker English conversations, Atlas spans 15+ languages with dialect and accent metadata, Chorus covers multi-speaker separation, and Dialog focuses on expert conversations. The site says customers can request samples, enter a use-specific license and receive off-the-shelf access within one to two days. Those terms make the first evaluation straightforward, while the buyer still needs to inspect consent, rights and annotation quality.
The founder context is relevant. The YC biographies describe Tomer Cohen as a former Chief of Staff at Scale AI and McKinsey consultant, and Ben Wiley as the former engineering lead for Scale’s Public Sector GenAI platform and a Microsoft engineer. The company says its datasets are used by Fortune 100 companies and research labs; that is company-reported customer context.
What I’d ask
What are the collection and consent terms for each language and speaker type, and how are diarization, accent and noise errors measured? I’d request samples with metadata and a license summary, run them through the target model, and compare signal quality to a smaller internally collected set before purchasing scale.
My editorial take
David AI is a strong fit for model teams that treat data design as research. The hypothesis-to-release process is more useful than a generic “audio dataset” catalog. The key purchase decision is provenance and measured task lift, not corpus size.
Quick facts
| Field | Sourced detail |
|---|---|
| Buyer fit | Speech recognition, translation, synthesis and conversational-AI teams |
| Dataset examples | Two-speaker, multilingual, multi-speaker and expert-conversation audio |
| Delivery | Samples, licensing, API/contact-led access; company-described |
| Public pricing | Not exposed in the sources checked |
Sources checked
Checked 2026-09-19.
| Source | Used for |
|---|---|
| YC company profile | Product, founders and company-reported customer context |
| David AI homepage | Dataset suite and research/production process |
| David AI news | Public company update surface |
Cohort context
David AI is listed in Summer 2024. In our 2026-09-18 directory snapshot, 161 of 248 listed companies in that cohort have YC’s primary industry label B2B (64.9%). This is a current-directory comparison, not an original intake count or a performance ranking. Nine-cohort dataset.
Public website snapshot
Observed 2026-09-19T16:17:00.072Z in raw homepage HTML. This records visible metadata and advertised links, not agent execution or product quality.
| Signal | Homepage observation |
|---|---|
| Product description metadata | Observed |
| Canonical link | Not observed in this response |
| H1 or H2 heading | Observed |
| Typed structured data | Not observed in this response |
| Docs/developer link | Not observed in this response |
| Pricing link | Not observed in this response |
| llms.txt link | Not observed in this response |
| Markdown alternate | Not observed in this response |
Public observations · Collection method. Missing links here do not establish that a capability or file is absent elsewhere.
