Company profile · 3 min read
.Wave: continuous inference for live voice and streaming AI sessions
Continuous-inference engine for streaming voice, speech, video, and custom stateful AI workloads with GPU scheduling and session management.
Published · Updated
.Wave is a continuous-inference engine for streaming AI models, aimed at teams building live voice, speech, video, or other stateful sessions. The buyer decision is whether the workload is constrained by per-session GPU cost and timing—not whether a conventional request-serving stack is simply “slow.”
What it does
.Wave says customers bring their own streaming model while the engine manages execution, session state, scheduling, and GPU sharing. Its public examples cover full-duplex voice and streaming speech recognition, with live-video and custom streaming architectures as adjacent uses (.Wave homepage; YC profile).
| Fact | What the public sources say |
|---|---|
| Buyer | Teams operating real-time voice, speech, video, or stateful AI sessions |
| Core job | Fit more concurrent sessions on a GPU while keeping each stream on time |
| Public case study | 56 concurrent Nemotron VoiceChat 11B sessions on one H100, with 0 late beats in 84,000 measured session-beats |
| Cost claim | The same case study reports approximately 98% lower GPU cost per conversation under its stated conditions |
| Founders | Thomas Minassian and Paul-Henri Biojout |
Why it fits
The useful distinction is the unit of work. A live session keeps producing small pieces of state; waiting for a batch request or redoing model setup can waste both time and GPU memory. .Wave's engine is designed to keep those sessions together, reuse model-weight reads, and make the schedule explicit. That is a good fit for a voice agent where a missed timing deadline is a broken conversation.
The performance numbers come from .Wave's own case study, which names the model, H100 hardware, 56-session profile, three runs, a 160 ms beat, and the fact that client transport and audio playout were excluded (.Wave case study). A buyer should reproduce the test with its model, context length, precision, transport, concurrency, and p99 target. The public sources do not show pricing.
The founder background fits the workload. YC describes Minassian and Biojout as the team behind Parler-TTS Mini Multilingual and Dollyglot's real-time video-avatar work (YC company profile). That is relevant model and streaming experience, not a guarantee that every architecture will see the same density.
Short version: .Wave is worth a technical evaluation for live AI products where GPU economics and deadline misses are the bottleneck. Bring a real workload and compare the whole serving path, not just the kernel benchmark.
Sources checked — 2026-09-19
Cohort context
.Wave is listed in Winter 2025. In our 2026-09-18 directory snapshot, 104 of 165 listed companies in that cohort have YC’s primary industry label B2B (63.0%). This is a current-directory comparison, not an original intake count or a performance ranking. Nine-cohort dataset.
Public website snapshot
Observed 2026-09-19T16:19:20.069Z in raw homepage HTML. This records visible metadata and advertised links, not agent execution or product quality.
| Signal | Homepage observation |
|---|---|
| Product description metadata | Observed |
| Canonical link | Observed |
| H1 or H2 heading | Observed |
| Typed structured data | Observed |
| Docs/developer link | Not observed in this response |
| Pricing link | Not observed in this response |
| llms.txt link | Not observed in this response |
| Markdown alternate | Not observed in this response |
Public observations · Collection method. Missing links here do not establish that a capability or file is absent elsewhere.
