LmCast :: Stay tuned in

Show HN: Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost

Recorded: Sept. 14, 2026, 6:09 p.m.

Original Summarized

Nari Labs Leads Coval’s Voice AI Benchmarks | Nari Labs

Nari LabsProductsBlogAboutLog InGet StartedModel APIsOptimized open-source modelsbehind simple production APIs.Nari Qwen3-TTS 1.7BPublic BetaNari Qwen3-ASR 1.7BPublic BetaInferenceBring your model.We deploy, optimize, and scale it.TrainingFine-tune multimodal modelswith an experienced team.BlogAbout

Research
Nari Labs Leads Coval’s Voice AI Benchmarks
By the Nari Labs Team · Sep 14, 2026

1st
Voice AI benchmarkSTT Latency, TTS WER

Table of Contents

TL;DR
Speech-to-Text
Text-to-Speech
Get Started

TL;DR
Nari Labs leads Coval’s voice AI benchmark by sitting on the quality-latency Pareto Frontier for both Text-to-Speech and Speech-to-Text. We also lead the latency-cost and quality-cost Pareto Frontier out of all publicly available models on the benchmark.
Coval is a leading provider of voice AI evaluation and benchmarks. They help speech AI agents perform better in production and publish one of the most widely cited benchmarks in the industry.
The Text-to-Speech (TTS) benchmark evaluates latency from text input to first audible chunk of audio (time-to-first-audio or TTFA) and Word Error Rate (WER). The Speech-to-Text (STT) benchmark evaluates latency from user’s finalize request to the final text output (time-to-final-segment or TTFS) and Word Error Rate (WER).
TTFA and TTFS are critical for voice agents, where latency can make a voice AI agent feel unresponsive. Low WER is an obvious key factor for model performance as well.
As of mid September 2026, Nari Labs tops both the Speech-to-Text and Text-to-Speech benchmarks. STT: #1 Latency, #2 WER. TTS: #2 Latency, #1 WER. Note that Coval’s benchmarks can fluctuate every 30 minutes*. We only include publicly available endpoints in our rankings and charts.
Speech-to-Text
Our Qwen3-ASR Fast model is ranked #1 in Time-to-Final-Segment (TTFS), at p50 of 44 ms and WER of 3.6%, placing #2 behind AssemblyAI’s Universal 3.5 Pro at 3.5%.
The pricing makes it even better. At $0.12 / hour, our Fast endpoint ties for the 2nd-lowest price among models with known public rates in Coval’s pricing directory. Universal 3.5 Pro costs 3.75× more, and Deepgram Nova 3 costs 2.4× more. Our Standard endpoint would be the cheapest at $0.06 / hour.
Nari Labs
Nari Labs
Text-to-Speech
Our Qwen3-TTS Fast model is ranked #2 in Time-to-First-Audio (TTFA), at p50 of 63 ms and WER of 3.8%, coming in at #1.
The only model with a lower median TTFA than ours is vui from Fluxions, at 49 ms. It is a 300M parameter model, compared to the 1.7B Qwen3-TTS that we serve.
At $10 per 1M characters, our Fast endpoint is tied for the #1 cheapest model on Coval’s pricing directory. ElevenLabs Eleven v3 Conversational costs 5x more, and Cartesia Sonic 3.6 costs 6.5x more. Our Standard endpoint would be the cheapest at $5 per 1M characters.
Nari Labs
Nari Labs
Interestingly, the official Qwen3 TTS Flash Realtime endpoint sits at 8.8% WER and 692 ms median TTFA. We both serve the same model.
We also surpass Baseten’s dedicated Qwen3-TTS endpoint, which records 6.0% WER and 101 ms median TTFA.
Get Started
Try both our Speech-to-Text and Text-to-Speech models for free for a limited period of time. We are moving our Public Beta APIs to a paid GA within this week and will provide $20 in credits for everyone who has created an account when the switch happens.
Try STT and TTS
Need help meeting the latency and capacity requirements of your voice application? Talk to our engineers
* Benchmark values in this post are based on Coval’s 1-day view as of September 14, 2026, at 15:00 UTC. WER is pooled across datasets. Rankings exclude dedicated inference endpoints. Prices compare Nari’s published rates with known public rates in Coval’s pricing directory.

← All posts

Nari LabsProductModel APIsDedicated InferenceTrainingModelsNari Qwen3-TTS 1.7BNari Qwen3-ASR 1.7BResourcesBlogGitHubHugging FaceCompanyAbout usContact UsLegalTermsPrivacy© 2026 Nari Labs · Backed by Y CombinatorMade with from SF & Seoul

Nari Labs has established leadership in voice artificial intelligence by leading the Coval voice AI benchmarks, focusing on the quality-latency Pareto Frontier for both Text-to-Speech and Speech-to-Text systems, as well as the latency-cost and quality-cost frontiers relative to publicly available models. Coval serves as a leading provider of voice AI evaluation, helping speech AI agents achieve better performance in production environments and publishing highly cited industry benchmarks. The Text-to-Speech benchmark evaluates two critical metrics: time-to-first-audio (TTFA), which measures the latency from text input to the first audible audio chunk, and Word Error Rate (WER). The Speech-to-Text benchmark evaluates time-to-final-segment (TTFS), measuring the latency from the user’s final request to the complete text output, alongside WER. Both TTFA and TTFS are deemed essential for voice agents, as latency directly impacts user experience, while low WER is fundamental for model performance. As of mid-September 2026, Nari Labs ranks first in both the Speech-to-Text and Text-to-Speech benchmarks, specifically ranking first in STT by latency and second in WER, and ranking first in TTS by WER and second in latency.

In the Speech-to-Text domain, the Qwen3-ASR Fast model demonstrates superior performance. It achieved the top rank in Time-to-Final-Segment (TTFS) with a median value of 44 milliseconds and a Word Error Rate of 3.6 percent, positioning it second only to AssemblyAI’s Universal 3.5 Pro at 3.5 percent. Furthermore, the model offers competitive pricing, with the Fast endpoint priced at $0.12 per hour, tying for the second-lowest rate among models with known public pricing in Coval’s directory, significantly undercutting options like Deepgram Nova, which costs four times more.

Regarding Text-to-Speech, the Qwen3-TTS Fast model ranks second in Time-to-First-Audio (TTFA), achieving a median of 63 milliseconds and a WER of 3.8 percent. This model is competitively priced, with the Fast endpoint costing $10 per million characters, positioning it as one of the cheapest options available, while competing providers like ElevenLabs charge five times this amount. A comparative analysis involving the same model indicates that the official Qwen3-TTS Flash Realtime endpoint has a median TTFA of 692 milliseconds and a WER of 8.8 percent. Nari Labs also surpasses Baseten’s dedicated endpoint in this category, recording a median TTFA of 101 milliseconds and a WER of 6.0 percent. Interestingly, when contrasting TTS latency, the only model demonstrating a lower median TTFA is vui from Fluxions, which is a 300 million parameter model, compared to the 1.7 billion parameter Qwen3-TTS model served by Nari Labs. The findings highlight Nari Labs’ position at the intersection of quality and speed for multimodal voice applications.