Show HN: Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost
Recorded: Sept. 14, 2026, 6:09 p.m.
| Original | Summarized |
Nari Labs Leads Coval’s Voice AI Benchmarks | Nari Labs Nari LabsProductsBlogAboutLog InGet StartedModel APIsOptimized open-source modelsbehind simple production APIs.Nari Qwen3-TTS 1.7BPublic BetaNari Qwen3-ASR 1.7BPublic BetaInferenceBring your model.We deploy, optimize, and scale it.TrainingFine-tune multimodal modelswith an experienced team.BlogAbout Research 1st Table of Contents TL;DR TL;DR ← All posts Nari LabsProductModel APIsDedicated InferenceTrainingModelsNari Qwen3-TTS 1.7BNari Qwen3-ASR 1.7BResourcesBlogGitHubHugging FaceCompanyAbout usContact UsLegalTermsPrivacy© 2026 Nari Labs · Backed by Y CombinatorMade with from SF & Seoul |
Nari Labs has established leadership in voice artificial intelligence by leading the Coval voice AI benchmarks, focusing on the quality-latency Pareto Frontier for both Text-to-Speech and Speech-to-Text systems, as well as the latency-cost and quality-cost frontiers relative to publicly available models. Coval serves as a leading provider of voice AI evaluation, helping speech AI agents achieve better performance in production environments and publishing highly cited industry benchmarks. The Text-to-Speech benchmark evaluates two critical metrics: time-to-first-audio (TTFA), which measures the latency from text input to the first audible audio chunk, and Word Error Rate (WER). The Speech-to-Text benchmark evaluates time-to-final-segment (TTFS), measuring the latency from the user’s final request to the complete text output, alongside WER. Both TTFA and TTFS are deemed essential for voice agents, as latency directly impacts user experience, while low WER is fundamental for model performance. As of mid-September 2026, Nari Labs ranks first in both the Speech-to-Text and Text-to-Speech benchmarks, specifically ranking first in STT by latency and second in WER, and ranking first in TTS by WER and second in latency. In the Speech-to-Text domain, the Qwen3-ASR Fast model demonstrates superior performance. It achieved the top rank in Time-to-Final-Segment (TTFS) with a median value of 44 milliseconds and a Word Error Rate of 3.6 percent, positioning it second only to AssemblyAI’s Universal 3.5 Pro at 3.5 percent. Furthermore, the model offers competitive pricing, with the Fast endpoint priced at $0.12 per hour, tying for the second-lowest rate among models with known public pricing in Coval’s directory, significantly undercutting options like Deepgram Nova, which costs four times more. Regarding Text-to-Speech, the Qwen3-TTS Fast model ranks second in Time-to-First-Audio (TTFA), achieving a median of 63 milliseconds and a WER of 3.8 percent. This model is competitively priced, with the Fast endpoint costing $10 per million characters, positioning it as one of the cheapest options available, while competing providers like ElevenLabs charge five times this amount. A comparative analysis involving the same model indicates that the official Qwen3-TTS Flash Realtime endpoint has a median TTFA of 692 milliseconds and a WER of 8.8 percent. Nari Labs also surpasses Baseten’s dedicated endpoint in this category, recording a median TTFA of 101 milliseconds and a WER of 6.0 percent. Interestingly, when contrasting TTS latency, the only model demonstrating a lower median TTFA is vui from Fluxions, which is a 300 million parameter model, compared to the 1.7 billion parameter Qwen3-TTS model served by Nari Labs. The findings highlight Nari Labs’ position at the intersection of quality and speed for multimodal voice applications. |