Chinese voice AI startup VUI Labs has placed its Luna-TTS model at the top of Hugging Face's TTS Arena, a community benchmark for text-to-speech systems. The model outperformed entries from ElevenLabs, MiniMax, and Cartesia to claim the number one position.

On Artificial Analysis' Speech Arena, a separate evaluation framework, Luna-TTS ranks third overall, placing ahead of Google's offering.

Luna-TTS is built on a diffusion architecture derived from Qwen3, a large language model family. The system achieves a first-packet latency of 41.6 milliseconds, a metric that measures the time until the first audio segment is delivered.

The TTS Arena on Hugging Face relies on crowdsourced pairwise comparisons where human listeners judge which of two generated speech samples sounds more natural. Artificial Analysis uses a combination of automated metrics and human evaluation to rank models on quality, latency, and other factors.

VUI Labs has not disclosed the full training data, compute budget, or licensing terms for Luna-TTS. The company has also not released independent verification of the latency figure outside of its own reporting.

The results indicate that a diffusion-based approach can compete with established proprietary systems on both quality and speed benchmarks. However, leaderboard positions can shift as new models are submitted and evaluation methodologies evolve.

Researchers note that arena-style benchmarks capture subjective preference but may not reflect performance in all real-world deployment scenarios, such as noisy environments or long-form narration.

VUI Labs has not announced a public release date, pricing, or API availability for Luna-TTS beyond the benchmark submissions.

Sources and further reading

China's 'Thinking Machines': VUI Labs' Luna-TTS Tops the Global TTS Arena, Beating ElevenLabs and MiniMax

This is an independent summary. The complete reporting, supporting context and any primary documents remain with Pandaily.