Compare transcripts side by side

One recording. Different transcripts.

Demo · Live provider results are not connected yet.

Reference transcript

0:08

Salom, bugungi uchrashuvni soat uchga ko‘chirsak sizga qulay bo‘ladimi?

Ha, taklifnomani yangilasangiz bo‘ladi.

Local
Global

Illustrative benchmark

1 of 8 models measured · 5 clips · 0:33 audio

Lower is better. WER and CER use case, punctuation, and apostrophe normalization. E2E is one full request; RTF is E2E divided by audio duration. Synthetic audio, dataset stt-demo-v1.

dataset=stt-demo-v1audio=synthetic_voice · PCM 24kHz · 16bit · mononormalizer=basic-text-v1reference=written_formtiming=client_e2e_single_request · mean_single_run
RankProvider and modelWER ↓CER ↓Coverage
24.6%17.6%5/5
0/5
0/5
0/5
0/5
0/5
0/5
0/5

Transcript

VoiceLab · Model not published

Public preview API · 2026-09-02

WER

0.0%

CER

0.0%

E2E

4.09s

RTF

0.54×

mode=batchlang=uztrials=1region=

Salom bugungi uchrashuvni soat uchga ko'chirsak sizga qulay bo'ladimi Ha taklifnomani yangilasangiz bo'ladi.

Reference transcriptTranscript

edits=0

Results by scenario

Test scenarioWERCER
Clear speech0.0%0.0%
Conversation0.0%0.0%
Names and places · Numbers and dates33.3%22.7%
Mixed language10.0%2.9%
Support order80.0%70.5%