Skip to content
Earlyn

Turkish speech-to-text on a Mac: Whisper vs Qwen3-ASR.

On 60 Turkish clips from Google's FLEURS test set, Whisper large-v3-turbo transcribed with an 8.2% word error rate against 10.9% for Qwen3-ASR 1.7B, both running locally on an Apple M2 Mac with 8 GB of memory.

Updated Read as MarkdownDownload the data (CSV)

Result

ModelDownloadWord error rateMedian seconds per clipLanguage detected as Turkish
Whisper large-v3-turbo (q5_0), language detected574 MB8.2% (84 errors / 1,024 words)4.260 of 60
Whisper large-v3-turbo (q5_0), Turkish forced574 MB8.2% (84 / 1,024)2.3n/a
Qwen3-ASR 1.7B (Q8_0)2.2 GB + 356 MB10.9% (112 / 1,024)3.660 of 60
Lower is better. Per-clip numbers: download the CSV.

Whisper stays in Earlyn. Qwen3-ASR reports better averages than Whisper large-v3 across many languages, but on Turkish it made a third more errors here, and its larger files do not fit next to Whisper's graphics memory on an 8 GB Mac: run on the graphics chip it stopped with an out-of-memory error, so we ran its audio encoder on the processor.

Method

  • Audio: the first 60 distinct clips of the FLEURS Turkish (tr_tr) test split, 16 kHz, read speech, 1,024 reference words in total.
  • Whisper: whisper.cpp v1.9.2 inside Earlyn, Metal, greedy decoding, the model already loaded; once with language detection and once with Turkish forced.
  • Qwen3-ASR: llama.cpp release b11222 (llama-mtmd-cli), Q8_0 weights, the audio encoder on the CPU, a new process per clip, so its times include loading the model.
  • Scoring: word error rate = (substitutions + deletions + insertions) / reference words, after lower-casing with Turkish rules (İ→i, I→ı) and removing punctuation.
  • Machine: Apple M2, 8 GB memory, macOS 27, on 1 October 2026.

What we learned for live meetings

Read speech in a test set is not a meeting. When Earlyn cut live audio into 3- to 12-second pieces so lines would appear quickly, Whisper's per-piece language detection misheard Turkish as English, Korean and French in one test meeting. Detecting the language once over the last 30 seconds of speech, then transcribing each piece in that language, fixed it. That is what Earlyn does now.

A second finding: a quiet microphone made every piece of a real meeting look silent, so nothing was transcribed even though the diagnostics counted 189 buffers with sound. Earlyn now treats a piece as silent only when no sample is audible, and raises quiet audio before transcription.

Translation: three local options on the same sentences

SentenceApple TranslationTranslateGemma 4BQwen3 1.7B (general model)
Ödeme entegrasyonu dün bütün testleri geçti.Payment integration passed all tests yesterday.The payment integration has successfully passed all tests yesterday.…passed all tests today.
Daniel will send the contract to Acme by Thursday.Daniel sözleşmeyi Perşembe gününe kadar Acme'ye gönderecek.Daniel, sözleşmeyi Çarşamba gününe kadar Acme'e gönderecek.…Çetin'e 4-5 Haziran tarihinde… (name and date invented)
Errors in bold. A handful of sentences is an example, not a benchmark.

Apple's on-device translation got the days and names right and took about 0.45 seconds a line once loaded (the first call 4 to 6 seconds). TranslateGemma 4B took about 2 seconds a line and is a 2.5 GB download. Earlyn therefore translates with Apple's engine when the language pair is installed and falls back to the answer model otherwise.

Questions

Is Whisper the best model for Turkish?
Of the two local models we measured on an 8 GB Mac, yes. Larger models or cloud services may do better; we did not test them here.
Can I reproduce this?
Yes. FLEURS is public, the per-clip results are in the CSV above, and the Earlyn debug build includes the command we used to run Whisper over a folder of clips.

Sources

  1. FLEURS dataset (Google), CC BY 4.0, accessed 2026-10-01
  2. Qwen3-ASR (Alibaba Qwen team), Apache 2.0, accessed 2026-10-01
  3. Qwen3-ASR 1.7B GGUF (ggml-org), accessed 2026-10-01
  4. TranslateGemma (Google), accessed 2026-10-01