How long will local speech-to-text take on your PC? The answer depends on the model, hardware, audio and other work running on the computer. This page separates an older one-clip timing table from a reproducible starter test with downloadable audio and raw results.
Historical results from July 4, 2026. Each listed GPU timing is shorter than this one clip. They exclude the app's recording, microphone handoff and text insertion. They do not measure a loaded GPU, long recordings or other computers.
| Model | CUDA time (11s clip) | CUDA speed | Vulkan time (11s clip) | Vulkan speed |
|---|---|---|---|---|
| tiny | 0.32s | 34.3x | 0.31s | 35.5x |
| base | 0.41s | 27.2x | 0.42s | 26.5x |
| small | 0.84s | 13.1x | 0.82s | 13.4x |
| medium | 1.91s | 5.7x | 2.02s | 5.4x |
| large-v3 | 3.90s | 2.8x | 3.73s | 3.0x |
These historical CPU timings use eight threads. A faster-than-audio result does not mean text arrives while you speak: the app also has to capture, process and insert it. Results from other CPUs, models or recordings can differ.
| Model | CPU time (11s clip) | CPU speed | Timing on this clip |
|---|---|---|---|
| tiny | 1.98s | 5.6x | Shorter than audio |
| base | 4.54s | 2.4x | Shorter than audio |
| small | 20.4s | 0.5x | Slower than real time |
| medium | 73.6s | 0.1x | Longer than audio |
| large-v3 | 133.5s | 0.1x | Longer than audio |
Start with a short sentence in the language you use, then check the inserted text and waiting time. Compare a smaller model if the delay is too long, and a larger available model if recognition errors matter more. Test again while your usual apps are open. Neither this one-clip table nor the starter lab establishes the best model for every accent, microphone or workload. Read the model guide and offline setup steps.
jfk.wav clip shipped with whisper.cpp, so anyone can run the same test.whisper-cli, with CPU, CUDA, and Vulkan builds. The original summary did not preserve the exact binary/model hashes or raw logs; it is not a fully reproducible build comparison.whisper_print_timings, best of two runs. It is not the complete process wall time or app stop-to-paste time. Best-of-two reporting can hide slower attempts. The speed multiple is nominal audio seconds divided by processing seconds; it is the inverse of the usual real-time factor.This is a small, inspectable engine test, not a marketing accuracy score. It uses six licensed English audiobook clips, three from LibriSpeech test-clean and three from test-other, selected before running recognition. The next section records the measured results and reproduction steps.
Seventy-two attempts completed: six recordings, two models, two backends, three repeats. The same CUDA-capable CLI build was used for both backends, with GPU use disabled for CPU runs and confirmed in each engine log. Each attempt started a new process. Timing includes process startup, model loading, decoding and shutdown, but not live recording or pasting into an app.
| Configuration | Median process time | Completed | First-attempt word edits / reference words |
|---|---|---|---|
| base, CPU | 1.488s | 18 / 18 | 18 / 203 |
| base, CUDA | 0.656s | 18 / 18 | 19 / 203 |
| small, CPU | 8.504s | 18 / 18 | 10 / 203 |
| small, CUDA | 1.080s | 18 / 18 | 10 / 203 |
Word edits count substitutions, deletions and insertions against the supplied reference after lowercase and punctuation normalization. The JSON includes word error rates, every transcript, per-clip minimum/median/maximum, engine logs, GPU snapshots, commands and binary/model hashes. We score the first attempt, not the best transcript. Just 203 reference words cannot establish an accuracy percentage for everyday use. A spelling split such as "goodwill" versus "good will" also affects this score.
The CPU small-model attempts varied from 3.519s to 15.330s. That spread is why a best-case number alone is misleading. CPU runs used the same CUDA-capable executable with GPU decoding disabled, so their process time still includes CUDA discovery and initialization. A CPU-only executable may differ. These timings cannot be compared directly with the historical table's internal engine times. This is a diagnostic starting point, not proof that one model, engine or product always wins.
Download audio, references and runner
Inspect all 72 raw results (JSON) Sample provenance and checksums Read the runner source
--backends cpu,cuda with --backends cpu. A Vulkan build can be tested separately with --backends vulkan; no fresh Vulkan, AMD or Intel results are claimed here.results.json in the output folder. A failed process, empty result or unverified backend remains a failure, not a fast success. Test your own ordinary dictations separately using the app's normal controls.node run-benchmark.mjs --binary "C:/whisper/whisper-cli.exe" --models-dir "C:/whisper/models" --models base,small --backends cpu,cuda --repeats 3 --threads 8 --out "results-my-pc"
The runner reads the included audio and references, your chosen executable, nearby DLL files and model files. It starts that executable, may query NVIDIA's installed system utility and writes a new results folder. The runner itself makes no network calls, records no microphone audio and changes no StarWhisper settings. It is not a security sandbox: the executable runs with your normal file and network access, so only use a build you trust. Inspect the results for private information before sharing them. Audio and models must already be downloaded. Tests can temporarily use substantial resources.
Selection rule: the first FLAC from each of the first three distinct speakers encountered in each official test-clean and test-other archive, before recognition results were known. These are read English audiobooks, not live dictation. Common public speech corpora may overlap model training data. No accent coverage, model independence or real-world accuracy claim is made.
Reference: CONCORD RETURNED TO ITS PLACE AMIDST THE TENTS
Reference: THERE WAS SOMETHING IN HIS AIR AND MANNER THAT BETRAYED TO THE SCOUT THE UTTER CONFUSION OF THE STATE OF HIS MIND
Reference: FOR GOD'S SAKE MY LADY MOTHER GIVE ME A WIFE WHO WOULD BE AN AGREEABLE COMPANION NOT ONE WHO WILL DISGUST ME SO THAT WE MAY BOTH BEAR EVENLY AND WITH MUTUAL GOOD WILL THE YOKE IMPOSED ON US BY HEAVEN INSTEAD OF PULLING THIS WAY AND THAT WAY AND FRETTING EACH OTHER TO DEATH
Reference: HE LAUGHED BUT IT WAS A CURIOUS KIND OF LAUGH FULL OF VEXATION INJURED AMOUR PROPRE AS THE FRENCH CALL OUR LOVE OF OUR OWN DIGNITY OF WHICH ARCHIBALD RAYSTOKE IN THE FULL FLUSH OF HIS YOUNG BELIEF IN HIS IMPORTANCE AS A BRITISH OFFICER HAD A PRETTY GOOD STOCK
Reference: PRESENTLY THE SHIP STRUCK THE MOUNTAIN AND BROKE UP AND ALL AND EVERYTHING ON BOARD OF HER WERE PLUNGED INTO THE SEA
Reference: HARK YOU MY MASTERS YOU THAT LOVE THE WINE COP'S BODY FOLLOW ME FOR SANCT ANTHONY BURN ME AS FREELY AS A FAGGOT IF THEY GET LEAVE TO TASTE ONE DROP OF THE LIQUOR THAT WILL NOT NOW COME AND FIGHT FOR RELIEF OF THE VINE
Audio and references: LibriSpeech by Vassil Panayotov, Guoguo Chen, Daniel Povey and Sanjeev Khudanpur, derived from LibriVox audiobooks, CC BY 4.0. We converted FLAC to mono 16 kHz PCM16 WAV without trimming, gain changes or noise reduction. References are unchanged. Attribution, archive URLs and source/converted hashes are in the manifest. The authors do not endorse StarWhisper.
Download StarWhisper for Windows
Local mode processes speech on your computer after the app and model are downloaded. Optional cloud transcription sends audio for remote processing. Offline transcription does not mean that account, download, payment or update features work without a connection. Watch the dictation workflow.