How fast is offline Whisper on Windows? A real speed benchmark (2026)

How long will local speech-to-text take on your PC? The answer depends on the model, hardware, audio and other work running on the computer. This page separates an older one-clip timing table from a reproducible starter test with downloadable audio and raw results.

Published July 4, 2026. Claims and methodology clarified September 19, 2026. The historical table uses a nominal 11-second jfk.wav clip on an Intel Core i9-13980HX (8 threads) and NVIDIA RTX 4090 Laptop GPU, best of two runs. Jump to the new test lab.

The short answer

0.84s
historical small-model CUDA timing for one 11-second clip
1 GPU
RTX 4090 Laptop tested, no cross-vendor conclusion
No WER
the historical table did not measure word errors

GPU results (RTX 4090 Laptop): CUDA vs Vulkan

Historical results from July 4, 2026. Each listed GPU timing is shorter than this one clip. They exclude the app's recording, microphone handoff and text insertion. They do not measure a loaded GPU, long recordings or other computers.

ModelCUDA time (11s clip)CUDA speedVulkan time (11s clip)Vulkan speed
tiny0.32s34.3x0.31s35.5x
base0.41s27.2x0.42s26.5x
small0.84s13.1x0.82s13.4x
medium1.91s5.7x2.02s5.4x
large-v33.90s2.8x3.73s3.0x

CPU results (Core i9-13980HX, 8 threads)

These historical CPU timings use eight threads. A faster-than-audio result does not mean text arrives while you speak: the app also has to capture, process and insert it. Results from other CPUs, models or recordings can differ.

ModelCPU time (11s clip)CPU speedTiming on this clip
tiny1.98s5.6xShorter than audio
base4.54s2.4xShorter than audio
small20.4s0.5xSlower than real time
medium73.6s0.1xLonger than audio
large-v3133.5s0.1xLonger than audio

What this means if you dictate all day

Start with a short sentence in the language you use, then check the inserted text and waiting time. Compare a smaller model if the delay is too long, and a larger available model if recognition errors matter more. Test again while your usual apps are open. Neither this one-clip table nor the starter lab establishes the best model for every accent, microphone or workload. Read the model guide and offline setup steps.

Methodology (so you can reproduce it)
Sample
The canonical 11.0-second jfk.wav clip shipped with whisper.cpp, so anyone can run the same test.
Engine
whisper.cpp whisper-cli, with CPU, CUDA, and Vulkan builds. The original summary did not preserve the exact binary/model hashes or raw logs; it is not a fully reproducible build comparison.
Timing
Internal total time reported by whisper_print_timings, best of two runs. It is not the complete process wall time or app stop-to-paste time. Best-of-two reporting can hide slower attempts. The speed multiple is nominal audio seconds divided by processing seconds; it is the inverse of the usual real-time factor.
Hardware
Intel Core i9-13980HX (8 threads), NVIDIA RTX 4090 Laptop GPU, 32 GB RAM, Windows 11, as recorded in the July summary.
Limits
One short English recording, one computer, no accuracy scoring and no load testing. Do not extrapolate the numbers to AMD/Intel GPUs, other languages, long files or live microphone reliability. The new lab below records each attempt and file hashes.

Windows dictation lab: reproducible starter test

This is a small, inspectable engine test, not a marketing accuracy score. It uses six licensed English audiobook clips, three from LibriSpeech test-clean and three from test-other, selected before running recognition. The next section records the measured results and reproduction steps.

Run September 19, 2026, 14:41-14:44 UTC. Intel Core i9-13980HX, 8 threads, 63.6 GiB system RAM reported to Node, NVIDIA RTX 4090 Laptop GPU (16 GB), driver 610.88, Windows build 26200. Normal desktop background work remained running. The GPU reported 1% use and about 5.3 GiB allocated before testing; this was not an isolated laboratory machine.

Seventy-two attempts completed: six recordings, two models, two backends, three repeats. The same CUDA-capable CLI build was used for both backends, with GPU use disabled for CPU runs and confirmed in each engine log. Each attempt started a new process. Timing includes process startup, model loading, decoding and shutdown, but not live recording or pasting into an app.

New-process timings from one CUDA-capable build, not pure inference speed or an overall model ranking. Median across 18 attempts per configuration; clips are 3.5-18.8 seconds long.
ConfigurationMedian process timeCompletedFirst-attempt word edits / reference words
base, CPU1.488s18 / 1818 / 203
base, CUDA0.656s18 / 1819 / 203
small, CPU8.504s18 / 1810 / 203
small, CUDA1.080s18 / 1810 / 203

Word edits count substitutions, deletions and insertions against the supplied reference after lowercase and punctuation normalization. The JSON includes word error rates, every transcript, per-clip minimum/median/maximum, engine logs, GPU snapshots, commands and binary/model hashes. We score the first attempt, not the best transcript. Just 203 reference words cannot establish an accuracy percentage for everyday use. A spelling split such as "goodwill" versus "good will" also affects this score.

The CPU small-model attempts varied from 3.519s to 15.330s. That spread is why a best-case number alone is misleading. CPU runs used the same CUDA-capable executable with GPU decoding disabled, so their process time still includes CUDA discovery and initialization. A CPU-only executable may differ. These timings cannot be compared directly with the historical table's internal engine times. This is a diagnostic starting point, not proof that one model, engine or product always wins.

Download audio, references and runner

Inspect all 72 raw results (JSON)   Sample provenance and checksums   Read the runner source

Reproduce it on Windows

  1. Download and extract the starter pack. Install Node.js if needed and obtain a compatible whisper.cpp CLI build and the base/small model files. The pack contains no installer, engine binary or model.
  2. Open a terminal in the extracted folder. Replace the example binary and model paths below with your local files. The output folder must not already exist.
  3. Run CPU first if you do not have a CUDA-capable build and compatible NVIDIA GPU: replace --backends cpu,cuda with --backends cpu. A Vulkan build can be tested separately with --backends vulkan; no fresh Vulkan, AMD or Intel results are claimed here.
  4. Read results.json in the output folder. A failed process, empty result or unverified backend remains a failure, not a fast success. Test your own ordinary dictations separately using the app's normal controls.
node run-benchmark.mjs --binary "C:/whisper/whisper-cli.exe" --models-dir "C:/whisper/models" --models base,small --backends cpu,cuda --repeats 3 --threads 8 --out "results-my-pc"

The runner reads the included audio and references, your chosen executable, nearby DLL files and model files. It starts that executable, may query NVIDIA's installed system utility and writes a new results folder. The runner itself makes no network calls, records no microphone audio and changes no StarWhisper settings. It is not a security sandbox: the executable runs with your normal file and network access, so only use a build you trust. Inspect the results for private information before sharing them. Audio and models must already be downloaded. Tests can temporarily use substantial resources.

Listen to the exact samples

Selection rule: the first FLAC from each of the first three distinct speakers encountered in each official test-clean and test-other archive, before recognition results were known. These are read English audiobooks, not live dictation. Common public speech corpora may overlap model training data. No accent coverage, model independence or real-world accuracy claim is made.

Clean 1: 6930-75918-0000, 3.505 seconds

Reference: CONCORD RETURNED TO ITS PLACE AMIDST THE TENTS

Clean 2: 1320-122617-0003, 6.285 seconds

Reference: THERE WAS SOMETHING IN HIS AIR AND MANNER THAT BETRAYED TO THE SCOUT THE UTTER CONFUSION OF THE STATE OF HIS MIND

Clean 3: 5639-40744-0032, 17.430 seconds

Reference: FOR GOD'S SAKE MY LADY MOTHER GIVE ME A WIFE WHO WOULD BE AN AGREEABLE COMPANION NOT ONE WHO WILL DISGUST ME SO THAT WE MAY BOTH BEAR EVENLY AND WITH MUTUAL GOOD WILL THE YOKE IMPOSED ON US BY HEAVEN INSTEAD OF PULLING THIS WAY AND THAT WAY AND FRETTING EACH OTHER TO DEATH

Other 1: 7902-96591-0008, 14.730 seconds

Reference: HE LAUGHED BUT IT WAS A CURIOUS KIND OF LAUGH FULL OF VEXATION INJURED AMOUR PROPRE AS THE FRENCH CALL OUR LOVE OF OUR OWN DIGNITY OF WHICH ARCHIBALD RAYSTOKE IN THE FULL FLUSH OF HIS YOUNG BELIEF IN HIS IMPORTANCE AS A BRITISH OFFICER HAD A PRETTY GOOD STOCK

Other 2: 7018-75788-0016, 7.540 seconds

Reference: PRESENTLY THE SHIP STRUCK THE MOUNTAIN AND BROKE UP AND ALL AND EVERYTHING ON BOARD OF HER WERE PLUNGED INTO THE SEA

Other 3: 4198-12281-0008, 18.810 seconds

Reference: HARK YOU MY MASTERS YOU THAT LOVE THE WINE COP'S BODY FOLLOW ME FOR SANCT ANTHONY BURN ME AS FREELY AS A FAGGOT IF THEY GET LEAVE TO TASTE ONE DROP OF THE LIQUOR THAT WILL NOT NOW COME AND FIGHT FOR RELIEF OF THE VINE

Audio and references: LibriSpeech by Vassil Panayotov, Guoguo Chen, Daniel Povey and Sanjeev Khudanpur, derived from LibriVox audiobooks, CC BY 4.0. We converted FLAC to mono 16 kHz PCM16 WAV without trimming, gain changes or noise reduction. References are unchanged. Attribution, archive URLs and source/converted hashes are in the manifest. The authors do not endorse StarWhisper.

Download StarWhisper for Windows

Local mode processes speech on your computer after the app and model are downloaded. Optional cloud transcription sends audio for remote processing. Offline transcription does not mean that account, download, payment or update features work without a connection. Watch the dictation workflow.