StarWhisper Windows dictation lab starter pack, September 19, 2026

Audio and reference text: LibriSpeech by Vassil Panayotov, Guoguo Chen,
Daniel Povey and Sanjeev Khudanpur, derived from LibriVox audiobooks.
Source: https://www.openslr.org/12
License: Creative Commons Attribution 4.0 International
https://creativecommons.org/licenses/by/4.0/

Changes: six FLAC clips converted to mono 16 kHz PCM16 WAV with FFmpeg.
No trimming, denoising or gain adjustment. Reference text is unchanged.
The manifest includes selection rules, archive URLs, source member names
and SHA-256 hashes before and after conversion.
No endorsement by the corpus authors or readers is implied.

Measured results: StarWhisper, CC BY 4.0 with attribution to StarWhisper.
https://starwhisper.ai/whisper-benchmark.html#dictation-lab

The runner is supplied for reproducing these tests with a trusted local
whisper.cpp executable. It does not include engine binaries or models,
record the microphone or change app settings. The runner itself makes no
network calls. It reads the supplied executable, nearby DLLs and model
files, may query NVIDIA's installed system utility, and writes results.
The executable runs with your normal file and network access, not in a
security sandbox. Inspect output for private information before sharing.
Use a new output directory. Instructions are on the benchmark page.

The historical upstream jfk.wav audio is not included. Its rights were
not verified and are not covered by the license for our measurements.

The September 19 measurements preceded hardening of log redaction and
the child environment in this runner. The decoding arguments, scoring
and timing definition are unchanged; the raw attempts are unmodified.

Limits: six English audiobook recordings, one computer and two models.
This is not a representative accuracy score, live-microphone acceptance,
a competitor comparison or a prediction of performance on other hardware.
Results may reflect normal background load and public-corpus training overlap.
