Frequently Asked Questions

Find answers to the most common questions about StarWhisper. Can't find what you're looking for? Contact support.

General Questions

Yes! StarWhisper has a permanently free tier with 500 words per day. This includes local transcription with the Tiny model, GPU acceleration, and all basic features. Pro is available as a subscription for unlimited usage, larger models, file transcription, and OpenAI API access.

Yes! After downloading an AI model, StarWhisper can transcribe completely offline using local processing. Your audio never leaves your computer, making it ideal for sensitive work or environments without internet access.

Whisper model accuracy on benchmark audio:

  • Tiny: ~85% (low-power devices, quick notes)
  • Base: ~90% (older hardware)
  • Small: ~94% (recommended default for dictation)
  • Medium (Pro): ~96% on long-form audio and non-Latin script languages
  • Large (Pro): ~98% on batch transcription of recorded files

For typical dictation (short clips), Small is the practical sweet spot. Medium and Large benchmark higher on long recordings but can hallucinate more on short voice-typing clips, so they are not the right default for live dictation. Real-world accuracy also depends on the language being set correctly, microphone quality, and background noise.

Usage Questions

There are two ways to start recording:

  • Click: Left-click the violet recording circle
  • Hotkey: Press Ctrl + Space

When recording, the circle turns orange and pulsates with your voice. Click again or press the hotkey to stop.

After transcription completes, the text is automatically copied to your clipboard. Simply click where you want the text and paste with Ctrl + V (or Cmd + V on Mac).

File transcription is a PRO feature. Pro users can transcribe audio and video files (MP3, WAV, M4A, MP4, AVI, etc.) up to 1 hour in length. Support for batch processing multiple files is included.

Technical Questions

Any microphone works with StarWhisper. For best results:

  • USB Condenser Microphones - Best accuracy (Blue Yeti, Audio-Technica AT2020)
  • Headsets - Consistent positioning (Jabra, Plantronics)
  • Built-in Laptop Mic - Works for casual use

Yes! If you have an NVIDIA GPU (GTX 900 series or newer), StarWhisper can use GPU acceleration for up to 10x faster transcription. GPU acceleration is optional and can be enabled/disabled in settings.

StarWhisper supports 50+ languages with auto-detection. Major languages include English, Spanish, Chinese, Hindi, Arabic, Portuguese, Russian, Japanese, Korean, French, German, and Italian. Setting a specific language improves accuracy by 2-3%.

Privacy & Security

Yes! StarWhisper has two modes:

  • Local Mode: Audio is processed entirely on your computer. Nothing is transmitted anywhere.
  • Cloud Mode: Audio is sent to OpenAI for transcription. Audio is not stored by OpenAI after processing.

You choose which mode to use for each recording.

StarWhisper's local-only architecture is designed to support HIPAA-sensitive workflows. In Local Mode, audio and transcripts never leave your device, which removes the cloud-transmission risk that drives most HIPAA concerns. StarWhisper itself is not HIPAA-certified (no such certification exists for software), and each practice should review the architecture with its own compliance program. Cloud Mode transmits audio to OpenAI and should be disabled for HIPAA-sensitive use.

Troubleshooting

Check these things:

  1. Ensure microphone is connected and not muted
  2. Check Windows microphone privacy settings
  3. Close other apps that might be using the microphone
  4. Try a different USB port for USB microphones
  5. Test microphone in Windows Sound Recorder app

Improve accuracy with these tips, in order:

  • Set a specific language (do not use auto-detect)
  • If you are on Tiny or Base, switch to Small. That fixes most accuracy issues.
  • Speak clearly at a moderate pace
  • Reduce background noise
  • Move microphone closer (6-12 inches)
  • Use a better quality microphone (USB or headset, not Bluetooth)
  • Only consider Medium for long-form audio or non-Latin script languages. On short dictation clips, Small typically gives better results than Medium or Large.