Developer tools

faster-whisper

Paying per minute to transcribe audio that never needed to leave the machine.

Replaces

Prices are quoted from the vendor's own pricing page on 2026-08-23 and link to it. Vendors change pricing; check before you decide.

Install

pip install faster-whisper
source
SYSTRAN
licence
MIT
verified
2026-08-23

A reimplementation of OpenAI's Whisper on CTranslate2. Same model weights, run on your own hardware.

The reason is not the money

R21 transcribes meetings and client calls with this and has a standing rule never to send that audio to a paid transcription API. At $0.006 a minute, an hour of recording is about thirty-six cents. Nobody would notice the bill.

The reason it matters is what the recordings contain. They are client conversations, and some of them touch medical practices. Audio that never leaves the machine cannot be retained by a third party, cannot be swept into a training set, and cannot be produced by someone else in response to a request. That is a much easier position to hold than a data-processing agreement, and it costs nothing to hold it.

The same reasoning rules out the convenient thing more than once. A meeting recording is the most sensitive artifact R21 handles routinely, and the default path for handling it is an upload.

What it actually runs on

R21 runs large-v3 on a 12 GB RTX 4070 Ti, feeding a research-ingest pipeline that turns recordings into structured notes. That combination matters more than the tool name:

  • The speed claim is a GPU claim. On CPU this is faster than the reference implementation but not fast, and the multiples people quote come from CTranslate2's int8 quantisation on a GPU.
  • Model size is the quality dial, and it is not subtle. Smaller models are noticeably worse on accented English and on proper nouns — which, for a Puerto Rico agency transcribing bilingual calls, is most of the content that matters.

Where it still gets things wrong

Names. Consistently. A transcript will render a person's name three different ways in one meeting, and every one of them will look plausible enough to copy into a document without checking. Any workflow that turns a transcript into a record needs a human pass on the proper nouns, or it will confidently publish a misspelled name that no amount of re-reading the transcript will catch.

Verified against the GitHub API on 2026-08-23: 25,046 stars, MIT.