Skip to content
Speech to textASRTranscriptionOpenAIWhisper

OpenAI Whisper: what it is, how to use it, cost

Whisper is OpenAI's free, MIT-licensed speech-to-text model. How to run it locally, which model to pick, API prices and its known limits.

John Jacob, Shantanu Nair

,

Updated 6 min read
OpenAI Whisper: what it is, how to use it, cost

OpenAI Whisper is an open-source speech recognition model. It turns speech into text in many languages and can translate non-English speech into English. The code and model weights are free under the MIT License (GitHub), so you can run it on your own machine with pip install -U openai-whisper and ffmpeg. If you don't want to run it yourself, OpenAI's API hosts it as whisper-1 at $0.006 per minute, next to newer transcription models such as gpt-transcribe at $0.0045 per minute (OpenAI pricing, as of September 2026). Two things to know before you start: open-source Whisper does not label speakers, and it sometimes writes text that was never spoken.

Whisper at a glance

FactDetail
Made byOpenAI
First releaseSeptember 2022; large-v2 in December 2022, large-v3 in November 2023, turbo in September 2024 (model card)
LicenseMIT, for both code and weights
What it doesTranscription, speech translation to English, language identification
Training data680,000 hours of audio: about 65% English, with the non-English audio covering 98 languages (model card)
Latest pip release20250625 (PyPI)
Hosted versionwhisper-1 in the OpenAI API, $0.006/minute

Whisper models: which one to pick

The open-source release comes in six sizes. This table is from the README; speeds are relative to large, measured on English speech on an A100 GPU, and your hardware will differ.

ModelParametersEnglish-only versionVRAM neededRelative speed
tiny39 Mtiny.en~1 GB~10x
base74 Mbase.en~1 GB~7x
small244 Msmall.en~2 GB~4x
medium769 Mmedium.en~5 GB~2x
large1550 Mnone~10 GB1x
turbo809 Mnone~6 GB~8x

The README lists turbo at 809 M parameters; the model card gives 798 M.

How to choose:

  • turbo is the default in the command-line tool. The README describes it as an optimized large-v3 with faster transcription and minimal loss of accuracy. It is the right starting point for most transcription.
  • turbo is not trained for translation. If you pass --task translate, it returns the original language. Use medium or large to translate into English.
  • English-only .en models tend to do better on English audio than the multilingual model of the same size, especially at tiny and base.
  • Accuracy varies a lot by language. The README links a word-error-rate breakdown by language for large-v3 and large-v2. Check your language before you commit to Whisper.

How to run Whisper locally

You need Python (the README says 3.8 to 3.11 are expected to work), PyTorch, and the ffmpeg command-line tool. A GPU helps a lot; on CPU, the larger models are slow.

  1. Install ffmpeg:
1# macOS (Homebrew)
2brew install ffmpeg
3
4# Ubuntu or Debian
5sudo apt update && sudo apt install ffmpeg
  1. Install Whisper:
1pip install -U openai-whisper

If the install fails while building tiktoken, the README says you may need Rust installed.

  1. Transcribe a file from the terminal:
1whisper audio.mp3 --model turbo

Whisper detects the language from the first 30 seconds. You can set it yourself, and translate with a multilingual model:

1whisper japanese.wav --language Japanese
2whisper japanese.wav --model medium --language Japanese --task translate

By default the command writes every output format (txt, vtt, srt, tsv, json). Use --output_format srt for subtitles only, and whisper --help for the full list of options. Word-level timestamps are available as an experimental flag, --word_timestamps True.

  1. Or call it from Python:
1import whisper
2
3model = whisper.load_model("turbo")
4result = model.transcribe("audio.mp3")
5print(result["text"])

result["segments"] holds the timed segments, and result["language"] the detected language. The model files download on first use to ~/.cache/whisper.

How to use Whisper through the OpenAI API

If you don't want to manage a GPU, the speech-to-text API runs Whisper and newer models for you. As of September 2026:

ModelBest forOutput formatsPrice per minute
gpt-transcribeGeneral transcription (OpenAI's current recommendation)JSON$0.0045
gpt-4o-mini-transcribeLower costJSON$0.003
gpt-4o-transcribeTranscriptionJSON$0.006
gpt-4o-transcribe-diarizeSpeaker labelsdiarized_json, JSON, text$0.006
whisper-1Subtitles, timestamps, translation to Englishsrt, vtt, verbose_json, JSON, text$0.006

Per-minute prices are OpenAI's estimates from its pricing page; some of these models are billed per token. Limits worth knowing, from the same guide:

  • Uploads are capped at 25 MB per file. Supported types are mp3, mp4, mpeg, mpga, m4a, wav and webm. Longer recordings need to be compressed or split.
  • Only whisper-1 returns srt/vtt and supports word or segment timestamps through timestamp_granularities. gpt-4o-transcribe-diarize doesn't support timestamp_granularities, but its diarized_json segments include start and end times.
  • The translation endpoint uses whisper-1 only and translates into English only.
  • whisper-1 prompts are limited to 224 tokens.

A minimal request that returns subtitles:

1curl https://api.openai.com/v1/audio/transcriptions \
2 -H "Authorization: Bearer $OPENAI_API_KEY" \
3 -F file="@audio.mp3" \
4 -F model="whisper-1" \
5 -F response_format="srt"

How Whisper works

Whisper is a standard Transformer encoder-decoder, described in the paper Robust Speech Recognition via Large-Scale Weak Supervision. Audio is cut into 30-second windows and converted to a log-Mel spectrogram. The encoder processes the spectrogram, and the decoder predicts text along with special tokens that tell it which task to do: identify the language, transcribe, translate to English, or add timestamps. One model replaces what used to be several separate pipeline stages.

Whisper architecture: 30-second audio windows become log-Mel spectrograms, pass through a Transformer encoder, and a decoder predicts text and task tokens

Image source: OpenAI

The difference from earlier systems was the data. Instead of training on small, carefully labelled datasets, OpenAI trained on 680,000 hours of audio paired with transcripts collected from the web, with a large share in languages other than English and a share used for translation to English.

Breakdown of Whisper's 680,000 hours of training data by task and language

Image source: OpenAI

Known limits

  • Hallucinations. The model card warns that output may include text not spoken in the audio. Users report this most often on silence, music or noise (see the GitHub discussion Hallucination on audio with no speech). Trimming silence or running voice activity detection first helps, and the CLI has a --hallucination_silence_threshold option (it requires --word_timestamps True).
  • Repetition. The model card also notes the architecture can get stuck repeating text.
  • Uneven accuracy. Performance is lower on low-resource languages and can differ across accents and dialects (model card).
  • No speaker labels. Open-source Whisper gives you text and timestamps, not "who said what". On the API, speaker labels come from gpt-4o-transcribe-diarize, not whisper-1.
  • turbo won't translate, and API uploads stop at 25 MB.
  • Not for high-stakes decisions. OpenAI cautions against using it in high-risk decision-making and against transcribing people recorded without their consent.

When a hosted tool makes more sense

Whisper is a model, not an app. Running it yourself means installing Python and ffmpeg, picking a model, having a GPU for speed, and then editing the raw text somewhere else. That is a good trade if you need privacy on your own hardware or you are building a product.

If you just need a transcript, subtitles or a summary, a hosted editor is less work. Disclosure: we make Exemplary AI. It transcribes audio and video in 99 languages in the browser, gives you a document-style editor where you can fix words and rename speakers, exports subtitles such as .srt, and can then translate, summarize or cut the recording into clips. There is a free plan; see pricing for limits.

Transcribe without the setup

Upload a file or paste a link and edit the transcript in your browser.

Try Exemplary AI free

FAQ

Is OpenAI Whisper free?

Yes. The open-source Whisper code and model weights are free under the MIT License, including for commercial use; you only pay for your own hardware. The hosted whisper-1 model in OpenAI's API costs $0.006 per minute as of September 2026.

Which Whisper model is the best?

For English and most transcription, start with turbo: it is the CLI default, and the README describes it as an optimized large-v3 that is much faster with minimal loss of accuracy. Use medium or large if you need translation into English, and a .en model if you only transcribe English on limited hardware.

Can Whisper identify different speakers?

No. Open-source Whisper and whisper-1 return text and timestamps without speaker labels. OpenAI's gpt-4o-transcribe-diarize API model returns speaker-labelled segments, or you can use a transcription tool with speaker editing.

Does Whisper work offline?

Yes. After pip install -U openai-whisper downloads the package and the first run downloads the model file, transcription runs entirely on your machine.

What's the difference between Whisper and gpt-transcribe?

Whisper is the open model you can run yourself, and whisper-1 is its hosted API version with subtitle, timestamp and translation output. gpt-transcribe is a newer hosted-only model that OpenAI's speech-to-text guide recommends for general transcription; it returns JSON and costs $0.0045 per minute.

What is the file size limit for Whisper?

The OpenAI API accepts files up to 25 MB. Running Whisper locally has no fixed file size limit; long files just take longer and need enough memory.

One upload. Every format you need.

Drop in a recording and work from the transcript — clips, captions, chapters, show notes and posts all come from the same file.