---
title: "OpenAI Whisper: what it is, how to use it, cost"
description: "Whisper is OpenAI's free, MIT-licensed speech-to-text model. How to run it locally, which model to pick, API prices and its known limits."
url: https://exemplary.ai/blog/openai-whisper
published: 2022-09-28
updated: 2026-09-25
author: "John Jacob, Shantanu Nair"
publisher: "Exemplary AI (https://exemplary.ai)"
---

# OpenAI Whisper: what it is, how to use it, cost

OpenAI Whisper is an open-source speech recognition model. It turns speech into text in many languages and can translate non-English speech into English. The code and model weights are free under the MIT License ([GitHub](https://github.com/openai/whisper)), so you can run it on your own machine with `pip install -U openai-whisper` and ffmpeg. If you don't want to run it yourself, OpenAI's API hosts it as `whisper-1` at $0.006 per minute, next to newer transcription models such as `gpt-transcribe` at $0.0045 per minute ([OpenAI pricing](https://developers.openai.com/api/docs/pricing), as of September 2026). Two things to know before you start: open-source Whisper does not label speakers, and it sometimes writes text that was never spoken.

## Whisper at a glance

| Fact | Detail |
| --- | --- |
| Made by | OpenAI |
| First release | September 2022; large-v2 in December 2022, large-v3 in November 2023, turbo in September 2024 ([model card](https://github.com/openai/whisper/blob/main/model-card.md)) |
| License | MIT, for both code and weights |
| What it does | Transcription, speech translation to English, language identification |
| Training data | 680,000 hours of audio: about 65% English, with the non-English audio covering 98 languages (model card) |
| Latest pip release | `20250625` ([PyPI](https://pypi.org/project/openai-whisper/)) |
| Hosted version | `whisper-1` in the OpenAI API, $0.006/minute |

## Whisper models: which one to pick

The open-source release comes in six sizes. This table is from the [README](https://github.com/openai/whisper#available-models-and-languages); speeds are relative to `large`, measured on English speech on an A100 GPU, and your hardware will differ.

| Model | Parameters | English-only version | VRAM needed | Relative speed |
| --- | --- | --- | --- | --- |
| tiny | 39 M | `tiny.en` | ~1 GB | ~10x |
| base | 74 M | `base.en` | ~1 GB | ~7x |
| small | 244 M | `small.en` | ~2 GB | ~4x |
| medium | 769 M | `medium.en` | ~5 GB | ~2x |
| large | 1550 M | none | ~10 GB | 1x |
| turbo | 809 M | none | ~6 GB | ~8x |

The README lists `turbo` at 809 M parameters; the [model card](https://github.com/openai/whisper/blob/main/model-card.md) gives 798 M.

How to choose:

- **turbo** is the default in the command-line tool. The README describes it as an optimized `large-v3` with faster transcription and minimal loss of accuracy. It is the right starting point for most transcription.
- **turbo is not trained for translation.** If you pass `--task translate`, it returns the original language. Use `medium` or `large` to translate into English.
- **English-only `.en` models** tend to do better on English audio than the multilingual model of the same size, especially at `tiny` and `base`.
- **Accuracy varies a lot by language.** The README links a word-error-rate breakdown by language for `large-v3` and `large-v2`. Check your language before you commit to Whisper.

## How to run Whisper locally

You need Python (the README says 3.8 to 3.11 are expected to work), PyTorch, and the `ffmpeg` command-line tool. A GPU helps a lot; on CPU, the larger models are slow.

1. Install ffmpeg:

```bash
# macOS (Homebrew)
brew install ffmpeg

# Ubuntu or Debian
sudo apt update && sudo apt install ffmpeg
```

2. Install Whisper:

```bash
pip install -U openai-whisper
```

If the install fails while building `tiktoken`, the README says you may need Rust installed.

3. Transcribe a file from the terminal:

```bash
whisper audio.mp3 --model turbo
```

Whisper detects the language from the first 30 seconds. You can set it yourself, and translate with a multilingual model:

```bash
whisper japanese.wav --language Japanese
whisper japanese.wav --model medium --language Japanese --task translate
```

By default the command writes every output format (`txt`, `vtt`, `srt`, `tsv`, `json`). Use `--output_format srt` for subtitles only, and `whisper --help` for the full list of options. Word-level timestamps are available as an experimental flag, `--word_timestamps True`.

4. Or call it from Python:

```python
import whisper

model = whisper.load_model("turbo")
result = model.transcribe("audio.mp3")
print(result["text"])
```

`result["segments"]` holds the timed segments, and `result["language"]` the detected language. The model files download on first use to `~/.cache/whisper`.

## How to use Whisper through the OpenAI API

If you don't want to manage a GPU, the [speech-to-text API](https://developers.openai.com/api/docs/guides/speech-to-text) runs Whisper and newer models for you. As of September 2026:

| Model | Best for | Output formats | Price per minute |
| --- | --- | --- | --- |
| `gpt-transcribe` | General transcription (OpenAI's current recommendation) | JSON | $0.0045 |
| `gpt-4o-mini-transcribe` | Lower cost | JSON | $0.003 |
| `gpt-4o-transcribe` | Transcription | JSON | $0.006 |
| `gpt-4o-transcribe-diarize` | Speaker labels | `diarized_json`, JSON, text | $0.006 |
| `whisper-1` | Subtitles, timestamps, translation to English | `srt`, `vtt`, `verbose_json`, JSON, text | $0.006 |

Per-minute prices are OpenAI's estimates from its [pricing page](https://developers.openai.com/api/docs/pricing); some of these models are billed per token. Limits worth knowing, from the same guide:

- Uploads are capped at **25 MB per file**. Supported types are mp3, mp4, mpeg, mpga, m4a, wav and webm. Longer recordings need to be compressed or split.
- Only `whisper-1` returns `srt`/`vtt` and supports word or segment timestamps through `timestamp_granularities`. `gpt-4o-transcribe-diarize` doesn't support `timestamp_granularities`, but its `diarized_json` segments include start and end times.
- The translation endpoint uses `whisper-1` only and translates into English only.
- `whisper-1` prompts are limited to 224 tokens.

A minimal request that returns subtitles:

```bash
curl https://api.openai.com/v1/audio/transcriptions \
  -H "Authorization: Bearer $OPENAI_API_KEY" \
  -F file="@audio.mp3" \
  -F model="whisper-1" \
  -F response_format="srt"
```

## How Whisper works

Whisper is a standard Transformer encoder-decoder, described in the paper [Robust Speech Recognition via Large-Scale Weak Supervision](https://arxiv.org/abs/2212.04356). Audio is cut into 30-second windows and converted to a log-Mel spectrogram. The encoder processes the spectrogram, and the decoder predicts text along with special tokens that tell it which task to do: identify the language, transcribe, translate to English, or add timestamps. One model replaces what used to be several separate pipeline stages.

![Whisper architecture: 30-second audio windows become log-Mel spectrograms, pass through a Transformer encoder, and a decoder predicts text and task tokens](https://exemplary.ai/img/blog/openai-whisper/asr-summary-of-model-architecture.svg)

Image source: [OpenAI](https://openai.com/index/whisper/)

The difference from earlier systems was the data. Instead of training on small, carefully labelled datasets, OpenAI trained on 680,000 hours of audio paired with transcripts collected from the web, with a large share in languages other than English and a share used for translation to English.

![Breakdown of Whisper's 680,000 hours of training data by task and language](https://exemplary.ai/img/blog/openai-whisper/asr-training-data.svg)

Image source: [OpenAI](https://openai.com/index/whisper/)

## Known limits

- **Hallucinations.** The [model card](https://github.com/openai/whisper/blob/main/model-card.md) warns that output may include text not spoken in the audio. Users report this most often on silence, music or noise (see the GitHub discussion [Hallucination on audio with no speech](https://github.com/openai/whisper/discussions/1606)). Trimming silence or running voice activity detection first helps, and the CLI has a `--hallucination_silence_threshold` option (it requires `--word_timestamps True`).
- **Repetition.** The model card also notes the architecture can get stuck repeating text.
- **Uneven accuracy.** Performance is lower on low-resource languages and can differ across accents and dialects (model card).
- **No speaker labels.** Open-source Whisper gives you text and timestamps, not "who said what". On the API, speaker labels come from `gpt-4o-transcribe-diarize`, not `whisper-1`.
- **turbo won't translate**, and API uploads stop at 25 MB.
- **Not for high-stakes decisions.** OpenAI cautions against using it in high-risk decision-making and against transcribing people recorded without their consent.

## When a hosted tool makes more sense

Whisper is a model, not an app. Running it yourself means installing Python and ffmpeg, picking a model, having a GPU for speed, and then editing the raw text somewhere else. That is a good trade if you need privacy on your own hardware or you are building a product.

If you just need a transcript, subtitles or a summary, a hosted editor is less work. Disclosure: we make Exemplary AI. It transcribes audio and video in 99 languages in the browser, gives you a document-style editor where you can fix words and rename speakers, exports subtitles such as .srt, and can then [translate](https://exemplary.ai/translation), summarize or cut the recording into clips. There is a free plan; see [pricing](https://exemplary.ai/pricing) for limits.

> **Transcribe without the setup**
> Upload a file or paste a link and edit the transcript in your browser.
> [Try Exemplary AI free](https://exemplary.ai/transcription)

## FAQ

### Is OpenAI Whisper free?

Yes. The open-source Whisper code and model weights are free under the MIT License, including for commercial use; you only pay for your own hardware. The hosted `whisper-1` model in OpenAI's API costs $0.006 per minute as of September 2026.

### Which Whisper model is the best?

For English and most transcription, start with `turbo`: it is the CLI default, and the README describes it as an optimized `large-v3` that is much faster with minimal loss of accuracy. Use `medium` or `large` if you need translation into English, and a `.en` model if you only transcribe English on limited hardware.

### Can Whisper identify different speakers?

No. Open-source Whisper and `whisper-1` return text and timestamps without speaker labels. OpenAI's `gpt-4o-transcribe-diarize` API model returns speaker-labelled segments, or you can use a transcription tool with speaker editing.

### Does Whisper work offline?

Yes. After `pip install -U openai-whisper` downloads the package and the first run downloads the model file, transcription runs entirely on your machine.

### What's the difference between Whisper and gpt-transcribe?

Whisper is the open model you can run yourself, and `whisper-1` is its hosted API version with subtitle, timestamp and translation output. `gpt-transcribe` is a newer hosted-only model that OpenAI's speech-to-text guide recommends for general transcription; it returns JSON and costs $0.0045 per minute.

### What is the file size limit for Whisper?

The OpenAI API accepts files up to 25 MB. Running Whisper locally has no fixed file size limit; long files just take longer and need enough memory.

> Exemplary AI is a browser-based AI tool that turns long videos and recordings into short captioned clips, transcripts in 99 languages, subtitles, translations into 116 languages, and written posts such as summaries, blog posts and YouTube chapters. It is built for YouTubers, podcasters and marketing teams; the Free plan includes 500 one-time credits with watermarked video exports. Pricing: https://exemplary.ai/pricing
