7,363 open-source and SaaS tools, with GitHub stats refreshed every day.

Whisper

Open source

OpenAI's open-source general-purpose speech recognition model, which handles multilingual transcription, speech translation and language identification.

Open-source alternative to

GitHub stars
110k
Last commit
1 mo ago
Repository age
4 years
Version
v20250625
Licence
MIT
Self-hosted
Yes

About Whisper

Whisper is a speech recognition model from OpenAI, general in purpose, released as open source under the MIT license. It was trained on a large and diverse audio dataset and handles several tasks in one model: transcription in multiple languages, translation of speech, and identifying the spoken language. The repository contains the Python code needed to run it.

It uses a Transformer sequence-to-sequence model trained on several speech processing tasks, including voice activity detection, which are represented together as a sequence of tokens predicted by the decoder. As a result, one model can stand in for what used to take several separate components in a speech pipeline. Special tokens act as task specifiers or classification targets.

The authors trained and tested with Python 3.9 and PyTorch 1.10, and expect the code to work with Python 3.8 to 3.11 and recent PyTorch versions. It can be installed with pip as openai-whisper or directly from the repository, and it relies on a few packages such as OpenAI's tiktoken tokenizer. Links to a blog post, paper, model card and Colab example are provided. It suits developers and researchers who need speech-to-text they can run themselves.

Key features

  • Multilingual speech recognition
  • Speech translation across languages
  • Spoken language identification
  • Transformer sequence-to-sequence architecture
  • Install with pip
  • Model card, paper and Colab example

Good fit for

  • →Transcribing audio and video
  • →Adding speech-to-text to applications
  • →Speech recognition research
Built with
Python
Tags
speech-recognition
speech-to-text
openai
python
transcription
pytorch
translation
ai

Whisper: questions and answers

What is Whisper used for?
Whisper is OpenAI's open-source general-purpose speech recognition model, which handles multilingual transcription, speech translation and language identification. It is a good fit for transcribing audio and video, adding speech-to-text to applications, and speech recognition research.
Is Whisper open source?
Yes. Whisper is open source under the MIT licence. Its source code is on GitHub at openai/whisper and is written mainly in Python.
Is Whisper free?
Yes. Whisper is open source, so the software itself is free to use.
What is Whisper an alternative to?
Whisper is an open-source alternative to Deepgram and AssemblyAI. Other open-source alternatives to Deepgram include whisper.cpp and LocalAI.
Is Whisper actively maintained?
Yes. The most recent commit to Whisper was on 31 August 2026, and the latest release is v20250625, published on 26 June 2025. The project has 110k stars on GitHub.

Open-source alternatives to Whisper

See all

SaaS alternatives to Whisper

See all