Menu
AI Speech-to-Text

Transcribe Audio to Text

Turn any recording into accurate text with Whisper AI. Word-level timing, TXT/SRT/VTT export, and automatic language detection across 95+ languages — MP3, voice memos, interviews, and more.

AI transcription runs on our secure servers.

Whisper-powered speech-to-text with word-level timestamps — export TXT, SRT, and VTT. It costs 1 credit per minute of audio, and a free account includes 6 AI audio credits per month — no card required. Files are processed over TLS and deleted automatically.

How it works

Step 1

Add your audio

Drop in an MP3, voice memo, interview, or any recording. The spoken language is detected for you.

Step 2

Transcribe with Whisper

Whisper AI writes out the words with timestamps down to the word — 1 credit per minute of audio.

Step 3

Export TXT, SRT, or VTT

Grab plain text for reading or timed captions for video. Your files are deleted automatically.

Frequently asked questions

How accurate is the transcription?

Transcription runs on Whisper AI with word-level timing, so it holds up well on clear speech — podcasts, interviews, voice memos, and lectures. It detects the spoken language automatically across 95+ languages, so you don't have to pick one. Clean audio transcribes most accurately, so trim out noise or run AI cleanup first for the best result.

What formats can I export?

Download plain text (TXT) for a clean read, or timed captions as SRT and VTT for video subtitles. Word-level timestamps are captured too, so captions line up tightly with the audio instead of drifting across long lines.

Is my audio private?

Transcription runs as an isolated server job over TLS, and your audio and transcript are deleted automatically after processing. Nothing is kept or used to train anything.

How much does it cost?

Transcription costs 1 credit per minute of audio. A free account includes AI audio credits every month — enough to transcribe a short recording, no card required — and Pro plans add a larger monthly allowance.

Cleaner audio transcribes better — try AI noise removal or remove silence first, or open the full Audio Studio.