Speech to Text

Commercial use OK 380+ models No watermark No sign-up needed
Model:
+ GPT-5, Claude, Gemini
Upload an audio or video file - or paste a URL - and get a clean transcript with timestamps. Speaker diarization, SRT/VTT subtitle export, 100+ languages with auto-detect. Cost scales exactly with clip length. Powered by Whisper large-v3 and Parakeet (self-hosted), plus premium Wizper and ElevenLabs STT.

Drag and drop audio/video, or click to browse

MP3, WAV, MP4, WebM, M4A - up to 500MB

- - -
Whisper large-v3 - 99 languages, best-in-class accuracy.
Token estimate for this clip
-
YouTube, Instagram, TikTok, Spotify, and 1,300+ platforms
URL transcription cost is based on the clip's actual duration - we quote after download. Expect ~500 tokens/minute on Whisper.
Recording: 0:00

Real-time transcription using your microphone

Transcript

Transcribing your audio...

This may take a moment for longer files.

What people transcribe with Free.ai

Interviews + podcasts

Diarization labels every speaker. Export SRT straight into your video editor, or plain text for an article writeup.

Auto captions + subtitles

Upload a YouTube upload or TikTok, pick SRT or WebVTT, and burn the subtitles on with /video/subtitle/. One-stop caption workflow.

Meeting notes

Upload a Zoom/Teams recording - get transcript + speaker labels. Pair with /write/summarize/ for bullet-point minutes.

Lectures + lessons

Transcribe a 90-minute lecture, then use /study/flashcards/ or /write/summarize/ to turn it into study material.

Foreign-language audio

Whisper auto-detects 99 languages. Transcribe in the original, then send the text through /translate/ to jump languages.

Legal + medical

Timestamps, speaker labels, JSON export with every word's start/end time - accurate court-reporter or clinical-note prep.

How Free.ai transcription compares

What you get Free.ai Otter.ai Descript Rev.com
Free daily usage5K+ tokens/day300 minutes/mo1 hr/month-
EngineWhisper large-v3, ParakeetProprietaryProprietaryHuman + AI
Languages99English-focused2230+
Speaker diarization
SRT / VTT exportPaidPaid
Public APILimitedLimited
Live streaming STT (free) Paid--
Sign-up requiredNoYesYesYes
Competitor figures reflect publicly listed free tiers as of 2026. Check each provider for current plans.
Advanced options
Result
Tokens running low. Get More Tokens
Want better results? Premium models (GPT-5, Claude, Gemini) deliver higher quality. View Plans

❤️ Love Free.ai? Tell your friends!

Sign up to get a referral link and earn 30,000 tokens per friend.

Want more? Sign up free for 30K tokens/day
Sign Up Free

Processing your request...

Best free speech to text tool. Upload MP3, WAV, MP4 or record live. Auto-detect language. Speaker diarization. No sign up required.

How to Use Speech to Text

1
Enter your input

Type text, upload a file, or describe what you want. No account needed.

2
Click generate

Our AI processes your request in seconds using the best open-source models.

3
Download & share

Download, copy, or share your result. Free for personal and commercial use.

Use this tool via API

Automate this tool from your own code. OpenAI-compatible REST endpoint, Bearer-token auth, no extra SDK required. Token costs match the web interface.

curl -X POST https://api.free.ai/v1/stt/ \
  -H "Authorization: Bearer sk-free-..." \
  -H "Content-Type: application/json" \
  -d '{"file": "@audio.mp3", "language": "auto"}'

Speech to Text - FAQ

Free.ai offers Whisper-powered speech to text with excellent accuracy, 99 languages, subtitle export, speaker detection, and live mic capture - completely free.

Upload an audio or video file (MP3, WAV, MP4, M4A), click Transcribe, and get accurate speech to text in seconds. Or record live from your microphone.

Yes. Paste any YouTube URL in the URL tab and the speech to text tool extracts the audio and converts it. Works with Instagram, TikTok, Spotify, and 1,300+ platforms.

Yes. Auto-detect or select from 99 languages. Our speech to text handles accents, background noise, and mixed-language audio well.

Yes. Select multiple audio files at once - each is sent through speech to text with progress tracking and the results are downloadable separately or combined.

Yes. The speech to text API at /api/ is OpenAI-compatible. Upload audio programmatically and receive JSON with the transcript, language, and timestamps.

Yes. Toggle Speaker Detection before uploading and the speech to text output is labelled per speaker (Speaker 1, Speaker 2…). Adds 50% to token cost.

Speech to text accepts files up to 500MB per upload. For multi-hour content, split the audio into chunks first.

Very accurate for clear audio - typically 95%+ word accuracy in English with our Whisper large-v3 backend. Quality depends on audio clarity, accent, and background noise.

Yes. The transcript is fully editable in-place. Fix errors, reformat, and copy/download as TXT, SRT, or VTT.

Yes. Audio is processed on our own GPUs and deleted after speech to text completes. Nothing is stored long-term, shared, or used for training.

Yes. Upload an audio or video file in /chat/ and ask the AI to transcribe it - combine speech to text with follow-up questions and summarization in one workflow.

Sign up free - 30,000 tokens/day

Create Free Account

No credit card required

How would you rate this tool?

Love Free.ai? Tell your friends!