Free Chinese Transcription

Transcribe Chinese audio and video to text with AI. Fast, accurate, and free.

How It Works

  1. Go to the Free.ai Transcriber
  2. Upload your Chinese audio or video file
  3. Our AI automatically detects Chinese and transcribes it
  4. Download your transcript as text or SRT subtitles

Chinese Transcription Features

  • Powered by faster-whisper (MIT licensed)
  • Automatic Chinese language detection
  • Supports MP3, WAV, MP4, M4A, FLAC, and more
  • Timestamps and subtitle export (SRT)
  • No file size limits on paid plans
  • Private and secure -- files are deleted after processing

Language Details

LanguageChinese
ISO Codezh
AI Modelfaster-whisper
PriceFree

More Languages

View All Languages

FAQ

Whisper large-v3-turbo lands in its top accuracy tier on Chinese — under 7% word error rate on standard benchmarks. In practice that means clean studio audio comes back near-perfect, and conversational audio is usable with minimal cleanup. (Tier A, under 7% word error rate on benchmark sets - we publish honest WER tiers rather than marketing claims.)

Yes - Chinese transcription draws from your daily free token pool first. Audio costs about 50 tokens per minute, so the anonymous daily pool covers a few hours of audio per day. Signed-in accounts get a larger 30,000-token daily pool. Past that, transcription is pay-as-you-go, with token top-ups from $1.

Pass language=zh for Mandarin (the default — simplified or traditional output depending on the source). For Cantonese use language=yue if your audio is Hong Kong / Guangzhou speech; Cantonese transcribed as zh will produce a Mandarin-orthography approximation that loses tones and slang.

MP3, WAV, M4A, FLAC, OGG, OPUS, and WEBM are accepted directly. For video (MP4, MOV, MKV) we extract the audio track server-side before sending it to Whisper - you do not need to convert anything yourself. Same pipeline regardless of source language, including Chinese.

Anonymous uploads cap at roughly 500 MB per file. Signed-in accounts go up to 2 GB. Duration is not a hard limit - long files are chunked automatically (30-second windows with overlap) and stitched back into a single transcript with continuous timestamps. Multi-hour Chinese recordings (podcasts, full lectures, meetings) work fine.

Yes - speaker diarization is on by default for every Chinese transcript. The output is segmented as Speaker 1 / Speaker 2 / Speaker 3 with timestamps, so interviews, panel discussions, and multi-party meetings come back labeled. Diarization runs on a separate model and works the same across all languages we support.

Yes - paste the URL into /transcribe/youtube/ for YouTube or /transcribe/podcast/ for podcast feeds (Apple, Spotify, RSS). We download the audio, run it through Whisper with language=zh, and return the transcript with timestamps and speaker labels. Typical Chinese content: podcasts, lectures, interviews, and long-form YouTube content in Chinese are the most common workloads we see.

Whisper costs about 50 tokens per minute of audio, so a one-hour recording is ~3,000 tokens. Most users never spend anything - the free daily pool of 30,000 tokens covers short clips, voice notes, and one-off podcasts. Beyond it, transcription is pay-as-you-go, with token top-ups from $1.

Yes - both segment-level (every ~10-30 seconds) and word-level timestamps are available. Word-level is the default for VTT/SRT subtitle export so the captions sync line-by-line. On the API set timestamps="word" in the request body. Chinese transcripts are returned in native Han characters (UTF-8) — simplified or traditional depending on the source audio and ISO code.

Yes. POST audio (multipart/form-data, field name "file") to /v1/transcribe/ with language=zh - or omit the language parameter to let Whisper auto-detect. Returns JSON with the transcript, segments, timestamps, and speaker labels. Full reference and SDK snippets at /api/.

Yes - once transcription finishes, click Translate or paste the text into /translate/. Chinese pairs with every other language we support (200+). For meeting minutes pipe the transcript through /summarize/; for dubbing send it to /voice/tts/ to render audio in the target language.

Whisper is trained on 680K hours of noisy real-world audio, so Chinese transcription is robust to background noise, music beds, and phone-quality recordings. Severe clipping or multiple overlapping speakers will still hurt accuracy. If a transcript comes back unusable, email contact@free.ai with the file - we will refund the tokens and look at whether a different engine handles your audio better.

Love Free.ai? Tell your friends!

Rate this page