Fal Speech-to-Text

Free.ai · stt · ~500 ഒരു സന്ദേശത്തിനും അടയാളങ്ങള്‍ നല്‍കുക minute

ഒരു ഓഡിയോ അല്ലെങ്കില്‍ വീഡിയോ ഫയല്‍ താഴെയിടുക അല്ലെങ്കില്‍ ഒരു യുആര്‍എല്‍ പകര്‍ത്തുക

~500 ഒരു സന്ദേശത്തിനും അടയാളങ്ങള്‍ നല്‍കുക minute
നമ്മുടെ GPUS-ല്‍ നിന്നും സ്വതന്ത്രമായി ഓടും. കൂടുതല്‍‌‌കൂട്ടുക Fal Speech-to-Text →

Fal Speech-to-Text is a സംസാരത്തിനുള്ള വാചക മാതൃക. ബാഹ്യമായ മോഡലുകള്‍ വഴി രഹസ്യമാക്കി — ~500 ചിഹ്നങ്ങള്‍ ഒരു മിനിറ്റ് (50% ല്‍ കൂടുതല്‍ വിലയ്ക്കു് മുകളില്‍ മാര്‍ക്ക് ചെയ്യുക).

_API വഴി ഉപയോഗിക്കുക

OpAI- യോജിപ്പുള്ള റേസ്റ്റ് API. ഒരു കീ ഉണ്ടാക്കൂ, ഈ മോഡലിനെ സെക്കന്‍ഡുകളില്‍ തന്നെ വിളിക്കുക.

curl -X POST https://api.free.ai/v1/stt/ \
  -H "Authorization: Bearer sk-free-..." \
  -H "Content-Type: application/json" \
  -d '{"model":"premium/speech-to-text","audio_url":"https://..."}'
എപിഐ സഹായക്കുറിപ്പുകള്‍ API കീ ലഭ്യമാക്കുക

പലപ്പോഴും ചോദിക്കപ്പെടുന്ന ചോദ്യങ്ങൾ

Fal Speech-to-Text transcribes spoken audio into text. Upload an MP3, WAV, M4A, or video file and Fal Speech-to-Text returns the full transcript plus optional SRT/VTT subtitles with timestamps.

Fal Speech-to-Text handles dozens of languages - Whisper-family models cover 90+, others vary. Pick "auto-detect" or specify the language for highest accuracy.

Word-error rate is 5-10% on clean English audio, 10-20% on noisy or accented audio. Large variants of the same architecture do meaningfully better on hard cases - pick larger when the audio is rough.

അതെ, ഓരോ ഭാഗവും തുടങ്ങുക/ തുടങ്ങുക. SRT അല്ലെങ്കില്‍ VTT എന്നോ കാലോപകരണം നിങ്ങളുടെ വീഡിയോയില്‍ നേരിട്ട് ലഭ്യമാക്കുക.

Fal Speech-to-Text is a premium transcription engine. About ~500-1,500 tokens per minute of audio, shown live before you generate. Self-hosted transcription runs free on your 30,000-token daily pool; premium models are pay-as-you-go, with top-ups from $1.

MP3, WAV, M4A, FLAC, OGG, plus video (MP4, MOV, WebM) - we extract the audio. Max 500 MB per upload. Longer files? Split with /audio/cut/ or use /v1/stt/batch/.

ശബ്ദകര്‍ത്താവ് (അടിസ്ഥാനങ്ങള്‍) എന്നറിയപ്പെടുന്ന ഒരു ഷീറ്ററിങ്ങ് ആണ് - /tranchannel /. Fal Speech-to-Text- ല്‍ "ഡിആറസി" എന്ന പരമ്പരയെ കൈകാര്യം ചെയ്യുന്നു; ഓരോ ഭാഗവും സ്പീഡര്‍ 1 / 2 / etc.

അതെ / bache/subject files. ഓരോ ഓഡിയോ- ശേഖരവും // കാഷ്/ സ്വീകരിയ്ക്കുക. /acam- config/? Tab=മുഴുവന്‍ പേരു്‌ കൂടെയുള്ളതാണു്. അറ- വൃക്ഷം സംരക്ഷിക്കുന്നതിനായി API ഉപയോഗിക്കുന്നു.

Yes - POST your audio to /v1/stt/transcribe/ with model="Fal Speech-to-Text". Returns JSON with text + segments + word-level timestamps. /api/ has the full reference.

GPUS-ല്‍ സ്വയമേയ മോഡല്‍സ് ഓഡിയോ സൂക്ഷിക്കുന്നു; ഒരു DPA വഴി കടന്നു പോകുന്നു. അതു് ഒരു പങ്കാളിത്ത-അന്‍-ആങ്കണത്തിനു ശേഷം (24HAn- 7d--ഇന്‍) നീക്കം ചെയ്യുന്നു. നിങ്ങളുടെ ഇന്‍പുട്ടുകളില്‍ ഞങ്ങള്‍ പരിശീലിക്കുന്നില്ല.

അതേ, Free.ai - ത്തോളം പേർക്ക് റെക്കോർഡ്‌ ചെയ്‌തിരിക്കുന്ന ഓഡിയോ - യുടെ (നിങ്ങളുടെ സ്വന്തം റെക്കോർഡ്‌, ലൈസന്‍സ്‌ ചെയ്‌തിരിക്കുന്ന വിവരങ്ങൾ അല്ലെങ്കിൽ സമ്മതത്തോടെ) വാണിജ്യ സാങ്കേതിക വിദ്യകൾ ലഭ്യമാണ്‌.

യഥാര്‍ത്ഥ- സമയം ആവശ്യത്തിനു് 0.05- 02× ആണ്. — 3-12 മിനുട്ടില്‍ 60-മിനിട പോസ്റ്റ് ട്രാന്‍ ബോര്‍ഡുകള്‍. പ്രീമിയം മോഡല്‍സ് പലപ്പോഴും വേഗത്തില്‍ പൂര്‍ത്തിയാക്കുന്നു. റെയിം ബട്ടണ്‍ ടാബ് അടയ്ക്കാന്‍ ഉപയോഗിയ്ക്കുക.

സ്നേഹം Free.ai, കൂട്ടുകാരോട് പറയൂ!

ഈ താള്‍ അനുബന്ധപ്പെടുത്തുക