mictoo
36 Languages · Auto-detect · Free

Multilingual Transcription
One platform for 36 languages and mixed-language audio

Choose one of 36 languages or use Auto-detect for bilingual interviews and recordings that switch languages. Review short switches and proper nouns after transcription.

Free
Auto-deleted
36 languages
AI summary

Drop your file here

or click to browse

MP3 · MP4 · WAV · M4A · OGG · WEBM · FLAC  ·  Max 25MB  ·  Max 30 min

180 min · Sign in

Got a bigger file? See how to compress.

Got a longer recording? See how to split.

🇬🇧English
🇫🇷French
🇪🇸Spanish
🇩🇪German
🌐32 more

How multilingual transcription works

1
Upload supported audio
Use MP3, MP4, WAV, M4A, and other common formats, then choose a language or Auto-detect.
2
Whisper detects and transcribes
Whisper large-v3 can follow many language changes, but short switches and names may need review.
3
Translate to your working language
One-click translation between any of the supported languages after transcription.

Example multilingual transcript & summary

TranscriptNotes
Search
00:00Good morning everyone. Let's get started with today's standup.
00:06Aus Berlin: die Serverauslastung war gestern sehr hoch.
00:14From Buenos Aires: acabamos de terminar el sprint de reportes.
00:22Tokyo team: 昨日のバグは今朝修正しました。
00:30Great. Let's sync on next week's release before we wrap.
00:35Thanks everyone. See you tomorrow.
AI Summary

Global team standup across Berlin, Buenos Aires, and Tokyo. Server load spike reported from Berlin. Reports sprint completed by the Argentina team. Yesterday's bug fixed by the Tokyo team this morning. Next-week release sync scheduled before wrap.

Translate to
Translate to English
Export
TXTSRTVTTDOCX
0:00 / 00:38

Interface preview. Upload a file above to use the transcript, translation, and export controls.

Why use Mictoo for multilingual audio?

36 selectable languages
Choose from the languages in the upload widget or use automatic detection for a first pass.
Mixed-language support
Many Spanglish, Franglais, and international-team recordings can be transcribed with Auto-detect.
Auto-detect or explicit choice
Let Whisper detect the language for short clips, or pin the language for cleaner first-pass results.
Translate between 36 languages
After transcription, choose any supported target language from the same 36-language list.

Multilingual audio that works well

Bilingual interviews
International standups
Global podcasts
MOOC lectures
Diaspora research
Cross-border legal

Tips for multilingual accuracy

  • For very short clips, pin the primary language explicitly.
  • For mid-sentence code-switching, use Auto-detect.
  • Review proper nouns and place names, which often cross languages.
  • Translate after transcribing rather than before.

What makes multilingual recognition difficult

Language ID on short clips
Very short clips may be assigned the wrong language.
"Choose the primary language"
Code-switching mid-sentence
Short switches can be missed or normalised, so bilingual output should be reviewed.
"Vamos a la meeting"
Named entities
Proper nouns can look like foreign words. Review them post-transcription.
"Björk / San José"
Non-Latin scripts
Japanese, Chinese, Korean, Arabic, and Cyrillic scripts are supported; review punctuation and names.
"昨日 / вчера"

Language families supported

VarietyExample differencesSupported
🌍European (Romance + Germanic)FR, ES, IT, PT, DE, EN, NL, SV, DA, NO...
🌏East AsianJA, ZH, KO
🌏SlavicRU, UK, PL, CS, SK, BG
🌍Middle EasternAR, HE, TR, FA
🌏South and Southeast AsianHI, BN, TA, UR, ID, TH, VI, MS

Transcribing French

Whisper handles French well, and the errors it does make cluster in three predictable places. Knowing them tells you what to skim when you proofread.

Liaison
Words run together in speech: les amis is spoken le-za-mi. The model usually gets it, but it can attach the linking sound to the wrong word on fast speech
Elision
J’ai, l’homme, qu’il. Apostrophes come back correctly in most cases; check them in proper nouns, where there is no dictionary to fall back on
Nasal vowels
Sans, bon, brun. These carry meaning and are the most common source of a wrong-but-plausible word
Varieties
Metropolitan, Quebec, Belgian, Swiss and West African French all work. Quebec vocabulary such as char for car or courriel for email transcribes fine; heavy joual is where it slips
Worth doing
Set the language explicitly rather than relying on auto-detect if the recording is under five minutes

Transcribing Spanish

The dialect spread is wider than in French, and Whisper is trained across it. What varies is which spelling it picks, not whether it understands.

Voseo
Vos tenes instead of tu tienes. Common in Argentina, Uruguay, Paraguay and much of Central America, and transcribed as spoken
Seseo and distincion
Gracias is spoken with an s across Latin America and with a th in most of Spain. Both come back spelled the same way, which is what you want
Yeismo
Calle sounds different in Madrid, Buenos Aires and Bogota. Spelling is unaffected
Varieties
Castilian, Mexican, Rioplatense, Caribbean, Andean, Chilean and US Spanish. Chilean and Caribbean speech at speed is the hardest case
Worth doing
Review names and place names. Regional vocabulary is handled better than proper nouns

Transcribing German

German gives the model two structural problems that no amount of audio quality fixes, because they are about how the language is written rather than how it sounds.

Compound nouns
Krankenversicherung, Geschwindigkeitsbegrenzung. Whisper mostly joins them correctly, but an unfamiliar compound can come back split into two words
Separable verbs
Ich rufe dich spaeter an. The prefix lands at the end of the clause, sometimes many words from the verb it belongs to
Umlauts and eszett
Returned properly rather than transliterated to ae, oe, ue and ss
Varieties
Standard German plus Austrian and Swiss. Austrian vocabulary such as Jaenner for January or Marille for apricot works; Swiss German dialect, as opposed to Swiss Standard German, is genuinely hard
Worth doing
Proofread technical compounds and any Swiss dialect passages first, they are where the errors concentrate

Frequently asked questions

Which languages are available?

The upload and translation controls currently offer 36 languages, including English, French, Spanish, German, Portuguese, Italian, Dutch, Polish, Russian, Ukrainian, Japanese, Chinese, Korean, Arabic, Hebrew, Hindi, Turkish, and others. See the language picker for the full list.

Does it handle code-switching?

It can transcribe many recordings that switch languages. Use Auto-detect for bilingual conversations and review short switches, proper nouns, and ambiguous phrases.

Can I translate the transcript?

Yes. After transcription, pick any target language and click Translate. GPT-4o-mini handles the translation and it appears alongside the original.

How accurate is auto-detect?

It works best when the recording contains enough clear speech. For short or noisy audio, choose the primary language explicitly and review the first draft.

Is multilingual transcription free?

Yes. Files up to 25 MB can be transcribed anonymously or up to 180 MB when signed in. No watermark and no per-minute fee.

Are my audio files stored?

No. Audio streams to the transcription provider, gets processed once, and is dropped. Transcripts persist only on signed-in accounts.

Transcribe audio in any language
One upload page for 36 languages, automatic detection, mixed-language audio, and translation.
Upload audio in any language

Explore other languages