No account required

Turn any audio into text

Upload your recording and get it back transcribed: with timestamps, every voice identified, and a summary. No account, no installation.

Drag your file in here

MP3, WAV, M4A, MP4, MOV… and any other audio or video, whatever the size

  • Every voice identified, and named where possible
  • A timestamp on every line: click it and listen to that moment
  • A summary with the agreements, pending tasks and key points
  • MP3
  • WAV
  • M4A
  • AAC
  • OGG
  • FLAC
  • WMA
  • AIFF

About this tool

Transcribing an hour of audio by hand can take several hours of work. With this tool it takes about three minutes, and the result isn't an unstructured block of text: it's organized by speaker, with a timestamp on every sentence, a built-in search box, and it's editable right in the browser.

It accepts any common format (MP3, WAV, M4A, AAC, OGG, FLAC), with a limit of 50 GB per file. The language is detected automatically.

Features

  • Detects the language automatically, out of more than 90 available.
  • Identifies the voices and can assign names to them from context.
  • A search box that reaches every transcript you have, not just this one.
  • Your files are stored on European Union servers and get deleted automatically.

How it works

  1. 1

    Upload the audio

    Drag it in or pick it from your computer. One file at a time, whatever the size.

  2. 2

    We transcribe it

    The process takes around 5% of the file's length: a sixty-minute meeting is ready in about three.

  3. 3

    Review and download

    Fix anything by clicking on it and download as DOCX, PDF, SRT, CSV or JSON.

Use cases

Interviews

Questions and answers separated by speaker, with a timestamp to verify each quote.

Meetings

The summary captures the decisions and the tasks, so nobody has to write the minutes.

Podcasts

Get the show notes, the article and the subtitles in a single pass.

In detail

What determines transcription quality

Three factors, and accent isn't one of them. The first is microphone distance: a phone in the centre of the table picks up far more room noise than the same phone placed near whoever is speaking, and that noise ends up as incorrect words. The second is overlapping speech: when two people talk at once, no system can distinguish both accurately.

The third is recording-specific vocabulary. The most common errors involve company names, surnames and industry acronyms, precisely because they appear nowhere else. These are corrected in a couple of minutes by editing the text, and that's where real accuracy is gained.

What to do once you have the text

The most-used feature is searching within the text: locating the moment a topic was discussed without listening to the recording again. If the voices are identified, you can also filter by speaker, which significantly cuts down review time for an hour-long interview.

The next step is exporting the result: a document to share, an SRT for the video, a spreadsheet to analyse speaking time per participant. None of those exports requires transcribing again: export in whatever format you need, as many times as you need.

Frequently asked questions

Need to transcribe a video?

Upload it and get the text with timestamps, each voice identified and a summary. No account needed to start.