How it works

From recording to text, step by step

The process has three steps. Here's each one in detail: what we guarantee, and what to keep in mind.

The three steps

  1. 01

    Upload your file

    Drag it onto the page or select it from your device. It uploads directly to our storage, bypassing the app, so a two-hour recording never stalls halfway.

    • The progress bar shows the real data uploaded.
    • You can close the tab: the upload continues and your text will be waiting.
    • Processing doesn't start until the file has fully uploaded.
  2. 02

    We process the audio

    We measure the real duration, detect the language, separate the voices and generate the text with an exact timestamp for every word.

    • Speakers start out as «Speaker A», «Speaker B»; rename them and the whole text updates.
    • Timestamps are word-level, not paragraph-level, which is why the player stays in sync.
    • The progress you see reflects the actual work happening.
  3. 03

    Review and download

    Edit any line directly on the text. Previous versions are kept, and the search index picks up your changes.

    • Search by content, not just by file name.
    • Download as DOCX, PDF, SRT, CSV or JSON, with subtitles ready to use.
    • If it was a meeting, generate a summary with the agreements and action items.

Under the hood

The technical decisions that matter

Three choices that explain the difference between a usable transcript and an unstructured block of text.

We measure duration ourselves

Minutes deducted come from measuring your file directly, not from the figure the provider reports. If they disagree, your file wins.

We keep the original result

We store what the model produced before any further processing. If we improve the system later, we can regenerate old transcripts at no extra cost.

Each speaker has a single record

Their name is stored once, not repeated on every line. That's why renaming is instant and there's never a mismatch between lines.

Formats and limits

What we accept

Audio

MP3, WAV, M4A, AAC, FLAC, OGG, OPUS

Video

MP4, MOV, WEBM, MKV, AVI

Max file size
50 GB per file
Max duration
10 hours per recording
Languages
90+, detected automatically
Speakers
Up to 32 per recording

Video is fine: we only extract the audio, at no extra cost.

Accuracy

What accuracy to expect

With clean audio in Spanish or English, accuracy runs around 98% word for word. With background noise, overlapping voices or strong accents, it's usually between 92% and 96%.

What helps

  • Keeping the microphone close to whoever is speaking.
  • Avoiding people talking over each other.
  • Setting the language manually if you already know it.

What hurts accuracy

  • Recording with the phone far from the voices.
  • Music or television playing in the background.
  • Recording in a noisy environment.

Your recordings

Where your data is stored

Fully within the European Union

Your recording and transcript are processed and stored on European servers.

Audio deleted automatically

After 14 days without a paid plan, or 90 days with one. The text is kept for as long as your account exists.

Your content stays yours

We never use your recordings to train models. Delete a transcript and its audio is removed that same day.

Try it with your own recording

Get a transcribed sample without needing a card.