-
1Pick a recording
Open Transcribe file in the sidebar and press Browse. Any common audio or video file works — mp3, m4a, wav, wma, flac, ogg, opus, aac, aiff, mp4 and more. Lectures, phone recordings, digitized tapes, meeting audio: no length limit, and nothing is uploaded anywhere.
The transcription runs on its own engine, so your push-to-talk key stays instant while a long file works in the background. Feel free to switch tabs — a toast tells you when it is done. Queue several files and they run one after another.
-
2Set the three options before you start
Under “Before it starts” you will find three pills. Switch them on or off per file:
- Enhance audio — levels and cleans the sound first. Quiet or distant speakers in interviews, lectures and phone recordings come through much better. Costs a few seconds per five minutes of audio and never makes things worse — if it cannot help, the original audio is used.
- Detect speakers — labels who said what. Ideal for meetings, interviews and calls. A recording with a single voice comes back as a plain transcript, without pointless labels.
- Clean up with AI — after transcribing, fixes misheard words and punctuation with your own AI key. Off until you add a key in Settings; see the AI clean-up guide.
-
3Watch the live timeline
Each phase shows up as its own row with a progress bar: read the file → enhance audio → transcribe (with a percentage) → detect speakers → clean up with AI. Text starts streaming into the page while the recording is still being processed, so you can begin reading right away.
Speaker detection runs alongside the transcription rather than after it, so it adds very little to the total time.
-
4Read along with the recording
When it finishes, the title reads something like “interview.mp3 · 12:41 · 3 speakers”, and a player appears with a waveform, elapsed time and playback speeds of 0.75×, 1×, 1.25×, 1.5× and 2×. Press Space to play or pause.
Click any word, or any paragraph timestamp, and the recording jumps to that exact moment — the fastest way to check a doubtful word against the original voice.
-
5Name the voices
With speaker detection on, paragraphs are prefixed with “Speaker 1:”, “Speaker 2:” and so on. Click a label, type a name, and every paragraph by that voice is renamed at once. The app also remembers the voice: next time that person appears in a recording, they are labeled by name automatically.
Remembered voices are listed as small chips under the Browse button — each has an × to forget it. Names survive the AI clean-up pass, and you can rename in the cleaned result too.
-
6Copy or save the text
Copy puts the whole transcript on the clipboard; Save writes a .txt file next to your recording (“interview - transcript.txt”). Speaker names are included as labels at the start of each turn. If you ran an AI clean-up, rewrite or summary, that result has its own Copy and Save.