Yiddish Tools
Toggle sidebar
Desktop app Download for Windows
Privacy Terms
Log in Register
All guides
Dictation (STT) · 5 min read

Transcribing recordings

Turn a lecture, interview or phone recording into text — with speaker detection, audio clean-up, a read-along player and optional AI tidy-up.

  1. 1
    Pick a recording

    Open Transcribe file in the sidebar and press Browse. Any common audio or video file works — mp3, m4a, wav, wma, flac, ogg, opus, aac, aiff, mp4 and more. Lectures, phone recordings, digitized tapes, meeting audio: no length limit, and nothing is uploaded anywhere.

    The transcription runs on its own engine, so your push-to-talk key stays instant while a long file works in the background. Feel free to switch tabs — a toast tells you when it is done. Queue several files and they run one after another.

  2. 2
    Set the three options before you start

    Under “Before it starts” you will find three pills. Switch them on or off per file:

    • Enhance audio — levels and cleans the sound first. Quiet or distant speakers in interviews, lectures and phone recordings come through much better. Costs a few seconds per five minutes of audio and never makes things worse — if it cannot help, the original audio is used.
    • Detect speakers — labels who said what. Ideal for meetings, interviews and calls. A recording with a single voice comes back as a plain transcript, without pointless labels.
    • Clean up with AI — after transcribing, fixes misheard words and punctuation with your own AI key. Off until you add a key in Settings; see the AI clean-up guide.
  3. 3
    Watch the live timeline

    Each phase shows up as its own row with a progress bar: read the file → enhance audio → transcribe (with a percentage) → detect speakers → clean up with AI. Text starts streaming into the page while the recording is still being processed, so you can begin reading right away.

    Speaker detection runs alongside the transcription rather than after it, so it adds very little to the total time.

  4. 4
    Read along with the recording

    When it finishes, the title reads something like “interview.mp3 · 12:41 · 3 speakers”, and a player appears with a waveform, elapsed time and playback speeds of 0.75×, 1×, 1.25×, 1.5× and 2×. Press Space to play or pause.

    Click any word, or any paragraph timestamp, and the recording jumps to that exact moment — the fastest way to check a doubtful word against the original voice.

  5. 5
    Name the voices

    With speaker detection on, paragraphs are prefixed with “Speaker 1:”, “Speaker 2:” and so on. Click a label, type a name, and every paragraph by that voice is renamed at once. The app also remembers the voice: next time that person appears in a recording, they are labeled by name automatically.

    Remembered voices are listed as small chips under the Browse button — each has an × to forget it. Names survive the AI clean-up pass, and you can rename in the cleaned result too.

    Built for interviews
    Short interjections — a “yes”, a “hm” — are snapped to the nearest real speaker instead of becoming phantom extra voices, so a two-person interview comes back as two people.
  6. 6
    Copy or save the text

    Copy puts the whole transcript on the clipboard; Save writes a .txt file next to your recording (“interview - transcript.txt”). Speaker names are included as labels at the start of each turn. If you ran an AI clean-up, rewrite or summary, that result has its own Copy and Save.

    Private end to end
    Reading the file, enhancing the audio, transcribing and detecting speakers all happen on this computer. Only the optional AI clean-up sends anything anywhere — the transcript text, to the provider you chose, with your own key, when you ask.