Transcription models and languages

updated

Oyma transcribes meetings on your Mac with Whisper speech models. Words appear while the meeting runs, and the full transcript is saved in the note. Models download once, on demand, and then work offline.

Live transcription

With Live transcription on (the default, in Settings › Meetings), text appears as people speak. Your audio is transcribed on your Mac, using the graphics processor when it can.

The first time you record, Oyma shows Set up live transcription before anything starts:

  • Download and record downloads the selected model once, then starts recording.
  • Record without live transcription records straight away. The audio is saved and can be transcribed after you stop.

The download never starts in the middle of a meeting. If live text is missing when you stop, Oyma transcribes the saved recording instead.

Choose a speech model

Open Settings › Meetings and pick a Speech model:

ModelDownloadNotes
Base148 MBReady a few seconds after you start recording. Good accuracy
Small (default)488 MBNoticeably better with names and punctuation. Takes minutes to prepare the first time
Large v3 Turbo1.6 GBMost accurate. Takes the longest to prepare and uses the most memory

The Model on this device row shows whether the selected model is downloaded. Use Download to fetch it ahead of your next meeting, or Delete to free the space. Models are stored with the app’s data on your Mac, not in your vault.

“Minutes to prepare” happens once per model, the first time it loads on your Mac. If live text lags in your first meeting with Small or Large v3 Turbo, that’s usually why. Base is the quickest choice if you need text immediately.

Other live transcription settings

All in Settings › Meetings:

  • Microphone: System default follows your macOS input setting, or pick a specific device. Changes apply to the next recording.
  • Transcribe my microphone: when off, your voice is still recorded but only the other participants are transcribed live.
  • Transcribe other participants: captures and transcribes the sound your Mac plays. It needs the Screen & System Audio Recording permission and macOS 14.2 or later. See Recording permissions on macOS.

Languages

Set Spoken language under Settings › Meetings › Transcription. Auto-detect works for most meetings. Choosing the language improves accuracy, especially for short recordings. The options are:

Auto-detect, English, Spanish, French, German, Italian, Portuguese, Dutch, Russian, Chinese, Japanese, Korean, Arabic and Hindi.

Use headphones

Oyma doesn’t cancel echo. Without headphones, your microphone hears the other participants through your speakers while call audio is also being captured, so their words can appear twice in the transcript. Wear headphones or earbuds for clean transcripts.

Transcribe an existing audio or video file

You can transcribe recordings made elsewhere, such as a Zoom recording, a voice memo or a lecture.

  1. Press ⌘O and run >Transcribe an audio or video file….
  2. Choose the file. Supported types include MP3, M4A, WAV, AAC, FLAC, OGG, WebM, MP4, MOV and M4V.
  3. Oyma copies the file into your vault, creates a meeting note, transcribes it, and adds a summary if one is set up.

Transcribing a file, or re-transcribing a meeting with >Re-transcribe current meeting, uses the Base model. If Base isn’t downloaded you’ll see “Recording saved. Set up the speech model…”. To fix it, select Base under Speech model, click Download, then switch back to the model you prefer for live meetings.

Speed and accuracy

  • Transcription runs fastest on the graphics processor. If it can’t use it, Oyma falls back to the processor, which is slower: live text can land several seconds behind speech. Run check in Settings › Meetings › Diagnostics shows Ready (GPU) or Ready (CPU — slower).
  • Bigger models are more accurate and slower to start. Small is a good balance for most meetings.
  • Transcripts aren’t split by speaker, and lines have no timestamps.

Local-only mode

With Local-only mode on, models don’t download automatically. Download the model once from Settings › Meetings, which uses the network with your consent, and transcription then runs fully offline.

If transcription is slow or the text lags, see Troubleshooting.

Questions

Does transcription send my audio anywhere?

Not by default. Transcription runs on your Mac with an open-source Whisper model. Audio only goes to a cloud service if you set Transcription to Cloud in Settings › Privacy and add your own OpenAI key.

Which languages are supported?

Auto-detect plus 13 languages: English, Spanish, French, German, Italian, Portuguese, Dutch, Russian, Chinese, Japanese, Korean, Arabic and Hindi.

How much disk space do the models use?

Base is 148 MB, Small 488 MB and Large v3 Turbo 1.6 GB. Nothing is bundled with the app; a model downloads only when you need it, and you can delete it in Settings › Meetings.

Your notes, in plain Markdown.

Free during the private beta. Apple silicon Macs, macOS 14 or later.

One email when your invite is ready. No newsletter.