Transcribe audio and video files on your Mac
In addition to recording live calls, Oyma can transcribe existing audio and video recordings stored on your Mac. Whether you have voice memos recorded on your phone, lecture recordings, customer research interviews, podcast drafts, or archived video conference clips, you can process them locally with embedded Whisper models without uploading megabytes of private recordings to cloud transcription services.
All media decoding, speech recognition, and formatting execute 100% locally on your machine, leveraging Apple Silicon’s Neural Engine and Metal Performance Shaders for lightning-fast turnaround.
Supported media formats and containers
Oyma uses native macOS AVFoundation and CoreMedia frameworks to demux, decode, and resample audio streams directly in memory. There is no need to manually convert files with third-party tools before importing.
Audio formats
- Apple Voice Memos (
.m4a,.aac): Drag and drop directly from the macOS Voice Memos app or Finder. - Waveform Audio (
.wav): Uncompressed 16-bit, 24-bit, or 32-bit float PCM audio from professional field recorders. - MPEG-3 Audio (
.mp3): Standard CBR and VBR encoded MP3 audio streams. - Free Lossless Audio Codec (
.flac): High-fidelity lossless multi-channel recordings.
Video formats
- QuickTime Movies (
.mov): Screen recordings, iPhone camera footage, and Final Cut exports. - MPEG-4 Video (
.mp4): Zoom cloud recordings, Meet downloads, and standard video conference archives.
When you import a video container, Oyma demuxes the primary audio stream into 16kHz mono PCM directly in RAM. The original video file on your disk remains completely untouched and unmodified.
Performance benchmarks on Apple Silicon
Local Whisper speech recognition is heavily optimized for Apple unified memory architecture. Because the CPU, GPU, and Neural Engine share access to the same high-bandwidth RAM pool, memory transfers are virtually instantaneous.
The table below provides typical transcription throughput across common Apple Silicon chips:
| Model Size | Model Disk Footprint | Processing Speed (M2 / M3) | 60-min Audio Duration | Recommended Use Case |
|---|---|---|---|---|
| Whisper Base | ~140 MB | 12x – 16x real-time | ~4 minutes | Quick drafts, clean podcasts, fast audio memos |
| Whisper Small | ~460 MB | 6x – 9x real-time | ~7 minutes | General meetings, multi-speaker interviews |
| Whisper Medium | ~1.5 GB | 3x – 5x real-time | ~14 minutes | Accented English, moderate background ambient noise |
| Whisper Large v3 Turbo | ~1.6 GB | 4x – 6x real-time | ~11 minutes | High-accuracy technical vocabulary & multi-language |
Step-by-step import workflow
You can transcribe files through multiple entry points depending on your preferred workflow:
Method 1: The Command Palette
- Open Oyma and press ⌘O to open the Command Palette.
- Type
>Transcribe an audio or video file…and press Enter. - Use the native macOS file dialog to pick an audio or video file.
- In the preflight configuration window:
- Language: Select the spoken language, or choose Auto-detect if the language is unknown or mixed.
- Model size: Choose between Base, Small, or Large v3 Turbo.
- Target vault folder: Choose where the resulting Markdown note should be saved (defaults to
Recordings/or your configured vault root).
- Click Start transcription.
Method 2: Drag and drop
- Drag any supported audio or video file from Finder into an open note or folder in Oyma‘s sidebar.
- A contextual prompt appears: “Import as asset” or “Transcribe to note”.
- Click Transcribe to note to launch the transcription engine immediately.
Background execution and queue management
Transcription runs asynchronously in a dedicated background worker process:
- Zero editor lag: The main text editing thread, graph rendering, and search indexing remain entirely smooth and responsive while Whisper executes.
- Dock and status indicator: A progress pill appears in Oyma‘s status bar indicating percentage complete, current playback position, and remaining time.
- Batch processing: You can queue multiple audio files consecutively. Oyma processes them sequentially to avoid overloading RAM or exhausting system thermal headroom.
Note layout and transcript structure
Once transcription finishes, Oyma creates a structured Markdown document in your vault:
---
title: "Interview with Product Strategy Team"
source_file: "Strategy-Sync-2026-10-08.m4a"
duration: "00:42:15"
language: "en"
model: "whisper-small"
date: 2026-10-08
tags:
- transcription
- interview
---
## Audio playback
![[Strategy-Sync-2026-10-08.m4a]]
## Transcript
[00:00:05] Good morning everyone, thanks for joining the roadmap session.
[00:00:22] We want to review the customer feedback from the beta rollout.
[00:01:14] The main takeaway was that local performance exceeded cloud latency by 4x.
Interactive audio scrub chips
Clicking any timestamp chip (such as [00:01:14]) in the editor immediately jumps the embedded audio player to that exact millisecond in the recording. You can listen back to verify ambiguous phrasing or add manual annotations without scrub hunting.
Generating summaries and action items
Once the transcript is generated, you can leverage Oyma‘s AI engine to extract insights:
- Click the Summarize button at the top of the transcript note.
- Choose your preferred summary template: Executive Brief, Action Items, or Thematic Outline.
- Select your AI engine:
- Run fully on-device with local Ollama models.
- Or use frontier models via your own Anthropic / OpenAI API keys.
- The generated summary is inserted above the transcript as clean, editable Markdown bullet points.
Troubleshooting media transcription
| Issue | Root Cause | Solution |
|---|---|---|
| Unsupported codec error | Rare proprietary video container or DRM-protected stream. | Open file in QuickTime Player, choose File → Export As → Audio Only, and import the exported .m4a. |
| High hallucination or silence loops | Prolonged background noise, music, or long periods of silence. | Enable Voice Activity Filter (VAD) in Transcription Settings before running Whisper. |
| Slow processing speed | Running Large models on Intel Macs or base 8GB machines under memory pressure. | Switch to Whisper Small or Base in Settings for substantially faster processing. |
| Inaccurate foreign language translation | Transcription language was left set to English default. | Set explicit spoken language in the preflight dialog instead of relying on auto-detection. |
To understand how Whisper model sizes trade off memory versus phonetic accuracy, review our guide to transcription models. For information on recording meetings live without third-party bots, see recording a meeting.
Questions
What audio and video file formats does Oyma support?
You can import MP3, WAV, M4A, FLAC, MP4, and MOV files. Video files are demuxed locally to extract the audio stream before running on-device Whisper transcription.
Are imported media files uploaded to the cloud?
No. All media processing and speech recognition run entirely on your Mac's Apple Silicon unified memory using embedded Whisper models. Zero audio bytes leave your device.
How long does it take to transcribe an audio file?
On an Apple Silicon Mac, local Whisper models typically process audio faster than real-time. A 30-minute podcast or recording usually takes 2 to 5 minutes using the Base or Small model.
Where does the transcribed text land?
The finished transcript lands as a clean Markdown (.md) note in your vault, complete with audio timestamp chips, meeting metadata, and an optional AI summary.
Can I export transcripts to subtitle formats like SRT or VTT?
Yes. From the transcript note menu, you can export synchronized captions in standard SubRip (.srt) and WebVTT (.vtt) file formats for video editing.