Feature · Built for Mac
Transcribe audio & video files
Convert existing recordings and voice memos into structured Markdown.
What this helps you achieve
- Transcribe MP3, WAV, M4A, MOV, and MP4 files directly from disk
- 100% on-device processing with zero cloud upload bandwidth required
- Generate timestamped Markdown records ready for local AI summarization
Not all spoken knowledge originates inside live video conferences. Many of your most consequential conversations occur during in-person executive briefings, conference keynote lectures, customer discovery sessions recorded on handheld devices, or audio voice memos dictated during your daily commute.
Furthermore, teams frequently possess extensive archives of existing MP4 webinar recordings, podcast interviews, and Zoom session archives that remain locked in opaque audio formats, inaccessible to text search and knowledge synthesis.
Traditional transcription services require you to upload multi-gigabyte video files to third-party web servers. This process consumes heavy network bandwidth, triggers long upload queues, and exposes confidential internal discussions to multi-tenant cloud storage.
Oyma eliminates this friction with native on-device file transcription implemented in src/renderer/workspace/Workspace.ts.
With the Transcribe an audio or video file… command, you can convert any existing media file on your Mac into a rich, timestamped Markdown document in minutes—completely locally and completely offline.
Fast local media ingestion on macOS
Transcribing an existing recording in Oyma requires zero complex command-line scripting or audio format conversions:
- Open the command palette (
⌘O) and choose> Transcribe an audio or video file…(or drag any media file into your vault). - Select your file in the native macOS picker. Oyma supports standard formats:
- Audio: MP3, M4A, AAC, WAV, FLAC, AIFF, OGG.
- Video: MP4, MOV, MKV, WebM.
- Choose your preferred Whisper transcription model (Base, Small, or Large v3 Turbo).
- Watch the progress bar advance. Audio streams are demuxed and decoded locally via native macOS AVFoundation APIs without generating bloated intermediate temporary files.
- In minutes, a fresh, beautifully formatted Markdown document appears in your active vault containing the full timestamped transcript, linked to the source audio.
Because the process runs entirely within macOS system frameworks, you avoid the tedious wait of uploading 2 GB video files over home or hotel Wi-Fi connections.
Batch processing for media libraries and interview archives
Researchers, journalists, legal professionals, and product managers often return from field interviews with dozens of voice memos.
Oyma‘s local media pipeline is built for high-throughput batch processing:
- Asynchronous queue management: Queue multiple interviews consecutively. Oyma processes files sequentially in the background while you continue writing notes in the live-preview Markdown editor.
- Automatic speaker turns and pacing: Whisper segments speech into logical conversational paragraphs, accurately inserting punctuation, commas, and question marks.
- Clickable playback offsets: Every paragraph includes an embedded audio offset (such as
[08:45]). Clicking any offset immediately cues the built-in media playback bar to that exact millisecond, allowing you to cross-verify disputed quotations effortlessly.
When processing completes, the newly synthesized text file is indexed immediately by our local full-text search engine, making every spoken word discoverable across your vault.
From raw speech to structured synthesis
Generating a verbatim transcript is only the first step. The true value emerges when speech transforms into structured intellectual capital.
Once an imported audio or video file is transcribed:
- Local AI summarization: With a single click, feed the transcript into local AI with Ollama to generate action item matrices, thematic breakdowns, and executive takeaway logs.
- Interactive questioning: Use our ask feature to interrogate the imported recording: “What concerns did the client raise regarding the Q4 rollout timeline?”
- Cross-linking and synthesis: Connect key interview concepts to existing architectural documentation or customer profiles using bidirectional wikilinks & backlinks.
Your archived podcasts, lecture recordings, and customer interviews cease to be static media files gathering digital dust; they become dynamic, searchable nodes in your graph view.
Robust media format decoding and storage management
Imported recordings originate from a wide array of recording hardware, ranging from high-end studio microphones to compressed phone voice memos.
Oyma‘s media demuxer handles format diversity gracefully:
- Lossless channel downmixing: Stereo and multi-channel audio files are downmixed to mono 16 kHz streams optimized specifically for the Whisper acoustic model architecture.
- Variable bitrate resilience: Files encoded with variable bitrates (VBR) or non-standard container metadata are normalized cleanly without pitch shifts or temporal drift.
- Zero redundant file duplication: When transcribing large video files, Oyma extracts only the audio stream into a temporary buffer, leaving the original video file in place on your disk rather than duplicating multi-gigabyte media assets.
Unmatched privacy for confidential recordings
Certain recordings can never legally or ethically be uploaded to external cloud services:
- Confidential merger and acquisition discussions.
- Patient consultations subject to medical privacy rules.
- Whistleblower and investigative journalism interviews.
- Internal human resources investigations.
Because Oyma executes transcription 100% on-device utilizing local Whisper weights, zero audio bits, transcripts, or metadata ever transit the public internet. You can perform transcription on a completely air-gapped Mac with Wi-Fi and Bluetooth disabled, guaranteeing total privacy compliance.
Turn your archive of voice memos and video recordings into searchable, structured knowledge with Oyma.
Questions
Related: Markdown basics, vault compatibility, and free browser tools.
Which audio and video file formats can Oyma transcribe?
Standard audio and video formats including MP3, M4A, WAV, AAC, FLAC, MP4, and MOV files are natively supported.
Does transcribing a video file upload the footage to cloud servers?
Never. The application uses local system media demuxers and on-device Whisper models. Not a single byte of video or audio leaves your Mac.
Can I transcribe recordings captured on an iPhone Voice Memo or handheld recorder?
Yes. Simply drop or import your voice memo file into Oyma, and it processes locally into a formatted Markdown note.
How long does it take to transcribe a 60-minute audio recording?
On modern Apple Silicon Macs (M1/M2/M3/M4), local Whisper processing typically transcribes an hour-long recording in less than 3 to 5 minutes.
Your notes, in plain Markdown.
Free during the private beta. Apple silicon Macs, macOS 14 or later.