Feature · Built for Mac
Live on-device transcription
Private real-time speech-to-text running entirely on Apple Silicon.
What this helps you achieve
- Real-time streaming speech-to-text in 13 languages plus automatic language detection
- 100% on-device Whisper model execution via Apple Silicon Metal acceleration
- Zero audio bytes or transcript snippets sent to third-party transcription servers
Speech is the fastest way humans exchange complex ideas, but audio recordings are notoriously difficult to reference after the call ends. Re-listening to a 45-minute recording to verify a single specification deadline or architectural decision takes substantial time.
Cloud-based transcription services promise convenience, but they come at an alarming privacy cost. Sending live audio streams from executive boardrooms, confidential HR reviews, or patent brainstorms to remote server farms creates severe regulatory and compliance vulnerabilities.
Oyma solves this problem by bringing OpenAI’s world-class Whisper speech-to-text engine directly onto your Mac. Proven in live-transcription-controller.ts and transcription-languages.ts, Oyma streams live, word-for-word transcriptions in real time during your calls—completely on-device, completely offline, and completely private.
Hardware-accelerated local transcription on Apple Silicon
Running deep learning neural networks in real time on a laptop used to be computationally impossible. Modern Apple Silicon architectures change the paradigm.
By leveraging the Apple Neural Engine, high-performance GPU cores, and high-bandwidth unified memory, Oyma executes optimized Whisper models with sub-second latency:
- Dual-channel audio capture: Audio is captured natively via zero-bot meeting recording, tapping both your microphone input and system output streams.
- Streaming chunking: Spoken audio is divided into overlapping 3-second acoustic windows, normalized, and fed into the local Whisper inference engine.
- Hardware acceleration: Matrix multiplications are offloaded directly to Apple Silicon Metal performance shaders, consuming minimal CPU and preserving your battery life.
- Live transcript display: As participants speak, transcribed words stream smoothly into the workspace sidebar, complete with accurate punctuation and sentence segmentation.
At no point during the meeting is a single packet of audio transmitted across the internet. You can turn off Wi-Fi entirely, and transcription continues without missing a syllable.
Choose the right Whisper model for your workflow
Different machines and workflows have varying performance priorities. Oyma lets you select between three open Whisper weights, downloaded once on demand:
- Base Model (148 MB): Blazing-fast inference with virtually zero memory footprint. Ideal for base M1 or M2 MacBook Air laptops during fast-paced 1-on-1 calls.
- Small Model (488 MB): The recommended balance for everyday business use. Excellent accuracy on technical terminology, accents, and cross-talk while maintaining rapid turnaround.
- Large v3 Turbo (1.6 GB): State-of-the-art accuracy for multilingual international conferences, complex legal jargon, and difficult acoustic environments. Highly recommended for M2/M3/M4 Pro and Max processors.
Models are stored permanently in your local application directory. You download them once; they belong to your Mac forever without recurring subscription fees or API billing meters.
Multilingual support across 13 major languages
Global teams rarely conduct all discussions in a single language. Oyma features native transcription support across 13 widely spoken languages:
- English, Spanish, French, German, Italian, Portuguese, Dutch
- Japanese, Korean, Mandarin Chinese
- Russian, Swedish, Polish
In addition to explicit language selection, Oyma includes an Automatic Detection mode. The model evaluates the acoustic features of the opening sentences, automatically configuring vocabulary mappings to match the spoken language without manual intervention.
Acoustic preprocessing and dual-track separation
Accurate transcription in real-world environments requires clean audio input. Spurious background noise, keyboard clicks, and cross-talk between call participants frequently degrade speech-to-text accuracy.
Oyma‘s audio engine executes pre-inference signal conditioning before frames enter Whisper:
- Voice activity detection (VAD): Silence and non-verbal pauses are filtered out, preventing the model from hallucinating repetitive phrases during quiet listening periods.
- Dual-channel isolation: By recording system audio and microphone input on separate audio channels, the model processes your voice and remote participants independently, preventing audio echo from confusing speaker turns.
- Automatic gain compensation: Low-volume speakers are dynamically boosted while loud audio spikes are gracefully compressed, ensuring consistent acoustic levels across all participants.
Interactive timestamps and Markdown integration
A static block of text is only the starting point. In Oyma, live transcription is deeply integrated with your note-taking environment:
- Clickable audio timestamps: Every transcript paragraph is tagged with an exact millisecond timestamp (such as
[14:32]). Clicking the timestamp instantly cues the local audio player to that precise moment in the call. - Searchable in full-text index: Transcripts are automatically indexed by our local full-text search engine, allowing you to locate any spoken phrase months later.
- Immediate AI synthesis: When the meeting concludes, your local transcript feeds directly into local AI with Ollama, synthesizing action items and decisions without typing a single prompt.
Transcripts are written to disk as clean plain Markdown files in your vault, ready to be reviewed in our live-preview Markdown editor or inspected inside Obsidian.
Experience private, lightning-fast on-device speech transcription on your Mac with Oyma.
Questions
Related: Markdown basics, vault compatibility, and free browser tools.
How are Whisper models downloaded and stored on my Mac?
Models download once on demand directly from open repositories to your local application data directory. They never expire and require no recurring network checks.
What hardware is required to run live Whisper transcription?
Live real-time transcription is optimized for Apple Silicon Macs (M1, M2, M3, M4) utilizing the Apple Neural Engine and unified memory.
Which languages are supported for live transcription?
Thirteen major languages are natively supported, including English, Spanish, French, German, Japanese, and Mandarin, along with automatic language detection.
Are transcripts saved alongside the audio recording in my vault?
Yes. Timed transcript segments are saved directly into your Markdown meeting note with clickable timestamps that link to the local audio file.
Your notes, in plain Markdown.
Free during the private beta. Apple silicon Macs, macOS 14 or later.