Research

The measurements behind the claims.

Benchmarks, model reports and the code behind our evaluations.

omi-med-stt-v1 same weights
CUDABF16 · batches
MLX q8local attention
CPUquiet cuts
Engineering note

Same weights. 21–23% lower WER.

How a better NeMo result led us through the CUDA, Apple MLX and CPU inference paths.

3 runtimes same weights Sep 2026
Read technical deep dive →
M-WER
1
2
3
Flagship model

omi-medical-1

Our flagship medical speech-to-text ranks #1 of 30 on the medical-first board. It recorded zero drug-name errors. Available through the EU API or private deployment.

#1 of 30 M-WER 0.94% EU cloud
Read article →
GUARD
Writer note
After Guard
520 missed
→ 0
12/12 flagged
Safety

AI Scribe Safety + Omi Guard

Eight frontier models tested with and without a deterministic safety layer. Guard recovered every measured miss and flagged all 12 confirmed hallucinations in this benchmark.

Omi Guard 8 writers Jun 2026
Read article →
0.6B
CC-BY
Parakeet On-device
Medication review
Radiology
Parakeet omi-medical-edge-1
Open model

omi-medical-edge-1

An on-device 0.6B medical speech-to-text model for Apple Silicon, NVIDIA CUDA and CPU.

0.6B CC-BY-4.0 M-WER 2.12%
Read article →
SAFETY
Grounded note
Risk signals
1.0x
3.1x
4.3x
Safety

Clinical SOAP Note Evaluation

A safety-first SOAP benchmark measuring hallucinations, evidence grounding, and clinical coverage.

Safety-first 6 models 300 dialogues
Read article →
3B
MIT
Phi-3 SOAP
Subjective
Assessment
Phi-3 Omi-Sum
Open model

Omi-Sum 3B

An open 3B clinical model for structured SOAP notes, released under the MIT licence.

3B MIT ROUGE-1 70
Read article →