Research

Open evaluation. Published models.

We believe healthcare AI should be verifiable. We publish our benchmarks, open-source our evaluation code, and release models under permissive licences.

M-WER
1
2
3
Flagship model

omi-medical-1

Our flagship medical speech-to-text — #1 of 28 systems on the sealed benchmark by Medical WER, with zero fabricated drug names. EU cloud or self-hosted.

#1 of 28 M-WER 0.94% EU cloud
Read article →
M-WER
1
2
OS
30
Speech

Medical Speech-to-Text Benchmark

28 systems ranked on medical audio using Medical Word Error Rate — sealed test set, open scorer, every major cloud API and medical variant.

28 systems M-WER Jul 2026
Read article →
0.6B
CC-BY
Parakeet On-device
Medication review
Radiology
Parakeet omi-medical-edge-1
Open model

omi-medical-edge-1

An on-device 0.6B medical speech-to-text model, benchmarked against 28 open and cloud systems.

0.6B CC-BY-4.0 M-WER 2.16%
Read article →
GUARD
Writer note
After Guard
520 missed
→ 0
12/12 flagged
Safety

AI Scribe Safety + Omi Guard

8 frontier AI models, with and without a deterministic safety layer. Scribes forget 43× more than they invent — in this benchmark Guard recovered every measured miss and flagged all 12 confirmed hallucinations.

Omi Guard 8 writers Jun 2026
Read article →
SAFETY
Grounded note
Risk signals
1.0x
3.1x
4.3x
Safety

Clinical SOAP Note Evaluation

A safety-first SOAP benchmark measuring hallucinations, evidence grounding, and clinical coverage.

Safety-first 6 models 300 dialogues
Read article →
3B
MIT
Phi-3 SOAP
Subjective
Assessment
Phi-3 Omi-Sum
Open model

Omi-Sum 3B

An open 3B clinical model for structured SOAP notes, released under the MIT licence.

3B MIT ROUGE-1 70
Read article →