Research

Clinical AI that works in any language — and knows when it might be wrong.

Allied-health therapy happens through conversation. When clinician and patient don't share a language, care suffers. We're building a calibrated, multilingual clinical-AI pipeline for PT, OT, speech and ABA practices — and measuring it like a research program, not a feature list.

What we are measuring

Our research aims

Three questions, each with a status that reflects where the evidence actually stands — not where we hope it will.

In progress

Accurate multilingual interpretation

Real-time interpretation that stays accurate on clinical terminology, disordered and pediatric speech, and rapid language-switching — where general-purpose translation breaks down.

In progress

Calibrated clinical-safety confidence

A per-utterance confidence signal that reliably flags unreliable output, so a clinician catches an error before it reaches a patient or a claim.

In progress

Documentation & coding accuracy

Turning the encounter into structured notes and confidence-scored codes — surfacing likely problems before a claim is ever submitted.

Evidence

Evidence to date

Verified results publish here as validation completes — every figure dated, with its sample size and a methodology note. Until then, the live telemetry below fills in automatically as the pilot runs.

Method

Our approach

Measured the way a clinical study is measured: against real encounters, graded by clinicians, with the failure modes counted.

Real encounters

De-identified evaluation corpora of real allied-health encounters, graded by clinical experts.

Clinical error rate, not word error

Accuracy measured against general-purpose baselines by clinically significant error rate.

Calibration that predicts

Confidence validated so that a low-confidence flag genuinely predicts a real error.

Clinician in the loop

By design: the system assists, a human decides.

Responsible AI

Safety and privacy are design requirements

Sefton is built for a clinical setting, so trust is not an afterthought.

HIPAA-aligned by construction

Research metrics are computed from counts, hashed identifiers and categorical outcomes — never patient records.

Aggregate, minimum sample, named reviewer

Published figures are aggregate-only with a minimum sample size, verified by a named reviewer before they appear here.

Measurement only

Every AI suggestion is measurement-only with a human in the loop; nothing bills or acts on its own.

Programs

Programs & recognition

Labels reflect current application status, not endorsement. Updated as programs progress.

Work with us

Partner on research, run a pilot in your clinic, or reach us for media. Preprints and white papers forthcoming.