DeafBench by Lantern Arc

Benchmark speech AI by what actually matters.

DeafBench measures accuracy, critical-information preservation, latency, and regressions across speech, captioning, and AI systems.

Beyond one score

A low error rate can still hide a high-impact mistake.

DeafBench keeps ordinary ASR metrics, but adds scoring for the information and timing that determine whether captions remain useful in real situations.

01

Typed critical information

Score categories such as names, numbers, times, dates, codes, money, and other information where substitution can change the outcome.

02

Reproducible evaluation

Use frozen manifests, model adapters, and repeatable inputs so results can be compared without quietly changing the benchmark.

03

Accessibility regressions

Compare model or product versions and surface cases where a release improves an aggregate score while making important scenarios worse.

Research workflow

Careful claims require visible methodology.

DeafBench is built to make its corpus, scoring, model comparison lanes, and limitations inspectable. Results should be reproducible before they become public claims.

01Freeze the input

Version benchmark manifests and test data.

02Run adapters consistently

Normalize model interfaces without hiding differences.

03Score multiple dimensions

Keep baseline error metrics beside accessibility-critical scoring.

04Publish limitations

State corpus scope, uncertainty, and what the evidence does not prove.

Open source

Review the benchmark, not just the result.

DeafBench is public so researchers, accessibility practitioners, and engineers can inspect the implementation and challenge the methodology.