Typed critical information
Score categories such as names, numbers, times, dates, codes, money, and other information where substitution can change the outcome.
DeafBench by Lantern Arc
DeafBench measures accuracy, critical-information preservation, latency, and regressions across speech, captioning, and AI systems.
Beyond one score
DeafBench keeps ordinary ASR metrics, but adds scoring for the information and timing that determine whether captions remain useful in real situations.
Score categories such as names, numbers, times, dates, codes, money, and other information where substitution can change the outcome.
Use frozen manifests, model adapters, and repeatable inputs so results can be compared without quietly changing the benchmark.
Compare model or product versions and surface cases where a release improves an aggregate score while making important scenarios worse.
Research workflow
DeafBench is built to make its corpus, scoring, model comparison lanes, and limitations inspectable. Results should be reproducible before they become public claims.
Version benchmark manifests and test data.
Normalize model interfaces without hiding differences.
Keep baseline error metrics beside accessibility-critical scoring.
State corpus scope, uncertainty, and what the evidence does not prove.
Open source
DeafBench is public so researchers, accessibility practitioners, and engineers can inspect the implementation and challenge the methodology.