AI dictation benchmark: test developer terms and meaning

Khoa Truong Nguyen AnhSources checked Updated

A transcript can read fluently and still change “do not deploy” into “deploy.” For developer work, that mistake matters more than a missing comma. A useful dictation benchmark measures whether instructions survive and how much work remains before the text is ready.

This is a protocol you can run with tools available on your machine. The downloadable dataset contains written prompts, not recorded speech. There are no measured product results or claims that one tool handles Vietnamese-accented English better than another.

Define what your dictation benchmark compares

Choose one task, such as entering a coding request into a local text editor. Compare two tools you can actually run on the same device, with the same microphone and target app. If you are still choosing candidates, the developer voice-tool guide separates dictation from call audio processing and task automation.

Decide whether you are comparing raw recognition or the finished text inserted by an app. Cleanup, dictionaries and surrounding text can affect the latter. Wispr's accuracy and limitations guide, checked September 13, 2026, describes cloud transcription and formatting that can remove fillers or clean up self-corrections. Comparing that output with a raw transcript answers a workflow question, not which underlying speech model is more accurate.

Record the voice model, cleanup model, context setting and dictionary for each configuration. Use none when a stage is disabled and unknown when the tool does not expose it. An unknown setting remains a limitation of the comparison.

Use a fixed set of sentences

Download the 20-sentence input CSV. It contains five examples in each of four groups: Vietnamese, Vietnamese with developer vocabulary, English, and switching languages between sentences. The Vietnamese dictation guide explains the language-setting questions behind those groups.

Read the sentences as written for the controlled part of the test. A spontaneous coding request can be a separate condition. Do not quietly combine scripted reading and improvised speech in one score.

The English group does not contain an accent. An accent belongs to the speaker's delivery, which has not been recorded here. Use anonymous speaker IDs and, if relevant and volunteered, a self-described language background in your notes. Do not infer nationality or other attributes from the audio. One person's result describes that person under those conditions.

The Whisper model card documents uneven performance across languages, accents and dialects for those models. That is a reason to test the intended use case; it does not establish the relative performance of the apps in your shortlist.

Keep the recording conditions comparable

For live dictation, alternate the order of tools between rounds. Three repetitions per sentence is a manageable starting design, not a statistical guarantee. Two tools, 20 sentences and three repetitions would produce 120 attempts per speaker. This is a proposed workload, not a completed experiment.

Keep mic distance, room, target document and speaking instructions stable. Note warm-up runs separately. The microphone comparison procedure helps distinguish missing audio from a recognition error.

If both tools accept the same audio file, replaying identical input removes differences between readings. Record that as file input. Do not mix its latency with live dictation, and do not assume a dictation-only app supports file import. For live runs, use live and retain recordings only where the app and your test permissions allow it.

Freeze a baseline before tuning vocabulary. Dictionary changes create a new configuration; the developer dictionary guide helps prepare a separate vocabulary condition.

Review meaning, terms and numbers separately

Use the benchmark results template, which extends the earlier worksheet with model, context, input-mode and reference-version fields. Every supplied row is not_run. Copy rows for each tool and repetition, and keep the original output before correcting it.

For each completed attempt, review three fields:

  • meaning_preserved: yes only if the instruction and its conditions survive. Dropping a prohibition or reversing the order of actions is no.
  • terms_preserved: use the input CSV's required terms and evaluation note. Identifier casing matters for userId; use not_applicable when the sentence has no required terms.
  • numbers_preserved: judge the value, unit, date or time. Equivalent forms such as “nine thirty” and “09:30” can be acceptable; a changed value is not. Use not_applicable when there is no numerical content.

A blank rating means not reviewed. It is different from no and from not_applicable. completed means the attempt finished, not that the output was correct. A completed empty transcript can therefore receive no for meaning. Use failed for a run that did not complete and record the reason in notes.

Measure latency_ms from releasing the dictation control to all text appearing in the target field. Measure correction_seconds from starting your review of that output until it is usable, including reading it. If a metric cannot be measured, leave it blank; zero means a measured zero. Describe any different timing definition as a separate condition.

Run the local summary script

Download summarize-dictation-benchmark.py into the same folder as the results CSV. It uses Python 3.9 or newer and only the standard library. After reviewing the source, run:

python3 summarize-dictation-benchmark.py developer-benchmark-results-template.csv > summary.json

The untouched template produces 20 not_run rows and an empty groups list. That is the expected starting state, not a zero-error result. Save completed measurements in a separate CSV and pass that filename when ready.

For attempted rows, fill the configuration fields, sample_id, positive integer trial and reference. The script groups by speaker, date, tool/version, device, mic, target app, language, condition, dictionary, configuration ID, models, context, input mode, reference revision and sentence group. Give a changed setup a new config_id; set reference_revision to a new value if you edit the input set.

Each group reports attempted/completed/failed counts, sample IDs, rating counts and median timings. A rating's yes_fraction uses only reviewed yes and no values. Missing and inapplicable ratings are shown separately. Timings cover completed attempts with recorded values; the script reports their observation counts and missing values. Failed attempts remain visible in the failure count and are not given invented completion times.

There is no combined winner score. Before comparing groups, check that sample IDs, repetitions and review coverage match. A high fraction based on one reviewed sentence is not comparable to a fully reviewed set. The script does not pair trials or calculate confidence intervals.

Why this script does not calculate WER

Word error rate measures edits between a reference and hypothesis; JiWER's documentation describes minimum-edit-distance measures and empty-reference behavior. WER is useful when normalization and tokenization are defined, but a meaning-preserving rewrite can receive word errors while a short, consequential negation error can look small.

If you add WER later, publish the exact preprocessing and score comparable language groups separately. Whitespace splitting Vietnamese and English does not automatically produce equivalent units. Keep the original transcripts and manual judgments alongside the metric.

The supplied script only summarizes your reviews and timings. It does not listen to audio, judge meaning, detect accents or call an API. Its behavior was checked with synthetic CSV fixtures, including missing values, failed attempts, multiple configurations and invalid data. Those checks validate the utility, not the recognition quality of any product.

Publish observations with their limits

A useful results update includes the configurations, input revision, permitted recordings where available, original outputs, manual ratings, failure reasons and timing definitions. Show the number of speakers, samples and repetitions, and keep failures beside successes.

Until two real tool runs and their reviews exist, this page remains a benchmark kit. Use it to answer a narrow question about your own workflow before extending any conclusion to other speakers, microphones or languages.

Some links may earn a commission if you subscribe, at no extra cost to you. Commercial relationships do not determine the recommendation; we state the constraints and alternatives.