Offhand

Blog · Measurements

How well does it hear you?

Eleven languages through the recogniser Offhand uses, three thousand adversarial cases aimed at what we suspected it got wrong, and five hundred spoken notes put through the extraction step. Numbers, dates and times survive every time. Names do not, and this is where we say so.

Two boundaries, before any number below. Every figure here was decoded on a Mac, not on an iPhone. The speech model and the language model are both managed by the OS and undisclosed; macOS 26 and iOS 26 need not carry the same revision, and we have not compared them — so these are Mac findings and phone hypotheses. And there is no timing here at all: nothing has measured how long Offhand takes on a phone, so we print no seconds, no percentiles and no battery figures rather than pass off a laptop's numbers as a phone's.

Three questions, asked separately, because they fail separately.

Can it hear the language?

Spanishin the app, word error 4.9%
Brazilian Portuguesein the app, word error 7.5%
Englishin the app, word error 7.9%
Best of the elevenItalian — not offered 3.5%
Simplified Mandarinmeasured, not offered 9.0%
Worst of the elevenTraditional Mandarin — withdrawn 9.2%
0% 5% 10% error rate
Figure 1Lower is better. Every language falls between 3.5% and 9.2% error. Recognition is therefore not what limits the app: the languages that perform badly in Offhand do so at the extraction step, not the transcription step. Chinese, Japanese, Cantonese and Korean are scored per character rather than per word, because they have no spaces to score against; for Korean, word scoring would measure a spelling convention rather than recognition.

Offhand listens in three languages, one at a time: English, Spanish or Brazilian Portuguese. The other eight rows below are research about the recogniser, not a list of what the app speaks, and a good score here is not a promise that a language is coming — Italian reads best of all eleven and Offhand does not speak it. Simplified Mandarin is the row to read carefully: it is measured, it was offered for a time, and it was taken out on 16 September 2026. Its number never moved. Only its scope did.

Language Error Metric Gave up on the note Verdict
Spanish 4.9% word 10% Ready — in the app
Portuguese (Brazil) 7.5% word 10% Ready — in the app
English (US) 7.9% word 0% Ready — in the app
Italian 3.5% word 15% Usable — not offered
Korean 5.4% character 25% Usable — not offered
Japanese 5.9% character 15% Usable — not offered
French 6.2% word 15% Usable — not offered
German 6.5% word 15% Usable — not offered
Cantonese 8.5% character 40% Not ready — withdrawn
Mandarin (Simplified) 9.0% character 15% Measured — not offered
Mandarin (Traditional) 9.2% character 30% Not ready — withdrawn

Table 1 — "Gave up on the note" is how often the model produced one summary rather than tasks. It is the column that decided every verdict, and the verdicts decided what ships: Traditional Mandarin and Cantonese fail this way two to three times as often as English, so they were offered, measured, and taken back out on 3 August 2026. All three languages in the app are rated ready, and English gave up on no note at all.

What it actually gets wrong

The first round of adversarial cases read as 13.82% word error. Scored again ignoring regions that differ only in notation, it was 5.29%. The recogniser was writing "gmail dot com" as gmail.com, "twelve hundred dollar" as $1,200, "the twenty second" as the 22nd — and being marked wrong for it. Most of an apparent error rate is not error. What is left, is names.

Numbers, dates, times, codes 100%
Small words“a”, “the”, absorbed or not 98.8%
Technical jargon 95.3%
Similar vowelsminimal pairs 92.9%
Homophones 74.2%
Product and brand names 66.7%
People’s namesthe largest loss in the product 38.2%
0 50% 100% survived intact
Figure 2How often the entity that mattered in a sentence came through intact — 490 cases across five voices, English. Numbers and dates are recognised reliably. Names are not, and a task naming the wrong person is not useful. The failure is not partial: across 51 names and products, 27 survived every utterance and 13 survived none. Almost nothing fell in between, because the question is whether a name is in the recogniser’s lexicon, not whether it is hard to pronounce.

What we tried

Apple documents two ways to tell a recogniser which words to expect. We ran both properly, on the corpus that fails.

Mechanism Transcripts it changed
contextualStrings — every proper noun in the corpus, correctly cased 0 of 620
A compiled SFCustomLanguageModelData with phrase counts, authorised with the extra permission it requires 0 of 75

Neither API changed any transcript, in either direction — the count is zero, not a small improvement. There is currently no way to give this recogniser a list of names, so Offhand does not ask for a second permission in order to appear to.

Said Heard
Deyu deu, day, dayu
Xiaoming zioning, ziming
Nguyen new yen
Saoirse sercha
Vercel versal
Supabase superbase
Raycast ray cast

Table 2 — out-of-lexicon names replaced by the nearest ordinary English words. If the people you record have names like these, this is where Offhand will fail most often, and we would rather you knew that before buying than after.

Does the note become tasks?

Output was well formed 97.4%
Right number of tasks 84.3%
Everything right at oncecount, schedule and shape 65.4%
0 50% 100% of 503 notes
Figure 3Two notes in three come out entirely right. This is the strictest figure on this page, and the one to judge the app by. Scoring gives no partial credit: a note that produced the right two tasks but put one of them in the wrong part of the day is counted as a failure here, as it would be to you. Where the model cannot place a phrase in time it files the task under Anytime rather than guessing an hour, and where it cannot find a task at all it says so on the receipt rather than inventing one.

What this page does not tell you