Two boundaries, before any number below. Every figure here was decoded on a Mac, not on an iPhone. The speech model and the language model are both managed by the OS and undisclosed; macOS 26 and iOS 26 need not carry the same revision, and we have not compared them — so these are Mac findings and phone hypotheses. And there is no timing here at all: nothing has measured how long Offhand takes on a phone, so we print no seconds, no percentiles and no battery figures rather than pass off a laptop's numbers as a phone's.
Three questions, asked separately, because they fail separately.
- Can it hear the language? Three hundred utterances per language from FLEURS, a public read-speech corpus, through the same recogniser the app uses. Read speech in a quiet room is the easy case — treat these as a ceiling, not an average.
- What does it get wrong? 3,162 synthesized cases written to attack specific suspicions: names, homophones, brands, jargon, numbers. Synthetic speech has no room, no microphone and no accent the model has not heard. A failure it reproduces is real; a failure it misses is not ruled out.
- Does the note become tasks? 503 written notes with the tasks a person would file recorded alongside, scored on whether the count, the schedule and the output shape all came out right. Scored strictly: one wrong field fails the case.
Can it hear the language?
Offhand listens in three languages, one at a time: English, Spanish or Brazilian Portuguese. The other eight rows below are research about the recogniser, not a list of what the app speaks, and a good score here is not a promise that a language is coming — Italian reads best of all eleven and Offhand does not speak it. Simplified Mandarin is the row to read carefully: it is measured, it was offered for a time, and it was taken out on 16 September 2026. Its number never moved. Only its scope did.
| Language | Error | Metric | Gave up on the note | Verdict |
|---|---|---|---|---|
| Spanish | 4.9% | word | 10% | Ready — in the app |
| Portuguese (Brazil) | 7.5% | word | 10% | Ready — in the app |
| English (US) | 7.9% | word | 0% | Ready — in the app |
| Italian | 3.5% | word | 15% | Usable — not offered |
| Korean | 5.4% | character | 25% | Usable — not offered |
| Japanese | 5.9% | character | 15% | Usable — not offered |
| French | 6.2% | word | 15% | Usable — not offered |
| German | 6.5% | word | 15% | Usable — not offered |
| Cantonese | 8.5% | character | 40% | Not ready — withdrawn |
| Mandarin (Simplified) | 9.0% | character | 15% | Measured — not offered |
| Mandarin (Traditional) | 9.2% | character | 30% | Not ready — withdrawn |
Table 1 — "Gave up on the note" is how often the model produced one summary rather than tasks. It is the column that decided every verdict, and the verdicts decided what ships: Traditional Mandarin and Cantonese fail this way two to three times as often as English, so they were offered, measured, and taken back out on 3 August 2026. All three languages in the app are rated ready, and English gave up on no note at all.
What it actually gets wrong
The first round of adversarial cases read as 13.82% word error. Scored again ignoring regions that differ only in notation, it was 5.29%. The recogniser was writing "gmail dot com" as gmail.com, "twelve hundred dollar" as $1,200, "the twenty second" as the 22nd — and being marked wrong for it. Most of an apparent error rate is not error. What is left, is names.
What we tried
Apple documents two ways to tell a recogniser which words to expect. We ran both properly, on the corpus that fails.
| Mechanism | Transcripts it changed |
|---|---|
contextualStrings — every proper noun in the corpus, correctly cased |
0 of 620 |
A compiled SFCustomLanguageModelData with phrase counts, authorised with the extra permission it requires |
0 of 75 |
Neither API changed any transcript, in either direction — the count is zero, not a small improvement. There is currently no way to give this recogniser a list of names, so Offhand does not ask for a second permission in order to appear to.
| Said | Heard |
|---|---|
| Deyu | deu, day, dayu |
| Xiaoming | zioning, ziming |
| Nguyen | new yen |
| Saoirse | sercha |
| Vercel | versal |
| Supabase | superbase |
| Raycast | ray cast |
Table 2 — out-of-lexicon names replaced by the nearest ordinary English words. If the people you record have names like these, this is where Offhand will fail most often, and we would rather you knew that before buying than after.
Does the note become tasks?
What this page does not tell you
- Nothing here was decoded on an iPhone. Both models are managed by the OS and undisclosed, iOS and macOS need not carry the same revision, and the silicon is unrelated. Until the on-device comparison is run, every number above is a Mac finding and a phone hypothesis.
- No latency, no battery, no thermal figures, anywhere. We have not measured them on a phone, and a laptop's numbers would not be an approximation of a phone's — they would be a different claim wearing the same units.
- The language figures are read speech in a quiet room, from a public corpus. Nobody in them is walking, driving, whispering at midnight, or talking over a kettle. Treat them as the best case.
- The adversarial cases are synthesized speech, not recordings of people. No room, no microphone, no accent the model has not already heard. A failure it reproduces is real; a failure it does not reproduce is not ruled out — which is the more important half.
- Two extraction sub-scores are deliberately not published. One is measured over sixteen cases, which is not a rate. The other has a denominator we could not yet explain, and a number nobody can explain is one that falls apart the first time it is questioned.
- No interviews, no field study, no usability sessions. We have not asked anyone anything. When we do, it will appear here and say so.