Offhand

Blog · Measurements

How often Offhand gets it right

How often a spoken note comes out entirely right, on everyday notes and on the hardest set we have. How the model inside the iPhone measures against 3 others its own size. And how well iOS hears you before any of that starts, including the thing it gets wrong most often: names.

Offhand turns a spoken note into tasks. The thinking is done by the model that ships inside the iPhone, called Apple Intelligence. Nothing is sent anywhere.

So the fair question is how well that works, and whether a bigger model would do better. This is what we measured: how often a note comes out right, how the phone's model compares with 3 others its own size, and how well iOS hears you before any of it starts.

Apple Intelligence, and Offhand's own rules

Two things do the work. Apple Intelligence — the model built into the iPhone — reads the note and decides what the tasks are. Offhand supplies the rest: the instructions that tell the model what counts as one task, and the code that reads the day and time out of the words. Every figure in this table is the two of them together.

What How often it is right Measured on
The everyday set, every field right at once: task count, day and time 224 of 261 (86%) 261 everyday English notes
A note that asks for one task comes out as one task 149 of 151 (99%) the everyday English set, 261 notes
Same, on our hardest set 85 of 93 (91%) 136 hard notes
When the task count is right, the date and time are right too 29 of 30 (97%) · 65 of 68 (96%) the everyday set · the hardest set
The day you named is the day the task lands on 114 of 123 (93%) every everyday note that named a day
A clock time you said is the time you get 30 of 33 (91%) every everyday note that said one
A clock time you did not say is never invented 980 of 980 every note in this study, both languages
"Next Monday", "this Friday", "tomorrow", "tonight", 下周三, 明天: the right day, every time 568 of 568 568 sentences, English and Chinese, read by code
Clean English speech, words right about 97 in 100 (3.3% word error) 800 audiobook sentences
Our own recorded to-do speech, English, words right about 95 in 100 (5.3% word error) 725 recordings

And the same answer every time. We ran the everyday set of 261 notes on 4 different nights, in different setups, and got the identical score each time. The part that reads the day and time out of the words was re-run after a rewrite and gave the same answer on all 397 notes.

Where it is under 9 in 10, we say so: splitting a note into several tasks, Chinese clock times, spoken names. Those are in the tables below.

How we tested the models

We wrote 104 situations and said each one twice, once in English and once in Simplified Chinese, so the two languages face the same test. The notes are short: in English, 8 words in the middle of the range and 15 at the longest.

Short is not the same as easy. The work is not hearing the words, it is deciding how many tasks are in them, and most of these notes are written to make that call hard. "Buy eggs, bread and soy sauce" is 1 task, not 3. "The app crashes on launch, especially after opening the settings page" is 1, not 2. "Pay the rent, call my mother, and book the train tickets" is 3.

For every note we wrote down the right answer: how many tasks it should make, and what day and time each one should get. A note only counts as right when everything is right: the number of tasks and the schedule on each one. One wrong field fails the whole note.

Then we gave the same notes, with the same instructions, to 4 models:

Longer notes, from our other sets.

The 104 above are the comparison. Our other test sets hold longer notes, the kind people leave when they are not being careful, and those are where you can see what Offhand does with a whole note rather than a score. All 3 below are rows from those sets, shown in full with what came out.

A 15-item list, said in one breath.

Everything the house needs before winter. Bleed the radiators upstairs. Service the boiler. Clear the gutters. Check the roof tiles after the storm. Replace the cracked bathroom tile. Draught-proof the back door. Lag the pipes in the loft. Order a cord of logs. Sweep the chimney on Wednesday. Test the smoke alarms. Replace the carbon monoxide detector. Book the electrician for the consumer unit. Fix the gate latch. Put the garden furniture in the shed. Drain the outside tap.

15 tasks, 15 things. The opening sentence is dropped because it is not a task. Only the chimney sweep gets a day, Wednesday, because it is the only one that named a day. Same result on today's model and on the next one.

A rambling note, filler words and all.

um so tomorrow i've got that email to sort for the school trip thing and i think i need to reply to dave's team and also i still haven't texted sam back have i, ugh, and gym maybe thursday or friday not sure yet

4 tasks: Email to sort for the school trip thing · Reply to Dave's team · Text Sam back · Gym. The "um", the "have i", the "ugh" are gone. One honest miss inside the pass: all 4 landed on tomorrow, including the gym that was "maybe Thursday or Friday". The word "tomorrow" at the front of the note was read as covering everything after it.

Changing your mind halfway through.

lunch with pat tomorrow. oh no, lunch with pat is next week. tomorrow is the review

2 tasks: review tomorrow on tomorrow, lunch with pat next week with no date. The correction was taken and nothing was written down twice. More of the same shape: "pick up the cake at 5. no, that's wrong, pick up the flowers at 5 and get the cake whenever" moves the 5 o'clock to the flowers; "order the blue cover for the sofa. no, the red one" leaves one task, the red one.

This is the hardest thing we test, and we say so: on our 70-note set of people changing their minds, about half come out entirely right. The commonest miss is the version you took back staying on the list as an extra row. You can see it and delete it, but you should know it happens.

What we found

The model on the phone did the job.

Model English, notes entirely right Chinese, notes entirely right
Apple's model, as shipped in Offhand today 78 of 104 (75%) 84 of 104 (81%)
Apple's model, next generation 85 of 104 (82%) 87 of 104 (84%)
Qwen 3.5, 4B 72 of 104 (69%) 70 of 104 (67%)
Gemma 4, E4B 71 of 104 (68%) 73 of 104 (70%)

Given our instructions, the model already in the phone scored highest in both languages, ahead of the 2 open models of its own size.

One honest caveat travels with that table. Most of the gap is one kind of note: the "buy eggs, bread and soy sauce" shape above, and the item added after a pause, "buy bread. Butter." We count those as one task, and Apple's model agrees. The other models made one task per thing. 31 of the 104 notes are that shape. Set them aside and the models are within a few notes of each other.

So the fair reading is not "the phone wins". It is: for this job the phone's own model is as good as anything its size, and it costs nothing, works with no signal, and never sends your words away.

The model never invented a time. Across every run of Apple's model, zero tasks got a clock time the note did not say. If you did not say a time, you do not get an alarm.

Hearing you

The words have to be right before any of this matters. Offhand uses the speech recognizer built into iOS. We ran it head to head against Cohere Transcribe, the model at the top of the public speech leaderboard, on 4,642 recordings in English and Chinese.

Apple's recognizer Cohere Transcribe
English, word error rate 5.5% 3.9%
Chinese, character error rate 9.0% 9.1%
English silence that got words invented 0 of 25 17 of 25
Chinese silence that got words invented 6 of 25 25 of 25

Cohere is better on English words, by 1.7 points, and level on Chinese. But it writes words over silence: 68% of the English silent clips and 100% of the Chinese ones came back with text in them. In a to-do app, a phrase invented from a door closing becomes a task nobody asked for. Apple's recognizer wrote nothing over English silence.

The recognizer's real weakness is names: fewer than 4 spoken names in 10 survive intact. Numbers, dates, times and prices survive every time. We say so on the box.

Before you trust these numbers

Measured August 2026 on macOS 26.5 and macOS 27.0, with the instructions that ship in Offhand. Every figure comes from a run we can repeat, and the notes and result rows are kept on file.