
We ran Apple's two on-device models through Artificial Analysis's mobile benchmark — five evaluations of general ability, run their way — and published the runner, the scorers and every answer. Offhand asks the model one narrow question per spoken note.
The general test
42.3AFM 3 Core Advanced
32.6AFM 3 Core · iPhone 15 Pro
The one job
86%of 261 everyday notes entirely right
The scores
Mid-table on the composite — and silent on the one job.
One question per note
One question per note — what are the tasks? — answered through guided generation as data, not prose. The model decides what the tasks are; code reads the day and the time.

The window is shared
The ceiling is 4,096 tokens — and instructions and generated items count against it, not just the note. Reckon on the output.
the 338-word, 45-item note: 4,090 — refused
Small beats patched
A year of measured edits grew to 547 words — and lost to 16 on every axis. Now every change is judged against a from-scratch floor; the first move on a failure is deleting a sentence.
scratch-zero and scratch-min as well as against what ships.Placement beats wording
Nineteen arms tried to move item boundaries by wording; the same disfluent notes broke in all of them. Moving only where it sat:
One sentence, four places; a filled dot is a broken control note.
The language of the turn
Told to answer in the note's own language, it never held. One line in that language, after the note, did:
What the platform does
The run taught as much about the platform as the model:
The OS picks the model by hardware; variant is read-only.
One request can restart the service — ~20 s of refusals.
A throwing tool erases the turn's tool calls.
The framework adds its own system prompt.
Greedy output repeats byte for byte — one pass stands for many.
Does it do the job?
The everyday set — 261 English notes, macOS 26's model:
It still errs on boundaries — one job can return as several tasks; three or more offers Merge.
Figures decoded on a Mac — the job's on macOS 26's model (the next generation: 218 of 261). The model half needs iPhone 15 Pro or newer; an older phone files each note as one task. English, Spanish or Brazilian Portuguese; the Chinese figures are research.
A composite answers a question nobody ships. The phone's model does the job.
A to-do list for you and your agents.