Offhand

Blog · Measurements

What we learned shipping Apple's on-device model

Modest scores on general tests, one narrow job done well: what building on Apple's on-device model taught us.

A scatter of moss-green ink spatter on paper, split by one flat horizontal band

We ran Apple's two on-device models through Artificial Analysis's mobile benchmark — five evaluations of general ability, run their way — and published the runner, the scorers and every answer. Offhand asks the model one narrow question per spoken note.

The general test

42.3AFM 3 Core Advanced

32.6AFM 3 Core · iPhone 15 Pro

The one job

86%of 261 everyday notes entirely right

Figure 1The composite measures a generalist; the product asks a specialist one question. Left: the study's composite scores. Right: the everyday set, count, day and time all correct.

The scores

Mid-table on the composite — and silent on the one job.

BFCLCore Advanced 61.7
BFCLCore 35.6
IFBenchCore Advanced 28.2
IFBenchCore 28.6
AA-OmniscienceCore Advanced 8.2
AA-OmniscienceCore 4.5
GPQA DiamondCore Advanced 39.4
GPQA DiamondCore 30.8
MATH-500Core Advanced 73.8
MATH-500Core 63.4
0 25 50 75 100 — score
Figure 2The five evaluations, run Artificial Analysis's way on macOS 27. Only the study's own scores appear — the board's medians, ranks and every other build are Artificial Analysis's to publish, not ours.

One question per note

One question per note — what are the tasks? — answered through guided generation as data, not prose. The model decides what the tasks are; code reads the day and the time.

A loose brush-stroke wave in moss green settling into three flat bars
Figure 3The job, as a pipeline. Code does what code does better: a cloud model writing its own dates got 52 of 98 English schedules right, where the parser reading the same titles got 88 — and on the sentences it is built for, code only, it does not miss: 232 of 232 weekday phrases.

The window is shared

The ceiling is 4,096 tokens — and instructions and generated items count against it, not just the note. Reckon on the output.

Figure 4Schematic, not to scale. Notes of 150–200 words returning 26 items were comfortable; the 338-word note hit the wall.

Small beats patched

A year of measured edits grew to 547 words — and lost to 16 on every axis. Now every change is judged against a from-scratch floor; the first move on a failure is deleting a sentence.

A year of measured editswords in the prompt 547
Written from scratch — won every axiswords 16
0 274 547 words
Figure 5The incumbent is never exempt: each change is measured against scratch-zero and scratch-min as well as against what ships.

Placement beats wording

Nineteen arms tried to move item boundaries by wording; the same disfluent notes broke in all of them. Moving only where it sat:

One sentence, four places; a filled dot is a broken control note.

Mid-paragraph
Paragraph end
Closing self-check
A different self-check
Figure 65, 1, 3 and 0 of the 5 control notes. Four different wordings at the bad spot broke the same five.

The language of the turn

Told to answer in the note's own language, it never held. One line in that language, after the note, did:

Figure 7The model reads the language to answer in off the words nearest the note — judged rows on macOS 26. On macOS 27's model the counts are already 0, 1 and 1. The Mandarin column is research about the model; the language is measured and not offered.

What the platform does

The run taught as much about the platform as the model:

01

The OS picks the model by hardware; variant is read-only.

02

One request can restart the service — ~20 s of refusals.

03

A throwing tool erases the turn's tool calls.

04

The framework adds its own system prompt.

05

Greedy output repeats byte for byte — one pass stands for many.

Does it do the job?

The everyday set — 261 English notes, macOS 26's model:

224 of 261entirely right — count, day, time
149 of 151one-task notes stay one task
0 of 980invented clock times — every run

It still errs on boundaries — one job can return as several tasks; three or more offers Merge.

Figures decoded on a Mac — the job's on macOS 26's model (the next generation: 218 of 261). The model half needs iPhone 15 Pro or newer; an older phone files each note as one task. English, Spanish or Brazilian Portuguese; the Chinese figures are research.

A composite answers a question nobody ships. The phone's model does the job.

A to-do list for you and your agents.