The cheapest, most private and fastest request is the one that never leaves the device.
Frontier models are extraordinary and expensive. Small local models are cheap, private and instant, and are perfectly adequate for a surprising share of real requests. Almost nobody measures that share honestly, because the incentive in the industry runs the other way.
We measure it. The research question is not "can a small model do this?" but "which requests, in a real workload, does a small model handle at parity — and how does a system tell them apart before it has an answer?"
A classifier decides, per request, whether the local model is sufficient. Where confidence is low the request escalates to a larger model, and where a local answer is produced it can be silently verified against a frontier model on a sample basis to keep the classifier honest over time.
The user is never asked which model to use. Asking would defeat the purpose.
Systems that cannot function without a network exclude anyone with a poor connection and fail at exactly the wrong moments. We work on graceful degradation: what an assistant can still do with no network, how it queues what it cannot do, and how it tells the user which of the two it is doing.
Data that never leaves the device cannot be intercepted, subpoenaed, retained or used for training. For a class of genuinely sensitive requests — health, finance, family, anything said at three in the morning — local inference is not a cost optimisation, it is the only acceptable answer.
We run self-hosted vision workloads on local compute rather than sending screen and camera frames to a third party. Continuous vision is the clearest case where sending everything to a remote provider is both the expensive option and the wrong one.