Edge models and on-device intelligence

The cheapest, most private and fastest request is the one that never leaves the device.

The problem

Frontier models are extraordinary and expensive. Small local models are cheap, private and instant, and are perfectly adequate for a surprising share of real requests. Almost nobody measures that share honestly, because the incentive in the industry runs the other way.

We measure it. The research question is not "can a small model do this?" but "which requests, in a real workload, does a small model handle at parity — and how does a system tell them apart before it has an answer?"

Hybrid routing

A classifier decides, per request, whether the local model is sufficient. Where confidence is low the request escalates to a larger model, and where a local answer is produced it can be silently verified against a frontier model on a sample basis to keep the classifier honest over time.

The user is never asked which model to use. Asking would defeat the purpose.

Offline capability

Systems that cannot function without a network exclude anyone with a poor connection and fail at exactly the wrong moments. We work on graceful degradation: what an assistant can still do with no network, how it queues what it cannot do, and how it tells the user which of the two it is doing.

Privacy through locality

Data that never leaves the device cannot be intercepted, subpoenaed, retained or used for training. For a class of genuinely sensitive requests — health, finance, family, anything said at three in the morning — local inference is not a cost optimisation, it is the only acceptable answer.

Vision at the edge

We run self-hosted vision workloads on local compute rather than sending screen and camera frames to a third party. Continuous vision is the clearest case where sending everything to a remote provider is both the expensive option and the wrong one.

Mofy AI Ltd Research Findings The Stack Tool use Memory Edge Efficiency Emotional awareness Identity Patents The Group About Contact

Mofy AI Ltd — registered in England and Wales, company number 16562138. An artificial-intelligence research company working on tool use, privacy-first long-term memory, edge models, token-efficient architectures and emotional awareness.

Part of a group of three UK companies: Mofy AI Ltd (research), Bonz-Ai Limited (applied), and Ask Zai (product).

Contact: [email protected]

Research Findings The Stack Patents The Group About Contact