Token-efficient architectures

Cost per useful answer is the real limit on who gets to use this technology.

Why this is a research programme and not an optimisation

If an AI assistant costs several pounds a day to run, it will be sold to enterprises and to nobody else. If it costs pennies, it can be sold to a household for about a pound a month. The difference between those two worlds is not a pricing decision, it is an architecture decision, and it has to be made at the start.

So we treat cost per useful answer as a primary metric, tracked in the same place as quality, rather than something reviewed when the bill arrives.

Context discipline

Most systems waste the majority of their context on material that does not change the answer: full tool schemas that will never be called, entire conversation histories where a summary would do, and raw tool output passed through verbatim.

We study pruning as an active decision at every turn — what to keep in full, what to summarise, what to drop entirely, and how to measure whether the pruning cost any answer quality.

Routing cheap work to cheap models

A large share of the calls in an agent system are mechanical: classify this, extract that, reformat this list. Sending those to a frontier model is the single most common source of waste we find. Routing them to a small model typically costs a fraction and loses nothing measurable.

Caching and reuse

Prompt caching, result reuse and deduplicating near-identical calls across a multi-agent system. In agent architectures the same context is often re-sent many times over; how a system is structured decides whether that is nearly free or ruinously expensive.

Calling the expensive model once

The largest savings are structural rather than incremental. A system designed so that one well-briefed call to a capable model replaces eleven poorly-briefed ones is both cheaper and better. Most of this programme is really about that: making the brief good enough that you only need to ask once.

Mofy AI Ltd Research Findings The Stack Tool use Memory Edge Efficiency Emotional awareness Identity Patents The Group About Contact

Mofy AI Ltd — registered in England and Wales, company number 16562138. An artificial-intelligence research company working on tool use, privacy-first long-term memory, edge models, token-efficient architectures and emotional awareness.

Part of a group of three UK companies: Mofy AI Ltd (research), Bonz-Ai Limited (applied), and Ask Zai (product).

Contact: [email protected]

Research Findings The Stack Patents The Group About Contact