Cost per useful answer is the real limit on who gets to use this technology.
If an AI assistant costs several pounds a day to run, it will be sold to enterprises and to nobody else. If it costs pennies, it can be sold to a household for about a pound a month. The difference between those two worlds is not a pricing decision, it is an architecture decision, and it has to be made at the start.
So we treat cost per useful answer as a primary metric, tracked in the same place as quality, rather than something reviewed when the bill arrives.
Most systems waste the majority of their context on material that does not change the answer: full tool schemas that will never be called, entire conversation histories where a summary would do, and raw tool output passed through verbatim.
We study pruning as an active decision at every turn — what to keep in full, what to summarise, what to drop entirely, and how to measure whether the pruning cost any answer quality.
A large share of the calls in an agent system are mechanical: classify this, extract that, reformat this list. Sending those to a frontier model is the single most common source of waste we find. Routing them to a small model typically costs a fraction and loses nothing measurable.
Prompt caching, result reuse and deduplicating near-identical calls across a multi-agent system. In agent architectures the same context is often re-sent many times over; how a system is structured decides whether that is nearly free or ruinously expensive.
The largest savings are structural rather than incremental. A system designed so that one well-briefed call to a capable model replaces eleven poorly-briefed ones is both cheaper and better. Most of this programme is really about that: making the brief good enough that you only need to ask once.