Five programmes. Each one exists because a real deployment hit a wall that published research did not answer.
We only start a programme when three things are true: the problem blocks a real system we are running, the published literature does not solve it, and we can measure whether we have improved it. That last test kills most ideas, which is the point.
Each programme runs against a live product rather than a benchmark suite. It is a slower way to do research and a much harder one to fool yourself with.
How a model finds and calls the right tool when there are thousands of them and only a few dozen will fit in its context at once. Covers search-before-invoke, progressive disclosure of tool schemas, tool budgets, and recovery when a tool call fails or returns rubbish. Read more.
Memory that lasts years, stays on the user's side of the trust boundary, and forgets on purpose. Covers what is worth storing, how to retrieve it without flooding the context, and how a person inspects and deletes what a system knows about them. Read more.
Which questions a small local model can answer well, which genuinely need a frontier model, and how a system routes between them without the user ever being asked. Covers hybrid routing, local fallback when the network is gone, and the privacy gain from never sending the question at all. Read more.
Cost per useful answer, treated as the primary metric rather than an afterthought. Covers context pruning, caching, routing cheap work to cheap models, and structuring systems so the expensive model is called once rather than eleven times. Read more.
Detecting emotional change, distress and gradual manipulation inside a conversation, at the level of individual chunks rather than the conversation as a whole. UK patent application pending. Read more.