These are not theoretical papers. Each one came out of a real deployment hitting a wall that published research did not answer, and each one runs in production today.
Models talking to models. A routing layer that pairs the right model to the right task, call by call, across several providers — and does it without the person asking ever seeing a model name, a configuration screen or an API key.
The research question underneath it is unglamorous: how do you decide, before you have an answer, which engine the question deserves? Get it wrong in the cheap direction and quality falls off a cliff. Get it wrong in the expensive direction and nobody can afford to use the thing.
Real-time monitoring for manipulation, emotional distress and self-harm — watching the conversation as it moves rather than scanning finished messages for banned words. It watches the AI's behaviour as well as the user's, because a system that drifts is as much of a risk as a person who is being led somewhere.
Because it tracks a direction of travel rather than a single line, it can raise a concern about a conversation in which no individual message would ever be flagged. You cannot be frog-boiled.
Generative interface work: a workspace that assembles itself around the task rather than being picked from a fixed menu of screens, and that adapts to the state the person is actually in.
For a neurodivergent user especially, an interface that changes shape between sessions is a tax. One that changes shape around what you are doing, while staying recognisable, is the opposite.
We have built, measured and torn down several memory architectures looking for one that lasts years, stays on the user's side of the trust boundary, and forgets deliberately rather than by accident.
The shape that has held up is mostly local, a little cloud, with a graph binding it together — the same answer the edge and efficiency programmes keep arriving at from their own directions. The cheapest configuration is also the most private one, and that is not a coincidence. Read more.
Live voice agents with emotional context and tool calling — real conversation rather than turn-based chat. Recognition, synthesis, natural turn-taking, and latency low enough that it does not feel like a walkie-talkie.
Speech is the fastest input most people have and the only one some people have, which makes it a research priority rather than a feature.
None of these were designed as products. Each is what was left standing after a programme ran long enough to rule out the alternatives, and they only look like a stack in hindsight.
What binds them is a single constraint: the result has to work for someone who cannot configure anything, cannot afford a retainer, and may be having the worst day of their life while using it. That constraint is why the memory is mostly local, why the routing never asks the user to choose, why the interface adapts, and why something is always watching the emotional direction of the conversation.
The stack is licensed within the group and, selectively, outside it — for partnerships, IP licensing and platform integration. Email [email protected].