Memory that lasts years, stays yours, and forgets on purpose

Almost every AI system begins each conversation as a stranger. The industry answer has been to stuff more history into a bigger context window, which is expensive, slow, and still forgets everything the moment the window rolls over.

The real questions are not storage questions

They are: what is worth remembering, who is allowed to see it, and how does a person find out what a system knows about them? Those decide whether memory feels like continuity or like surveillance, and no amount of context length answers them.

We built several architectures, and tore most of them down

This programme has been through more designs than any other. Different splits of what to hold, different retrieval strategies, different rules for what decays and what persists — each built properly, measured against a live system, and in most cases abandoned.

The shape that has held up is mostly local, a little cloud, with a graph binding it together. Most of what a system needs to know about a person never has to leave their machine; the graph is what lets a small local model find the right thing without reading everything; and the cloud is reserved for the narrow slice that genuinely benefits from it.

What surprised us is how well that arrangement scores on every axis at once. It is the cheapest configuration, the fastest, and the most private — the same conclusion the edge and efficiency programmes reached from their own directions. Three independent routes to one answer is the closest thing research gives you to a result you can trust.

What we are still working on

Recall gating. Retrieval is not free. Every recalled memory costs context, and an irrelevant one actively damages the answer. Deciding before retrieval whether a turn needs memory at all, and how much, is a first-class part of the system rather than a tuning parameter.

Deliberate forgetting. A store that only accumulates becomes both a liability and a performance problem. Principled decay by relevance rather than recency alone, and hard deletion that is genuinely hard, including from every derived index.

Inspectability. A person should be able to read, correct and delete every stored fact about themselves, in plain language, without a support ticket.

Privacy by architecture. Privacy achieved by architecture survives a change of ownership. Privacy achieved by policy does not. That is why the default is local rather than a promise not to look.

Applied in

Deployed in Ask Zai, where conversations resume mid-thought and projects build on each other across months — and where the cheapest tier is also the one that keeps the most on the user's own machine.

Mofy AI Ltd Research Findings The Stack Tool use Memory Edge Efficiency Emotional awareness Identity Patents The Group About Contact

Mofy AI Ltd — registered in England and Wales, company number 16562138. An artificial-intelligence research company working on tool use, privacy-first long-term memory, edge models, token-efficient architectures and emotional awareness.

Part of a group of three UK companies: Mofy AI Ltd (research), Bonz-Ai Limited (applied), and Ask Zai (product).

Contact: [email protected]

Research Findings The Stack Patents The Group About Contact