
Imagine an investment committee meeting on a difficult Monday morning. Inflation has surprised, sterling is weakening, gilt yields are rising, and the firm must decide whether to change its hedge ratio, raise liquidity, or slow its commitments to private markets.
An AI system has done impressive work. It has assembled the data, compared historical episodes, run scenarios, challenged the house view, and produced a recommendation before most committee members have finished their coffee. The prompts are stored. The outputs are timestamped. The model version is known. The meeting is transcribed. The committee approves a modified course: retain the hedge, raise the liquidity buffer, and continue private-market commitments subject to conditions.
Six months later, the decision has become a problem. The model has been replaced. The chief investment officer has left. One of the assumptions proved wrong. A regulator, trustee, or new committee asks four ordinary questions: Who decided? How did the conclusion become binding? Which dissent was overruled? Can the judgement now be reopened without inventing a new history?
The institution has perfect logs, yet no adequate answer. It remembers what the machine said, but not quite what the institution did. The model is transparent. The institution is opaque.
This is institutional amnesia, and more data alone will not cure it.
My previous essays examined the wiring through which shocks travel across financial markets. This essay turns to the wiring of decision itself: how analysis becomes commitment, how that commitment survives, and when it must be reopened.
Much of AI governance begins with the model. Was it accurate? Explainable? Biased? Secure? Aligned? These are necessary questions, but financial institutions do not act through models alone. They act through mandates, committees, delegated authorities, challenge functions, risk limits, and named people who carry responsibility.
An output may contain a recommendation. A recommendation is not a decision. And a decision does not become an institutional judgement merely because enough people clicked “approve”.
The missing object is the institution’s commitment: a persistent record of what was accepted, by whom, under which authority, for what scope and period, on what evidence, with which unresolved objections, and with what consequences. It marks the boundary between analysis and accountability.
This distinction matters because aggregation cannot manufacture authority. Ten models agreeing do not create a mandate. A majority vote among ineligible participants does not become valid because it is mathematically decisive. An eloquent answer does not own the action it induces. The institution must do that.
The shift sounds small, but it changes the target of governance. Instead of asking only whether an AI can explain its answer, we must ask whether an institution can explain how an answer became its judgement.
The natural response is to retain everything: prompts, outputs, documents, votes, emails, model cards, and meeting recordings. Storage is cheap; forgetting feels reckless.
Yet an archive is not necessarily a memory. A warehouse full of records may still fail to answer the question a future authorised reviewer is entitled to ask.
Suppose the committee rejected a minority view because a particular risk model was then approved. Later, the model is withdrawn. Can the original judgement be reconsidered? Only if the institution retained more than the final vote. It may need the rejected alternative, the evidence supporting it, the dependency on the old model, the authority under which the choice was made, and the route by which the judgement entered later decisions.
Now suppose the law changes. A reviewer may acquire powers that no one possessed at the time. The institution cannot reconstruct what it failed to preserve merely because it is now legally permitted—or required—to look.
The problem is therefore not maximal retention but sufficient retention. Financial firms understand this idea in another setting. They hold liquidity reserves not because every cash demand is known, but because some future demands must remain answerable. Institutional memory needs a similar discipline: preserve enough state today to meet the legitimate questions, challenges, and repairs that may arise tomorrow.
This is not an argument for indiscriminate data hoarding. Retention has costs, privacy limits, and its own risks. The task is to identify which distinctions must survive: authority, evidence, assumptions, dissent, dependencies, alternatives, and downstream effects. If those distinctions are erased, no amount of later computing can recreate them honestly.

There is a further complication. We often evaluate an AI model as though it carries a stable quality score from one setting to another. In an institution, performance depends on role.
A model that is useful as an analyst may be dangerous as a chair. One that generates imaginative alternatives may be poor at verifying facts. A cautious model may strengthen a risk function but paralyse a trading decision. A model praised for agreeing with an organisation’s constitution may also suppress the very dissent that makes a committee worth having.
The relevant unit is therefore not simply the model, but the configuration of humans, models, roles, rules, and communication paths around it. Replace one model, change who sees which evidence, or move an AI from adviser to gatekeeper, and the organisation may no longer be running the same decision procedure—even if the workflow screen looks identical.
This is especially important in multi-agent systems. More agents do not necessarily produce more intelligence. They can amplify confidence, imitate one another, conceal common dependencies, or manufacture a theatrical disagreement whose conclusion was fixed by design. Diversity of model names is not diversity of judgement.
Models should be tested not only for general capability, but for fitness in particular institutional roles and for their effect on the committee as a whole. The question is not merely, “Is this model better?” It is, “What kind of institution do we create by putting this model here?”
Michael Mainelli recently described the road beyond large language models as a tension between “primordial ooze” and “intelligent design”: greater scale and emergent complexity on one side; more deliberate theories and architectures on the other. His conclusion favours a mixture, with specialised systems interacting in a multi-agent future.
For high-consequence finance, intelligent design must move up one level. We need not only to design better artificial agents, but to design the institutions around them.
An aligned model is not yet a governed institution. A model may follow a corporate constitution while the organisation remains unable to identify who authorised its use. It may give transparent reasons while the committee quietly discards dissent. It may be individually reliable while its repeated use across connected decisions creates a common point of failure.
Doubt is not noise to be optimised away. Properly placed, it is institutional infrastructure. A robust committee does not merely produce an answer; it preserves the conditions under which that answer may cease to deserve commitment.
Preventing institutional amnesia requires more than preserving a transcript. Before an institution relies on AI in a high-consequence decision, it should be able to pass three tests.
These tests do not require autonomous machines to become legal persons, nor do they require every committee to become a software project. They require institutions to treat judgement as something designed, carried through time, and owned.
Long Finance asks: when would we know our financial system is working? For an AI-assisted institution, the answer cannot be “when the model is usually right”. We would know more when the institution can show how a conclusion became binding, preserve enough to revisit it, and recognise when changes to its human–machine organisation have altered the judgement itself.
Intelligence may emerge from primordial ooze or intelligent design. Accountability will not. It has to be built.
Dr Yuan Gao is Founder and CEO of Pangura Limited. This essay draws on her ongoing research programme on how institutions form, preserve, and revisit judgements in human–AI systems. All three manuscripts are currently under peer review; the first working paper is available on SSRN.