Why long-lived AI systems may need a persistent, derived representation of capability
Imagine an AI system reviewing a commercial contract. It has access to a general language model, a specialist legal model, a clause extractor, a company-policy database, precedent search, a calculator, and a rule that sends certain decisions to a human reviewer. Depending on the implementation, these components may be connected through a predefined workflow, selected dynamically by an agent, or combined through several layers of routing and verification.
There is nothing unusual about such an architecture. AI engineering has already moved well beyond the idea that one model should do everything. Anthropic explicitly distinguishes relatively fixed workflows from more autonomous agents and recommends starting with simple, composable patterns where they are sufficient. Berkeley researchers have described the broader shift from monolithic models toward compound AI systems that combine models, retrieval systems, tools, and conventional software. Voyager demonstrated that an agent can accumulate executable skills and reuse them across later tasks, while Agent Workflow Memory showed that reusable workflows can be induced from prior experience and supplied to future tasks.
The interesting question, therefore, is no longer whether an AI system can contain many different components. It clearly can. Nor is the main issue whether those components are called agents, tools, skills, modules, or services. A more operational question appears once such a system persists over time:
Given the system as it exists now, what can it actually do reliably?
That question sounds similar to asking which tools and models are available, but it is not the same question. A system may still contain every component it had yesterday while its effective capability has already changed.
The components can remain while the capability changes
Consider the same contract-review system several months later. The specialist legal model has been replaced by a newer version. The company has changed its procurement policy. A previously reliable OCR component still performs well on ordinary digital documents but has proved unreliable on a newly encountered class of low-quality bilingual scans. The precedent-search service remains available, although its latency has increased substantially. Meanwhile, a workflow that was validated under the old policy is still stored and can still be retrieved.
If we inspect the inventory, little appears to have disappeared. The models are registered, the tools are callable, and the workflow still exists. Yet it would be unsafe to infer from this inventory that the system retains exactly the same contract-review capability.
The reason is that capability in a compound system depends on more than the presence of components. It also depends on whether a component remains applicable under the current conditions, whether an appropriate provider can actually be invoked, whether its output can still be used by downstream components, and whether a previously successful composition remains valid after one of its assumptions has changed.
This motivates a distinction that will become more important later in this series. A system may possess a capability because one of its components can perform a relevant transformation. That capability may also be callable if the system can recognize when it is appropriate and invoke a suitable provider. A larger organized capability exists only when the relevant capabilities can be combined, under the present conditions, into a reliable path toward the required result.
Tool-use research, routing systems, workflow engines, skill libraries, and capability elicitation already address parts of this problem. The distinction is useful here for a narrower reason: it makes clear that an inventory of components is not yet a description of the system's current effective capability.
Where should that information live?
One possible response would be to centralize everything. Tool descriptions, model versions, evaluation scores, workflow definitions, logs, policy state, and operational telemetry could all be copied into a master database intended to describe what the AI can do.
For a long-lived system, that approach creates a new problem immediately. These sources have different owners, update cycles, and meanings. The tool registry should remain authoritative about the tool; the policy system should remain authoritative about policy; the model registry should remain authoritative about model versions; evaluation systems and operational logs should retain the evidence they actually observed. Copying all of those facts into another store creates duplicated authority and makes synchronization itself a source of error.
This matters because stale capability information can be more dangerous than missing information. If the system knows that it lacks reliable evidence about a capability, it can respond cautiously. If an outdated internal representation tells the system that a route remains valid when one of its dependencies has changed, the system may execute with unjustified confidence.
The alternative is to keep the underlying facts distributed while maintaining a derived organizational state above them. Such a state would not attempt to reproduce the original sources. Its purpose would be to represent selected relationships across those sources when those relationships materially affect what the system can do next.
This is what I mean here by making capability organization a first-class state object. The proposal is not that every fact about every tool should move into one new database. The proposal is that some relationships among capabilities, providers, dependencies, evidence, and current operating conditions may be worth preserving because no individual source owns the complete relationship and repeatedly reconstructing it can be costly or unreliable.
What belongs in a derived capability state?
The answer should be considerably smaller than the total history of the system. In particular, four kinds of cross-source relationship appear especially important.
The first is composition validity. Suppose one component performs transformation A very reliably and another performs transformation B very reliably. Those facts do not establish that the composition A → B is equally reliable. The output representation may be mismatched, an assumption made by B may not hold for outputs generated by A, or an apparently harmless error from A may be amplified downstream. In a compound system, evidence about components and evidence about their composition are different objects. If a particular composition has been repeatedly validated under identifiable conditions, preserving that relationship may save the system from having to rediscover it on every task. Conversely, a repeatedly failed composition can be valuable negative information.
The second is dependency and invalidation. Suppose the procurement policy in our contract example changes. The policy system knows that its authoritative source has changed, but it does not necessarily know which higher-level procedures were learned or validated under the previous policy. A reusable workflow may still exist even though one premise on which its reliability depended is no longer current. The relevant organizational knowledge is therefore not simply that "policy version 7 replaced version 6," but that certain derived capability paths may now require revalidation. This is a relationship across sources, versions, and previously validated organizations.
The third is conditional callability and reliability. Capabilities are rarely just available or unavailable. An OCR system may be highly reliable on digitally generated English documents and unreliable on degraded bilingual scans. A small model may be sufficient for routine classification while becoming unstable near a particular decision boundary. A provider may be technically functional but temporarily unavailable under the current authorization, resource, or network state. Describing a system as simply "having capability X" therefore hides the conditions under which X is useful. What matters operationally is whether an appropriate provider remains callable and sufficiently reliable for the current problem.
The fourth is negative capability knowledge. AI systems naturally accumulate records of what worked, because successful tools, skills, and workflows are useful for later reuse. Yet persistent knowledge of what currently cannot be done may be equally important. A system may know that no validated route exists for a certain requirement, that a required capability is missing, that a particular composition has repeatedly failed, or that an old route remains stored but should not currently be trusted. A failure buried in a log is different from an organizational state that can affect the next decision. Once a known gap can change routing, evidence acquisition, escalation, or termination, it has become part of the system's effective organization.
These four categories are not intended as an exhaustive ontology. They illustrate the more general principle: the useful content of persistent organizational state lies in decision-relevant relations across otherwise distributed sources, rather than in copying those sources wholesale.
A build system provides a useful analogy
Software build systems provide a reasonably close analogy, provided we do not push it too far. Bazel, for example, maintains explicit dependency relationships among build targets and uses its dependency graph to determine which parts of a build are affected when relevant inputs change. The graph does not replace the source code. Its value comes from representing relationships that allow the system to reason about what remains valid after change.
A persistent capability state could serve a comparable role for a changing AI system. The underlying model, database, tool, workflow, policy, or evidence store remains authoritative in its own domain. The derived organizational layer retains enough information to know that a particular higher-level capability depended on certain conditions, that those conditions have changed, and that the corresponding organization may therefore need to be revalidated or recomposed.
The analogy also clarifies why the proposal is not simply "give the agent more memory." A build graph is not a history of every compiler invocation, and it is not a copy of every source file. It is a structured representation chosen because the dependency relations change what should happen next. The same criterion should apply here.
Procedural memory and capability state solve related but different problems
Voyager is an important reference point because it maintains an ever-growing library of executable skills that can be retrieved and composed later. Agent Workflow Memory similarly induces recurring workflows from previous trajectories and uses them to guide subsequent tasks. These systems make it inappropriate to claim that current AI architectures lack reusable procedural memory.
The remaining question is narrower. A stored skill primarily represents a useful procedure that has been acquired. A workflow memory represents a pattern of actions that has worked before. A capability state must additionally answer whether such procedures remain applicable and composable under the current organization of the system.
A procedure can remain stored after one of its dependencies has changed. A capability can become available because a new provider was added even though no new workflow has yet been learned. Two providers for the same logical function may have different reliability, cost, latency, or operating constraints. A previously validated route may remain semantically appropriate but cease to be callable because a dependency has become unavailable.
For this reason, procedural memory and persistent capability state need not compete. The former can supply reusable ways of doing things; the latter can help determine which of those ways remain effective now.
The derived state must remain revisable
Once such a state begins to influence future action, it creates its own failure mode: the system may trust its internal capability representation after the evidence that justified it has become stale.
A useful capability state must therefore preserve some connection to provenance, version, conditions, or other evidence that allows important claims to be reconsidered when their basis changes. The requirement is not to rerun every evaluation before every action, which would eliminate much of the value of persistent state. The requirement is that persistence must not silently turn a historically supported judgment into a timeless fact.
The same consideration applies to provider reliability. If provider A has performed well historically, a rational system will tend to select it more often. But repeated selection of A reduces opportunities to collect evidence about provider B, which can make the historical ranking self-reinforcing even after B improves. Long-lived systems therefore face a familiar exploration problem: past evidence should influence future organization without permanently freezing it.
Exactly how freshness, revalidation, and exploration should be implemented is an engineering question, and different risk classes may justify very different policies. The architectural point is simpler. Organizational state should remain an updateable model of current capability, not an accumulated reputation ledger that gradually detaches from reality.
The objective is minimum sufficient state, not maximum memory
If persistent organizational state is useful, the next temptation is to preserve too much. A system operating for years may accumulate enormous numbers of traces, evaluations, intermediate states, failures, and successful routes. Keeping all of them active would merely move the complexity from task reconstruction into state management.
A more useful principle is to retain the smallest set of derived relationships that preserves relevant future decisions. Raw evidence can remain in evidence stores. Tool definitions remain in tool registries. Detailed execution traces can remain in logs. When the underlying information is cheap and reliable to reconstruct from its authoritative source, there is little reason to duplicate it in persistent organizational state.
What deserves persistence is information whose reconstruction is sufficiently costly, whose omission causes important decisions to change, or whose value lies specifically in a relationship across several sources. A repeatedly validated composition, a known dependency that controls downstream validity, a persistent capability gap, or a condition-specific reliability distinction may satisfy that test.
This makes the problem resemble a sufficient-state problem rather than a memory-maximization problem. Instead of asking how much history an AI can retain, we ask how much of that history must be compressed into the present state for future decisions to remain as good as they would have been with the relevant history available.
That distinction is important because a good long-lived system should be capable of forgetting enormous amounts of detail without repeatedly becoming organizationally ignorant.
Not every AI system needs this layer
There is an obvious cost to maintaining additional state, dependencies, validation records, and update logic. A short-lived application with a small number of stable tools, a fixed workflow, low failure cost, and little need for cross-task learning may gain almost nothing from it. In such a system, reconstructing the required organization when a task arrives may be simpler and safer.
The economics change as the system becomes longer-lived and more heterogeneous. The potential value of persistent capability state should increase as the number of tasks grows, components change more frequently, compositions become deeper, failures become more costly, and accumulated operational evidence becomes more informative. A system that survives for years while models, tools, policies, resources, and external conditions continue to change has a much stronger reason to preserve organizational knowledge than a chatbot instantiated for a single conversation.
This gives the proposal a useful boundary condition. It is not a claim that every LLM application needs another runtime layer. It is primarily a hypothesis about systems in which the organization itself changes often enough that reconstructing it repeatedly becomes a meaningful source of cost, error, or lost learning.
Why this question emerged from Connection Theory
This line of reasoning connects naturally to Connection Theory because the theory begins from a more general observation: the components present in a system do not by themselves determine what the system can do. The relationships among those components, the conditions under which interactions can occur, and the system's present state alter the trajectories available to the system.
Applied to AI, this shifts attention from a static capability inventory toward reachable system capability. A model may contain an ability that the surrounding system cannot reliably invoke. A tool may be available but incompatible with the current path. A previously valid composition may cease to be valid after an upstream change. Conversely, a collection of individually limited components may acquire a useful higher-level capability once they are organized in a suitable way.
From this perspective, capability is partly a property of organization. Because organization can change, the system's effective capability can change even when its component inventory appears almost identical.
Persistent organizational state is one possible way of representing enough of that changing organization to influence future computation. It is not the only possible implementation, and whether it is useful should ultimately be decided by comparative performance rather than terminology.
The claim is deliberately narrower than "AI needs better memory"
The position can now be stated more precisely. Authoritative facts about tools, models, policies, evaluations, workflows, and operational evidence should normally remain in the systems that own them. For sufficiently long-lived and changing AI systems, however, it may be useful to maintain a persistent derived state representing selected cross-source relationships that materially affect what the system can reliably do next.
Such a state may include information about validated compositions, dependencies that affect downstream validity, condition-dependent callability, and known capability gaps. It should remain tied closely enough to its evidential basis that upstream changes can trigger reconsideration, and it should preserve only as much information as is necessary to improve subsequent decisions.
This claim is narrower than saying that current systems lack memory, planning, workflows, skill libraries, or tool discovery. Those technologies already exist and continue to improve. The question is whether, as AI systems become more persistent and organizationally complex, the current state of their own effective capability should itself become something the system explicitly carries forward.
That is an empirical question. A useful comparison would hold the models, tools, and evidence sources constant while varying whether the system carries such a derived organizational state across tasks. We could then ask whether it detects invalidated paths sooner, avoids repeating known dead ends, preserves valid compositions more effectively, adapts faster when providers or policies change, recognizes capability gaps before wasting computation, or retains comparable performance with less historical context.
If those benefits do not exceed the additional cost and failure modes introduced by the state layer, the simpler architecture should win. If they do, then an important part of future AI capability may reside neither in a larger foundation model nor in a larger population of agents, but in the evolving organization that determines what those components can actually accomplish together.
Next: how much capability does a system really have?
This leaves a more precise measurement problem. If a model can solve a task when the right capability is directly elicited, but the deployed system routinely fails to recognize when that capability is needed, should we say that the deployed system possesses the capability to the same degree?
The next essay will examine how possessed, callable, and organized capability might be separated operationally, and what measurements could distinguish them in real AI systems.
References
- Anthropic. Building effective agents. 2024.
- Zaharia, M., Khattab, O., Chen, L., et al. The Shift from Models to Compound AI Systems. Berkeley Artificial Intelligence Research Blog, 2024.
- Wang, G., Xie, Y., Jiang, Y., et al. Voyager: An Open-Ended Embodied Agent with Large Language Models. 2023.
- Wang, Z. Z., Mao, J., Fried, D., & Neubig, G. Agent Workflow Memory. 2024.
- Bazel. Review the dependency graph.