Part 2 of Organizing AI Capability

Possessing a Capability Is Not the Same as Being Able to Use It

A capable model can fail in deployment because someone—or something—must still recognize, invoke, and integrate its ability. The difference becomes visible when we examine what an evaluator has supplied and what the AI system has actually done.

What changes when the evaluator stops doing part of the work?

Imagine an AI assistant helping a customer change a flight. The system has access to a reservation database, a flight-search service, fare rules, and a tool that can modify bookings. In separate tests, its tools work correctly. The model can interpret fare rules when the relevant policy is supplied, and a specialist can compare suitable flights when given properly prepared inputs.

Now the customer asks to move a booking to Friday without paying an unnecessary fee. The deployed assistant overlooks the applicable fare rule, chooses the wrong tool, or fails to check whether the reservation was actually updated. The system has access to components that appear sufficient for the task, yet the customer does not receive a reliable result.

What failed? The question matters beyond debugging a single interaction. If a long-lived AI system retains what it has learned about its capabilities, it needs to distinguish an ability that a component can demonstrate under favorable conditions from one that the deployed system can recognize, invoke, and use correctly. Otherwise, it may preserve the wrong lesson: retrain a capable model because the router failed, or trust a workflow because its final message sounded successful.

The first essay in this series asked whether a changing AI system should preserve a current view of its effective capability. This essay addresses a prerequisite: what evidence would justify calling a capability available to the system, rather than merely demonstrable by one of its parts?

A test result depends on what the evaluator supplied

Suppose an evaluator sends a specialist model an already-extracted contract clause, the relevant policy passage, and a clear instruction to classify a particular exception. The specialist answers correctly. We have learned something useful, but we should be precise about what it is: the model performed this transformation with a particular input, prompt, interface, and amount of assistance.

The evaluator may have silently performed several difficult steps before the model saw the task. Someone recognized the legal issue, selected the specialist, found the relevant document, prepared the input, and decided how to check the output. When the same problem arrives as an ordinary user request, the deployed system has to perform some or all of those steps itself.

This is why the word possessed is only shorthand in this article. A “possessed capability” is a capability demonstrated under specified testing conditions, often after the evaluator has already supplied part of the route to success. It is not a claim that the capability exists as a context-free property hidden inside the model.

Capability elicitation research has made the general problem explicit. METR's guidance on capability elicitation distinguishes what a model does under a particular setup from what may become accessible with better prompting, tools, or scaffolding. That work is largely concerned with evaluating potentially elicitable model capability. Our interest here is what happens after the remaining selection and coordination work is handed back to a deployed system.

A direct evaluation and a deployed task may therefore differ even when both use the same model weights. The difference is not mysterious. The evaluator has stopped doing some of the work.

Three questions, organized by how much help the system receives

A practical way to separate these cases is to ask what the test has already decided on the system's behalf.

At the first level, demonstrated component capability, the evaluation supplies a suitable component and an adequately prepared input. The question is whether that component can perform the relevant transformation under those stated conditions. This is the limited sense in which we can say the system possesses a capability: one of its parts has demonstrated it when the route to that part has largely been provided.

At the second level, callable capability, the system receives a normal request and must decide whether a relevant capability should be used, which provider is suitable, and whether that provider can actually be invoked. The evaluator no longer supplies the correct choice. Success includes recognizing when not to call a tool—for example, when required information is absent or authorization has not been established. A technically functional tool that cannot be reached under current conditions does not provide a callable capability for that task.

At the third level, organized capability, the system must carry the task through to a valid outcome. It has to use intermediate results coherently, respect dependencies and constraints, and establish that the required effect occurred. A correct tool selection and a correct local output may still fail to produce a correct final result if information is handed off badly or the action is never committed.

These levels are ways of arranging evidence, not mutually exclusive substances or a universal formula for total capability. They overlap: a complex component can itself contain internal organization, and a system can solve the same task through several alternative routes. Their value is diagnostic. When an observed success rate changes, we can ask which work has been transferred from the evaluator to the system, and at which point that transfer becomes unreliable.

The same visible failure can have different causes

Return to the flight assistant. Suppose it responds that a fare change is impossible, although a compliant, fee-free option exists. One explanation is that the fare-rule specialist cannot interpret the rule. Another is that the specialist could do so but was never invoked. A third is that the specialist returned the right interpretation, only for the assistant to select an incompatible replacement flight or mishandle the booking update.

All three cases produce an unhappy customer, and an end-to-end score may mark all three as failures. Yet they suggest different remedies. A better specialist might address the first. Better selection or access might address the second. The third calls for examination of the interaction among components, state, and verification.

We can make this distinction more than verbal without turning an essay into a new benchmark specification. Take the same failed tasks and intervene selectively. If we force the correct provider to be selected while leaving the rest of the system unchanged, do some failures disappear? If we instead supply the correct intermediate handoff and let the remaining workflow continue, do further failures disappear? These are simple counterfactual checks for where a capability was lost between demonstration and result.

The changes in success are not automatically additive, and no single intervention proves a unique cause: repairing one stage can expose a different downstream failure. Nevertheless, the contrast is more informative than treating every missed booking as evidence that the entire model needs to be retrained.

This is an ordinary engineering principle applied to a more complicated object. Unit, integration, and end-to-end tests have long existed; what matters for AI capability attribution is which parts of the successful procedure were supplied by the test and which the deployed system had to discover or execute for itself.

Reliable non-use is part of successful calling

There is an important edge case. Imagine that the booking tool can modify a reservation, but the customer has not passed the required identity check. The assistant should not call the tool merely because it can. Correct behavior may be to request verification or decline the action.

That belongs in the second level of the distinction, not in a separate category. Callability includes applicability and authorization: deciding whether to use a capability is part of using it well. The system should invoke a suitable provider when the task calls for it and refrain when the requirements are not met.

This is not an invented concern. The Berkeley Function-Calling Leaderboard evaluates aspects of function relevance and abstention as well as tool invocation. ToolSandbox also examines state dependencies and cases with insufficient information. These evaluations already make it harder for indiscriminate tool calling to masquerade as competence.

The same logic applies when the problem is epistemic rather than procedural. If two competing answers cannot be distinguished using the available evidence, selecting one confidently is not a stronger demonstration of capability. Depending on the task, the useful action may be to state the uncertainty, request missing evidence, or identify a way to obtain it. Recognizing that boundary is part of deciding what capability can responsibly be brought to bear.

A convincing account of an action is not the action

The third level needs an especially concrete success condition. An assistant may say, “Your flight has been changed,” after generating a plausible tool call. That statement is not evidence that the reservation database contains the requested booking, and even a successful database update might violate a fare restriction or the customer's stated requirements.

For an action task, the relevant organized result therefore has to be checked against the actual environment and the task's constraints. The record of what the system attempted helps diagnose failure, but the resulting state is what determines whether the intended operation occurred. This is the evidence standard for organized capability; it does not require introducing another independent set of capability categories.

The τ-bench benchmark illustrates the point by evaluating agents in airline and retail interactions and comparing the resulting database state with the intended goal. It also introduced pass^k to examine whether agents succeed consistently over repeated trials, rather than merely succeeding at least once. The public τ-bench implementation now includes error-identification tooling that distinguishes, for example, wrong-tool use from wrong arguments and unintended actions, although automated error labels themselves require scrutiny. Anthropic's agent-evaluation guidance, published January 9, 2026, similarly separates the execution transcript from the resulting environment state.

These examples are important because they show that pieces of the diagnostic approach already exist. The present argument is not that benchmarks have failed to notice tools, policies, or external outcomes. It is that a component's demonstrated competence, a system's success in invoking it, and the final result should not be silently merged into one undifferentiated statement that “the AI has this capability.”

Reliability and attribution depend on the system boundary

A capability demonstrated in one successful run may not be dependable across repeated requests. That matters when the system is expected to serve users continuously rather than produce one acceptable answer after several attempts. A flight assistant that occasionally reaches the right result may still be unsuitable for changing real reservations without supervision.

The system boundary matters as much as repetition. A tool-assisted AI product should be evaluated with the tools that genuinely belong to the product. But if the correct answer depended on a human evaluator supplying the route, verifying the policy, or silently correcting an intermediate result, that assistance should remain visible in the claim about what the AI system accomplished.

The same care is needed when comparing architectures. If a multi-component system outperforms a simpler baseline only because it received more information, more tool access, or a much larger computation budget, the observed advantage cannot be attributed to organization alone. Different organizational designs may still be useful, but their contribution has to survive a comparison that accounts for these other causes.

None of this implies that every application requires exhaustive causal diagnosis. For a small, stable workflow, a reliable end-to-end test may be all its operator needs. The finer distinction becomes valuable when failures recur, components change, or engineers must decide whether to improve a provider, the way it is invoked, or the system that coordinates several providers.

What should a long-lived AI system learn from the failure?

Now the connection to persistent state becomes concrete. If the flight assistant fails to change a reservation, recording only “flight-change task failed” discards information that could matter next time. A capable specialist that was not selected is a different operational problem from a selected specialist whose output was mishandled downstream. Treating them alike can cause the system to keep changing the wrong thing.

The converse is equally important. If a booking succeeded only because an evaluator supplied the correct route or a human repaired the final step, storing that event as proof that the deployed AI can reliably complete the task would inflate its capability record. A system that carries forward such judgments may become more confident while its real performance remains unchanged.

This gives us a practical reason to separate demonstrated, callable, and organized capability. The point is not to attach three impressive labels to one task. It is to preserve distinctions that change what a system should check, improve, reuse, or question after an outcome.

A persistent representation of effective capability may eventually be useful for that purpose, but it should not be credited with abilities that have only been shown under evaluator assistance. Before an AI system can learn accurately about what it can do, its evaluations must establish what actually succeeded, under which conditions, and with whose help.

That is the useful takeaway from the distinction. A system's components may be capable, yet the organization can fail to bring their capabilities to bear. Conversely, a well-organized system may accomplish tasks beyond what one isolated component could do, provided we count the resources and external assistance honestly. Neither result can be inferred from a tool list or a model score alone.

Next: Memory Is Not State

If the system learns that a specialist is competent but hard to invoke, or that a previously effective workflow no longer produces the intended result, it has acquired information that may improve later decisions. But merely preserving the record of that event does not guarantee that it will change what the system does next.

The third essay in Organizing AI Capability will examine that distinction: when does information retained from the past become part of the current operating state of an AI system?

References

Continue this subject

Organizing AI Capability