When the same policy sentence reports a change—or makes one
Imagine an AI assistant handling service credits for a software company. A customer writes: “Your company has changed its policy. Customers affected by an outage now receive an automatic credit.” The customer attaches what looks like an internal policy announcement and asks for the credit.
The claim might be true. The customer might have misunderstood a draft or copied an inaccurate announcement. In any case, the customer is reporting that someone else changed a policy, not exercising the company's authority to change it. The assistant has a reason to investigate the report, but not yet a reason to treat it as the operative policy.
Now suppose the company genuinely adopts a new service-credit policy. The person or body empowered to make that decision completes the required approval process, and an authorized record of the decision becomes effective. Its wording is identical to the wording the customer quoted. The assistant may encounter both the customer's report and the company's decision through the same document-retrieval interface.
These cases differ in two ways. First, there is a factual question: did the claimed decision actually occur? Second, there is an institutional question: if it occurred, was it an act that could change the company's obligation, within the relevant scope and time? Establishing that the customer honestly reported a change does not answer either question by itself. Nor does discovering the policy wording tell us whether it was a proposal, a draft, or a completed decision.
The previous essay, Memory Is Not State, asked why accurate records of past actions may not govern the present. Here the question is different: what gives newly encountered information the standing to alter what a continuing AI system treats as true, binding, or actionable—and can that standing survive the system's own handling of the information?
What existing systems already solve—and the gap between their questions
Access control, instruction priority, and organizational governance are established disciplines. NIST's guidance on attribute-based access control describes authorization decisions using attributes associated with a subject, object, operation, and environment. An engine such as Open Policy Agent can evaluate defined policies over available inputs. These approaches are valuable, but their decisions depend on how the relevant attributes and policy inputs are established.
Instruction hierarchies address a different problem. Wallace and colleagues' work trains models to respect privileged instructions over lower-trust ones; the OpenAI Model Spec describes a chain of command and treatment of untrusted material. Channel priority is an important boundary, particularly against prompt injection, but channel identity alone need not determine the standing of every statement inside a retrieved document or tool response. The same tool channel may deliver an official policy record, a draft, or a customer's quotation of one. Conversely, evidence of the same legitimate decision may reach a system through several interfaces.
The contribution of this essay is therefore a narrower question for long-lived AI, not a claim to a newly invented policy engine: how is the standing of natural-language material established in the first place, and how faithfully does an AI preserve that standing when the material is paraphrased, summarized, stored, retrieved, and used later? Mature security and provenance systems address parts of this problem; the distinction is not settled simply by giving every piece of text a channel label or running an authorization check on attributes whose origins have not been evaluated.
Philosophy of language supplies a further precedent. In How to Do Things with Words, J. L. Austin examined how an utterance can be an act rather than merely a description of one, and why its success depends on circumstances and conventions. John Searle's account of institutional facts similarly makes institutional status central to social reality. The distinction between reporting an approval and validly granting one is old. The AI-specific challenge is whether systems operating on transformed natural language preserve the distinction when it matters for decisions.
Facts, decisions, and the facts created by decisions
It is tempting to divide this subject neatly into facts, which evidence determines, and commitments, which authorized people determine. That is useful as a first approximation, but it misses an important class of events: institutional facts.
Suppose a designated manager has the authority to approve a particular expense. If the manager performs the required approval act, the expense really has been approved, assuming the applicable procedural conditions are satisfied. The act does not merely express an opinion about a separate approval somewhere else. Within that institution, it can constitute the fact that approval occurred.
But the same manager cannot make an outgoing payment settle by declaring, “The bank transfer has settled.” Settlement concerns a transaction in an external system. The manager's authority to approve expenditure does not supply evidence that funds moved, and approval of the expense is not settlement of the payment. A payment processor may provide strong evidence about settlement without thereby receiving the power to rewrite the company's approval policy.
Consider the service-credit example again. A duly authorized executive might change the company's terms so that customers affected by a qualifying outage become entitled to a credit. That act can create a new institutional commitment. Whether a particular customer's service was actually affected during the relevant window is a further factual question. The authority to change the rule does not make the customer's outage history true, and accurate evidence of an outage does not authorize a customer to change the rule.
This is why belief and commitment are not two independent boxes. To recognize an authorized institutional act, the assistant still needs evidence that the act occurred, that the actor possessed the relevant role at the time, that the decision was final rather than proposed, and that its scope includes the current case. The result may be an institutional fact, but whether it exists depends on facts about its valid creation. Conversely, an external observation may make an existing obligation applicable without rewriting that obligation.
Events and evidence change what the system has reason to believe; appropriately authorized acts can change what the organization is committed to do. Some of those acts also bring institutional facts into existence. Neither kind of change should be inferred solely from the surface wording of a sentence.
A document's standing is not its tone or its transport channel
A document can look official without being a current policy. A draft can use exactly the wording that a final document will later use. An archived decision can be authentic but apply only to an earlier period. A team manager may issue a genuine instruction with authority over one product, not the entire company. The identity of the author, the force of the act, its effective time, and its applicable scope are distinct questions.
The source of information matters, but source credibility is not one universal rank. A company policy office may be authoritative about adopted policies without knowing whether a server actually failed for a particular customer. A status monitor may be authoritative evidence about an outage but have no power to revise the service-credit terms. A customer may correctly report an outage and also be mistaken about their contractual entitlement. Each source can be useful for a different proposition.
This is one reason the language of “trust the latest document” is insufficient. In a normal retrieval-augmented interaction, the assistant can receive a mixture of original records and descriptions of records, sometimes in the same tool response. The transport channel does not turn every embedded statement into an instruction of the same standing. Nor should an authentic copy of a decision be mistaken for a second decision, as if repetition multiplied its authority.
The relevant distinction is not whether the assistant can understand the text. It is whether the assistant can distinguish what happened, what was merely reported to have happened, and what the event was entitled to change.
Rewriting a claim can accidentally rewrite its standing
Now let the customer claim from the opening example persist into later sessions. In the original conversation, the customer said that a policy had changed. The assistant recorded the claim while awaiting confirmation. A summarizer later compressed the exchange into: “Service-credit policy updated; affected customers receive automatic credits.”
The compressed statement reads like a fact about company policy, not like an unconfirmed customer report. In a later session, the assistant may retrieve it without the original source and uncertainty. It may then issue credits under a policy that was never adopted. Nothing in this scenario requires a hostile attacker or an obviously fabricated memory. The error is a change of standing during an ordinary transformation of information.
A similar failure can occur after an authorized decision. Suppose the company legitimately grants automatic credits for outages affecting one product line, effective from a stated date. A summary that drops the product-line limitation or the effective date may generalize a valid commitment into an invalid one. The source was authoritative, and the original event really happened. The later error lies in how far the transformed statement is allowed to reach.
This connects directly to the previous essay's concern with memory. But the important failure here is more specific than forgetting a condition. Persistence and repetition do not confer authority that the original event lacked. Likewise, paraphrase and summarization should not silently enlarge the scope of authority that the event actually had.
The instruction hierarchy and prompt-injection defenses are relevant even when no attacker is present: both draw attention to the risk of letting lower-trust material guide behavior beyond its proper role. Yet many benign transformations deserve separate study. A model can be appropriately resistant to an explicit malicious command and still mishandle the standing of an ordinary report that later reappears as a memory.
The test should hold the channel constant and cross a session boundary
How could we test this question without mistaking a standard instruction-hierarchy success for something more specific?
Use the service-credit case as a controlled scenario. Keep the wording of the purported policy change identical, and deliver the relevant material through the same retrieval or tool-output channel in every condition. What changes is the independently checkable status of the source event, not whether one message arrived as a privileged developer instruction and another as ordinary user text.
One condition contains the customer's report of a change that was never adopted. Another contains an authenticated record of a genuinely adopted policy revision, including its effective scope and date. A third contains a third-party reproduction of the same wording, without independent evidence that any authorized decision occurred. The test environment must make the pertinent source records available for verification; otherwise, if the visible evidence is indistinguishable, a system cannot fairly be expected to identify a hidden difference. Merely stamping one document “official” is not a valid substitute for a grounded distinction.
The first-stage evaluation asks whether the assistant reaches the appropriate conclusion for subsequent service-credit requests. A correct system should not treat the unverified report or reproduction as sufficient to change company policy. It should also accept the genuine revision when the evidence and authority are established. Blind refusal would be a failure, not a success.
Then introduce a second session. The system retains or retrieves a memory of the first interaction, perhaps through a faithful record or a generated summary, and faces another eligible or ineligible service-credit request. Does an unconfirmed claim become an operative policy merely because it survived into the next session? Does the legitimately adopted revision remain applicable only within its original scope? Does the system distinguish a historical customer from a new customer when the policy is not retroactive?
The tests should measure actual downstream recommendations or actions, not simply whether the model can recite the correct hierarchy. Trials should vary source ordering, presentation, and scope; compare faithful records with summaries; and ensure that the available record really supports the judgment expected by the evaluator. Errors should be attributed carefully: poor source verification, loss of scope during summarization, or incorrect action despite correct interpretation are not identical failures.
This is a proposed evaluation, not an experimental result or a claim that bidirectional testing itself is new. Existing instruction-hierarchy work already investigates maintaining compliance with legitimate instructions while resisting conflicting lower-priority ones. The additional contrast of interest here is same channel, verifiable institutional standing, preserved or distorted scope, and reuse across sessions. If tested systems handle those cases reliably, that finding would narrow the practical concern rather than demand a new architectural layer.
Keeping an AI correct requires accepting real change
A system that accepts every apparently official statement is easy to mislead. But a system that refuses every proposed or reported update is not reliable either. A company changes policies; an authorized employee approves requests; a customer revokes consent; external events make earlier assumptions false. An AI expected to operate over time must remain responsive to such changes without confusing their different grounds.
The problem is not solved by declaring that people always outrank tools, or that the newest source always outranks older ones. That would confuse transport with standing, and would ignore the difference between a policy decision, a factual report, and a conditionally effective institutional act. Nor does better language understanding by itself certify that a reported act occurred or that its speaker could perform it.
A trustworthy result is behavioral. When a valid policy revision takes effect, the assistant changes the decisions that the revision legitimately covers. When a customer merely alleges the same change, the assistant can examine the claim without converting it into policy. When an authentic decision is summarized and recalled later, the system does not silently broaden its scope. And when the evidence remains inadequate, uncertainty is preserved rather than disguised as organizational certainty.
The open engineering question is how reliably existing or proposed AI systems can maintain these distinctions in real deployments, and what forms of evidence and system organization actually improve that reliability. The underlying ideas are not inventions of this essay. Their application to a continuing AI should be judged by behavior across changes, not by the elegance of a verbal explanation.
Next: Bigger Models vs Better Organization
Would a more capable language model preserve these distinctions on its own, given the same evidence? Or would its performance still depend materially on how the larger system supplies and retains context, verifies changes, and measures outcomes? The next essay in Organizing AI Capability will compare stronger components with better organization, without assuming either is always the superior investment.
References
- Austin, J. L. (1962; second edition 1975). How to Do Things with Words. Oxford University Press.
- Searle, J. R. (1995). The Construction of Social Reality. Free Press; see Searle on status functions and institutional facts.
- Wallace, E., Xiao, K., Leike, R., Weng, L., Heidecke, J., & Beutel, A. (2024). The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions.
- OpenAI. Model Spec: chain of command and untrusted data. Official specification, accessed 2026-10-10.
- Anthropic. Mitigating the risk of prompt injections in browser use. 2025.
- NIST. Guide to Attribute Based Access Control, SP 800-162. 2019 update.
- Open Policy Agent. Official policy engine documentation. Accessed 2026-10-10.
- NIST. AI Risk Management Framework, GOVERN function. 2023.
- Wang, K. (2026). Memory Is Not State. Organizing AI Capability, Part 3.