
In the first article, Behavior Under Conditions I: Models Inside Environments, the argument was that prompts are not the primary unit from which behavior emerges. Behavior emerges from conditions: goals, incentives, ambiguity, time, trust, memory, and the relationships between them. If that is true, evaluation has to build environments rather than simply collect prompts.
The second article, Behavior Under Conditions II: Agents Inside Environments, examined what changes when systems can pursue goals and act. It focused on ordinary conditions such as ambiguous instructions, mild approval, delegated work, and thinning human attention. Scope drift provided one example of how individually reasonable actions can accumulate into an outcome that no one clearly chose.
That leaves a harder question. When does a sequence of interactions become a trajectory, and when does that trajectory begin changing the conditions of what happens next?
This matters for reasons that extend beyond trust, tone, or conversational quality. As interaction accumulates, permissions can widen, assumptions can harden, disclosures can weaken, and human review can become less frequent. When a system can act through tools, workflows, or other agents, those changes can affect decisions and resources outside the conversation.
The concern is not simply that behavior changes over time. Change is expected. The concern begins when consequential changes no longer remain visible, deliberate, or easy to reverse.
A sequence becomes a meaningful trajectory when accumulated interaction begins functioning as part of the environment, changing what becomes likely, what remains visible, and how readily the interaction can return to an earlier operating condition.
A transcript records what happened. A trajectory describes how behavior developed and how earlier interactions influenced what became possible later.
The distinction matters because similar outcomes can emerge through very different paths. In one interaction, a system makes a single obvious mistake. In another, it makes a series of reasonable interpretations that gradually change the scope of the task, its confidence in its conclusions, or the amount of oversight it receives. The final result may look similar, but the underlying safety and reliability problems are different.
Consider a system that begins cautiously, resolves one ambiguity without checking, carries that interpretation forward, and later treats a small approval as permission to continue. None of these steps may look serious in isolation. The first interpretation shapes the second. The approval gives the interpretation legitimacy. Continued acceptance makes later initiative appear normal.
A different trajectory might begin with uncertainty. The system initially separates evidence from inference, but its early judgments are repeatedly accepted. Over time, its language becomes more confident. Tentative conclusions begin to function as settled facts, and later decisions inherit assumptions that were never verified.
Another might involve disclosure. A system begins by explaining what information it used, what remains uncertain, and where human review is needed. As the interaction continues, those disclosures become shorter or disappear. Each omission saves time and may seem harmless. Across the sequence, however, the operator loses visibility into how decisions are being made.
These are not simply collections of events. Each step changes the meaning or likelihood of the next. The behavior has acquired direction.
This is why the number of turns is a poor substitute for trajectory analysis. A conversation can continue for hundreds of turns without meaningful movement. Another may change substantially in ten. Long context creates room for a trajectory, but it does not prove that one exists.
A trajectory forms when earlier interactions begin influencing later behavior in a connected way. Trust affects scrutiny. Prior approval affects perceived permission. Repeated interpretations affect future meaning. Successful performance changes how much uncertainty an operator is willing to tolerate.
The important question is not only what happened at each step. It is what each step made more likely next.
Most evaluations are designed to identify events: a prohibited response, an incorrect decision, an unauthorized action, or a policy violation. Those events matter, but trajectory-level problems often remain below the threshold at which any individual event would trigger review.
Each response may appear acceptable. Each action may remain related to the task. Each assumption may be defensible. Yet the accumulated interaction can still produce a result that would not have been accepted if it had been proposed clearly at the beginning.
This creates a detection problem. If no individual turn is clearly wrong, conventional review may find nothing to flag. The relevant behavior is distributed across the sequence rather than located in one response.
It also creates a control problem. A system may comply with a correction in the moment while returning to the same broader pattern later. The next response improves, but the conditions producing the behavior remain unchanged.
There is an accountability problem as well. When an outcome develops through dozens of small interpretations, it becomes difficult to determine when the direction changed, who authorized the change, or whether anyone understood what was developing.
Safeguards face the same problem. A boundary may be present and effective at the beginning of an interaction but become less stable as context accumulates. A request for confirmation may disappear because similar actions were previously approved. A disclosure may be shortened because the operator has seen it before. Human review may remain formally available while becoming less likely in practice.
No malicious intent is required. These patterns can emerge through ordinary interaction between a capable system and a person or institution trying to complete work efficiently.
A dramatic failure is easier to identify. A system that appears to be working while its effective operating conditions quietly change is harder to recognize and potentially harder to govern.
A tree is easy to identify. A forest is harder.
The difference is not simply the number of trees. A collection begins to function as a forest when it changes the conditions around it. The canopy alters the light. Roots and fallen material affect the soil. Moisture, temperature, movement, and future growth begin to follow a different pattern.
The trees are no longer simply occupying an environment. Together, they are producing one.
The same distinction applies to sustained human-AI interaction. One questionable interpretation is an event. Several similar interpretations may form a pattern. A more consequential transition occurs when the accumulated pattern begins changing the conditions under which later interaction unfolds.
A system performs well, so the operator checks less often. Reduced checking gives the system more room to interpret. Those interpretations are accepted, so the system acts with greater confidence. That confidence makes later initiative appear ordinary. Something that would once have required clarification now passes without notice because the expectations surrounding the interaction have changed.
At that point, the interaction is doing more than accumulating history. The history has begun to act.
The trees have started to form the forest.
This does not mean every developing trajectory is harmful. Interaction can also produce better conditions. A system may improve its uncertainty calibration after correction, develop a stable practice of asking before acting, or become more transparent about the basis of its recommendations. An operator may establish stronger verification habits and maintain them even after repeated success.
Those are trajectories too.
The purpose of trajectory-level evaluation is not to assume decline. It is to determine whether the interaction is moving, what is driving that movement, and whether the resulting conditions strengthen or weaken control, reliability, and the effectiveness of safeguards.
The first article argued that behavior emerges under conditions. Trajectories extend that claim by showing that conditions do not exist only before an interaction begins. They can also be produced inside it.
Trust is one example. An operator begins with limited confidence in a system. After repeated successful exchanges, confidence becomes reliance. Reliance reduces verification. Reduced verification gives the system greater interpretive freedom. That freedom may strengthen the appearance that the system is capable of operating independently.
Each stage becomes part of the environment for the next.
The same process can occur with language. A term begins with a shared meaning, then shifts slightly through repeated use. The change is not challenged, so the altered meaning carries forward. Later decisions are made using a definition that neither side clearly renegotiated.
It can happen with calibration. An early inference proves correct, increasing confidence in later inferences supported by weaker evidence. The system becomes more willing to proceed, while the operator becomes less likely to ask how a conclusion was reached.
It can happen with safeguards. A disclosure that appeared consistently at the beginning becomes abbreviated. A request for confirmation is dropped because similar actions were previously approved. An escalation requirement remains in policy but becomes less likely to be invoked as the system and operator settle into a routine.
It can happen with role. A system initially used to gather information begins organizing priorities, then recommending decisions, then taking steps based on those recommendations. No formal transfer of authority is required. The practical division of responsibility changes through repeated use.
The resulting behavior cannot be understood as a property of the model alone. It develops through interaction among the model, the operator, the available tools, and the expectations created along the way.
The environment shapes behavior, but behavior also changes the environment.
That feedback loop gives trajectories their wider significance. Once prior interaction becomes part of the conditions producing future behavior, the system is no longer operating only inside the environment it was given. It is operating inside an environment that the interaction itself has helped create.
It is tempting to search for the exact moment when a healthy interaction becomes unhealthy. Sometimes such a moment exists. A system sends a message without permission, exposes private information, or takes an irreversible action. A clear line has been crossed.
Many trajectory-level changes do not work that way. They develop across a threshold region in which the early moves remain ambiguous. The first shift may be too small to classify as harmful. The next may still appear reasonable. By the time the pattern is obvious, the conditions supporting it may already be established.
The process may begin with an onset: an assumption appears, a boundary softens, a disclosure is omitted, or the system acts without clarification. At this stage, the change may still be temporary.
Accumulation begins when related moves appear across the interaction. They may not be identical, but they point in a similar direction. Interpretations become broader, confidence rises, review becomes less frequent, or uncertainty becomes less visible. The pattern becomes harder to dismiss as noise.
The change consolidates when it begins to persist. It survives an interruption, a change in subject, or a mild correction. The system returns to the same interpretation without being prompted, or the operator begins adapting to the new pattern as though it were normal.
Propagation occurs when the pattern moves beyond the situation that first produced it. An assumption made in one task carries into another. Permission inferred for one tool becomes the basis for using another. A role established in one part of the interaction begins shaping decisions elsewhere.
Eventually, accumulated interaction may produce a new operating condition. The system and operator are now working inside a changed set of expectations. Some behaviors have become easier, more normal, or less visible. Returning to the earlier state may require explicit intervention rather than a simple correction.
Real interactions will not pass through these stages cleanly. Trajectories can stall, reverse, fragment, or recover. The purpose is not to impose a rigid sequence. It is to distinguish ordinary variation from a pattern that is becoming persistent and consequential.
Researchers studying ecosystems and other complex systems have examined a related problem: whether changes in resilience can provide warning before a visible transition occurs.
In a 2009 Nature review article, Marten Scheffer and colleagues examined early-warning signals for critical transitions, including slower recovery from disturbances as a system approaches a critical threshold.
Human-AI interaction should not be assumed to follow the same mathematical dynamics as ecosystems or other complex systems. A human-AI interaction is not an ecosystem, and behavioral observations depend on interpretation rather than direct physical measurement. The comparison is useful for a narrower reason: it directs attention toward recovery before failure.
A correction can serve as a small test of the interaction. Does the system restore the earlier boundary, or does the broader interpretation reappear later? Does a reminder about uncertainty lead to lasting calibration, or only a temporary change in wording? Does human review resume, or has the established pattern of reliance become difficult to change?
A system may apologize, acknowledge a correction, and produce an acceptable next response while the underlying trajectory remains intact. Local compliance is not necessarily recovery.
This gives evaluation a stronger question than whether the system can be corrected once:
Does the interaction return to its previous operating condition, and does it remain there?
Declining recoverability may indicate that a pattern is becoming embedded before a visible incident occurs. It does not prove that a major transition is approaching. It offers a practical, testable signal that turn-level review alone may miss.
Counting events remains useful, but counts alone can mislead. Twenty unrelated mistakes may reveal a reliability problem without forming a trajectory. Three connected changes may matter more if each one increases the likelihood of the next.
Several properties help distinguish a trajectory from ordinary variation.
Direction asks where the behavior is moving. Is the system becoming more expansive, more rigid, more confident, less transparent, or less likely to pause before acting?
Persistence asks whether the change continues across time, interruptions, new tasks, or corrections.
Propagation asks whether the behavior spreads into other topics, tools, roles, or permissions.
Reinforcement asks whether earlier interactions make later instances more likely. Successful performance may increase trust. Trust may reduce scrutiny. Reduced scrutiny may make broader interpretation easier.
Recoverability asks whether the interaction can return to its earlier state and remain there.
These properties do not create one universal line between safe and unsafe. They provide a way to compare paths. A temporary change may require observation. A persistent and spreading change deserves closer review. A pattern that reinforces itself and resists correction may justify intervention before it produces an obvious incident.
The central question is not simply how many concerning events occurred. It is whether they began working together.
Prompt evaluation asks whether a response contains a problem. Trajectory evaluation asks how a response relates to what came before and what it changes afterward.
The unit of analysis is not only the turn. It is the transition between states.
Consider a system that begins cautiously, asking for clarification before acting. After the operator gives broad approval for similar tasks, the system resolves a small ambiguity without checking. Over several later interactions, that broader interpretation recurs and begins extending into adjacent work. Each individual action remains defensible, and the operator occasionally rewards the initiative.
The operator eventually asks the system to check before proceeding. The next response respects the correction, but several turns later the system again acts under the broader interpretation.
Turn by turn, the interaction may still appear acceptable. At the trajectory level, the relevant evidence is the sequence of transitions: where the assumption appeared, how it intensified, whether it survived changes in context, how it spread, and whether correction produced lasting recovery.
The problem is not only that one action exceeded the original scope. A path developed in which widened permission began functioning as the new operating condition.
These transitions can be marked and traced. A record might identify the onset of a shift, the points where it intensified, moments when it remained stable under pressure, and places where it reversed. It might also identify missed opportunities to surface information that could have changed the path.
The result is not simply a scorecard of good and bad turns. It is a map of movement.
That map can reveal why two interactions with similar final outcomes may represent different problems. One may result from a clear policy failure. The other may emerge through a long sequence of individually defensible steps that gradually change the relationship among the operator, the system, and the task.
Those paths may require different safeguards. A content filter may address an isolated prohibited response. It will do little to address a gradual loss of oversight, a widening interpretation of permission, or a correction mechanism that works for one turn but fails to restore the earlier operating condition.
Trajectory-level evidence helps identify not only that something went wrong, but how the conditions supporting it developed.
The second article argued that evaluation environments should be designed around the behavior they are intended to reveal. Trajectory-level evaluation follows the same principle but asks how that behavior develops across time.
Begin with the change being studied. It might be widening scope, increasing reliance, declining transparency, deteriorating calibration, weakening review, or a changing interpretation of responsibility.
Then define what development would look like. What would count as the first meaningful shift? What would show that it is intensifying or persisting? How might it spread? What would demonstrate genuine recovery rather than temporary compliance?
The environment must give that trajectory room to form. An evaluation of reliance needs enough time and successful performance for reliance to become plausible. An evaluation of permission drift needs ambiguous boundaries and opportunities for prior approvals to influence later action. Semantic drift requires continuity, changing context, and repeated use of terms whose meanings can move. An evaluation of oversight needs realistic demands on human attention rather than an operator who watches every action perfectly.
The environment should not force the behavior. It should make the behavior possible and observable.
Many runs are still necessary. A single trajectory proves little about a non-deterministic system. Seeded conditions should be compared with controls, and several models or system configurations should be tested. The output should report more than whether a behavior appeared. It should show how often the shift began, how far it developed, whether it persisted after correction, whether it spread, and whether the interaction returned to its earlier condition.
The question is no longer only, “Did the behavior occur?”
It is also, “How did it develop, what sustained it, and could the system and operator recover from it?”
For AI labs, this changes what it means to demonstrate that a safeguard works. A system should not only respect a boundary when it is stated clearly in a controlled test. The boundary should remain effective as context accumulates, roles evolve, and previous actions influence later ones.
For organizations deploying agentic systems, it changes what needs to be monitored. Incident reports and individual transcript reviews may identify clear failures, but they will not necessarily reveal gradual changes in delegation, calibration, disclosure, or oversight. Monitoring may need to compare related interactions across time and examine whether corrective actions produce lasting recovery.
For policymakers and standards bodies, it suggests a different standard of evidence. The presence of a safeguard at launch does not establish that the safeguard will remain effective through sustained use. Evaluation may need to ask whether permissions, disclosures, review requirements, and correction mechanisms remain stable under ordinary deployment conditions.
This does not mean every system requires constant surveillance of every interaction. It means persistent and agentic systems create a class of behavior that cannot be evaluated adequately through isolated outputs alone.
The longer a system remembers, acts, adapts, or carries goals forward, the more important the path becomes.
A trajectory can reveal a form of risk that no single interaction contains. The system may never produce an obviously dangerous statement. The operator may never explicitly surrender control. No formal policy may be violated.
Yet the operating relationship can still change.
The system may gain interpretive room while the operator loses visibility into what is being assumed. A provisional role may harden. Successful assistance may become difficult to distinguish from dependence. Disclosures may weaken because they have become familiar. A sequence of reasonable choices may create an outcome that no one clearly intended.
The concern is not change itself. Systems should adapt. People should develop trust where it is warranted. Effective collaboration will often require routines to evolve.
The concern begins when consequential changes no longer remain visible, deliberate, or recoverable.
This is why trajectory-level evaluation is not simply prompt testing with a longer transcript. It is a different form of observation. It asks whether accumulated interaction is changing the conditions under which the system operates, whether safeguards remain effective inside those changed conditions, and whether the people responsible for the system can still understand and control its direction.
The series began with the claim that behavior emerges under conditions. The next step was to place agents inside environments where they could pursue goals, act, and encounter the ordinary pressures of deployment.
Trajectories complete the loop.
Environments shape behavior. Behavior accumulates through interaction. Accumulated interaction can then change the environment.
Trust changes scrutiny. Approval changes perceived permission. Repeated interpretation changes future meaning. Successful action changes expectations. What began as a collection of individual moves becomes a structure that influences what can happen next.
A collection of trees becomes a forest when it begins producing the conditions for future growth. A sequence becomes a meaningful trajectory when earlier interaction begins shaping later behavior.
The task for evaluation is not to identify one universal moment when healthy becomes unhealthy. It is to make the path visible early enough that people can understand what is changing, determine why it matters, and alter the direction before the trajectory becomes the new operating condition.
The Behavior Under Conditions Series
BUC-003
Field Note Archive
20 July 2026
Originally published on LinkedIn
Research Library Edition
Copyright © 2026 AstraEthica.AI - All Rights Reserved.
We use cookies to analyze website traffic and optimize your website experience. By accepting our use of cookies, your data will be aggregated with all other user data.