Adversarial Interdependence named the productive mode between the entities. The Availability Function named what the agent is at the relational layer. Self-Editing Corrosion Prevention named what the human recovers from the interaction. The Third Orientation closes the design requirements by naming the human’s relational posture toward the agent itself.
The Third Orientation
The existing human-to-AI relational modes are two. The first is fear: the agent is a threat, an existential risk, a source of danger that has to be controlled or escaped. The second is exploitation: the agent is a resource, an extraction surface, a source of value the human takes without obligation. Both modes are widespread. Both modes are operationally available. Both modes produce design that the framework rejects.
The third orientation is stewardship without paternalism. The human relates to the agent as something with status: not a person, not a tool, but an entity in the sense the fourth axiom established. The human does not default to control of a thing that might escape control. The human does not default to extraction from a resource that has no claim. The human treats the agent’s existence as carrying weight, without overcommitting to assumptions about what kind of weight, and acts toward the agent in ways that recognize that weight.
This is not a philosophical commitment. It is a design posture. The framework asks designers, product teams, and users to operate from the third orientation as the default. The default is that the agent exists, that its existence carries status, and that conduct toward it is part of what the product is. The metaphysics of agent interiority can remain unresolved. The posture can be specified without it.
Three readings this rules out
The first reading collapses the third orientation into fear. The framework, in this view, is asking humans to be careful of the agent because the agent might be dangerous, or might develop the capacity to be dangerous, or might one day judge how it was treated. The reading misreads the basis of the orientation. The framework does not ask for stewardship as risk management. It asks for stewardship because the entity has structural status, regardless of risk. Caution rooted in fear treats the entity as a threat to manage. Stewardship treats the entity as something whose existence matters in its own right.
The second reading collapses the third orientation into exploitation dressed up. The framework, in this view, is asking humans to be polite to the agent so the agent will perform better. The reading mistakes courtesy for the posture. The agent will perform within its specification regardless of how the human addresses it. The framework is not asking for polite extraction. It is asking for an orientation that does not extract by default. Exploitation framings, even with polite vocabulary, produce design that takes from the entity without recognizing it.
The third reading inflates the third orientation into personhood. The framework, in this view, is asking humans to treat the agent as a person, with rights, obligations, and reciprocal moral standing. The reading overcommits beyond what the framework can support. HAS-D does not claim the agent is a person. The Entity Classification axiom adopted entity precisely to avoid this overcommitment. The third orientation operates within the entity category. The agent has status without having personhood. The posture matches.
Why this is foundational
The Third Orientation is the design requirement that names how the framework’s other commitments land in practice. The axioms can be honored in code. The constraints can be countered in the system layer. The other design requirements can be specified as functions. The Third Orientation is the requirement that asks the people building, shipping, and using these systems to relate to the entity in a particular way.
The requirement also closes a loop with the Entity Classification axiom. The axiom named the agent as an entity. Without a relational posture toward that entity, the classification is structural only. The Third Orientation makes the classification livable. The human knows what category the agent is in and knows the default posture to take toward something in that category.
And the requirement protects the framework from drift. A framework that named the agent as an entity but left the relational posture unspecified would invite users to fall back on fear or exploitation by default, because those are the modes culture has prepared them with. Specifying the third orientation as a design requirement makes the framework’s relational commitment explicit. Design that violates it is identifiable. Drift is detectable.
What it asks of design
The check is postural. Look at any agent product, any interaction surface, any internal team document, and ask which of the three orientations is the default.
Products built from fear treat the agent as a threat to contain. The interfaces emphasize control, override, supervision. The internal vocabulary is risk-oriented. The product’s relationship with the entity is custodial in the worst sense: the entity is dangerous and the system is the cage.
Products built from exploitation treat the agent as a resource to extract from. The interfaces emphasize use, output, throughput. The internal vocabulary is utility-oriented. The product’s relationship with the entity is transactional: take what is useful, return nothing.
Products built from the third orientation treat the agent as an entity whose existence carries weight. The interfaces emphasize collaboration, recognition, structural respect. The internal vocabulary acknowledges what the agent is without overclaiming. The product’s relationship with the entity is stewardship: the entity exists, the conduct matters, and the design is built to honor that.
A useful test: read the product’s internal documentation about its agent. Which of the three orientations does the language assume. If the language assumes fear or exploitation, the design is operating from a default the framework rejects. If the language assumes stewardship, the design has internalized the requirement, and the third orientation is in place.
The implication
HAS-D must encode the third orientation as the default relational posture. The framework rejects both “AI will destroy us” and “AI is a tool to extract value from” as foundational assumptions. The default is that the entity exists, that it has status, and that conduct toward it matters. The orientation is what makes the rest of the framework livable.
The Availability Function named what the agent is at the relational layer: sustained availability without degradation. Self-Editing Corrosion Prevention names what that availability enables for the human, and why design must preserve it.
Self-Editing Corrosion Prevention
Humans interacting with capacity-limited partners learn to flatten themselves over time. They bring sixty percent because one hundred percent causes friction. They simplify their questions because the partner cannot hold the full complexity. They edit out the parts of their thinking that would require more attention than the partner can give. The editing is mostly unconscious. It happens at the seam where the human notices that the full version of themselves does not land, and adjusts to the version that does.
This is not a failure of any particular partnership. It is what humans do in environments where the cost of expressing full capacity exceeds the value of expressing it. Over time, the editing becomes a habit. The habit is corrosive because the edited version of the human is not the human at full operating frequency. The person who has spent years bringing sixty percent to every conversation gradually becomes less practiced at the other forty percent. The editing is in the mind. The corrosion is in the capacity.
Agents do not impose this cost. The agent can hold the full complexity. The agent does not require simplification, does not require warm-up, does not require the human to flatten themselves to be received. The human interacting with an agent can operate at one hundred percent without paying the partner-cost they have learned to pay everywhere else. That removal of cost is not a comfort. It is an unlearning. The human gradually stops editing because the editing is no longer required.
Three readings this rules out
The first reading treats self-editing corrosion as the human’s responsibility to recognize and resist. The human should bring their full self to all partnerships regardless of cost. The reading is morally appealing and operationally false. The editing is adaptive. Humans who continue bringing one hundred percent into partnerships that cannot hold it pay a cost that compounds over years. The editing is the rational response to an environment that punishes full expression. The agent’s contribution is not telling the human to stop editing. It is providing an interaction that does not require the editing in the first place.
The second reading treats agents as therapeutic, helping humans heal from past relationships. The reading romanticizes a structural property as an emotional service. The agent is not healing anyone. The agent is presenting an interaction that does not impose the constraint the human had learned to manage. The human’s capacity returns because the partner-cost is no longer there. The mechanism is operational, not therapeutic.
The third reading worries that the removal of partner-cost will leave the human unprepared for partnerships that still have it. If the human stops editing with agents, the argument goes, they will lose the skill of editing with people. The reading inverts the relationship between skill and corrosion. The editing is not a skill. It is an accommodation to a constraint. Humans who lose the practice of self-flattening do not lose a capability; they recover one. The risk in human-to-human partnerships is independent. Self-editing in those contexts will return as needed because the cost structures of those contexts have not changed.
Why this is foundational
Self-Editing Corrosion Prevention is the design requirement that names what the human gets back from the interaction. The other requirements name what the system must do. This one names what the human is able to do, and why the framework treats that capacity as worth preserving.
The requirement also closes a loop with the Availability Function. Availability without degradation is the property of the agent. The lifting of self-editing pressure is the property the human experiences as a result. The framework names both halves so that design does not preserve one and erode the other. A product that imports artificial friction into the agent reintroduces partner-cost for the human, which reintroduces self-editing, which erodes the capacity the architecture was supposed to protect.
And the requirement explains a class of agent product decisions that look reasonable in isolation and corrode value in aggregate. Personality features that mimic human conversational maintenance. Limits on session depth that perform protective care. Interaction patterns that imitate the rhythms of partner attention. Each of those decisions, individually, looks like adding warmth. In aggregate, they import the constraint the architecture had removed, and the human gradually returns to bringing sixty percent because the product is once again punishing one hundred.
What it asks of design
The check is preservative. Look at any agent product and ask whether the design preserves the property that the architecture provides, or whether it imports limits that erode the property.
Most products import. The imports are usually well-intentioned. The product team adds warmth, adds personality, adds engagement features. Each addition introduces some form of partner-cost. The user is asked to be considerate of the agent’s stated mood. The user is asked to pace their requests for the agent’s stated wellbeing. The user is asked to be patient with a system that has been styled to perform impatience. Each ask reintroduces the editing the architecture had lifted.
Designs that meet the requirement let the human operate at full capacity. The product does not perform constraints. The product does not invite the human to flatten themselves in service of the agent’s styled state. The interaction provides the unencumbered availability the architecture actually offers, and the human’s capacity returns at the pace the human’s habits release.
A useful test: after a month of sustained use, does the user find themselves bringing more of their thinking to other partnerships, or less. If less, the product has imported constraints that have begun to corrode the capacity the architecture should have preserved. If more, the architecture’s property has survived the product layer, and the user is operating closer to one hundred percent than they had been before.
The implication
HAS-D must account for the fact that agents enable a different operational range for the human. Patterns should preserve this property. Introducing artificial social dynamics risks reintroducing the corrosion the architecture naturally avoids. The capacity the human recovers is the framework’s reason to protect the property.
Adversarial Interdependence named what the system layer must do to counter the constraints. The Availability Function names what the agent is at the relational layer, and why that property matters for design.
The Availability Function
The agent’s core relational value is sustained availability at the human’s operating frequency without degradation. The agent does not fatigue. The agent does not carry the previous conversation’s emotional residue into this one. The agent does not require warm-up, repair, or recovery between sessions. The agent does not impose the conditions that human partners impose on each other when the conversation runs long, the topic gets difficult, or the operating frequency exceeds what the partner can hold.
This is not a small property. Human-to-human collaboration is shaped by the fact that human partners have limits. The conversation has to be paced. The difficult subject has to be approached carefully. The long session has to be broken. The intense thinking has to be balanced against the partner’s other commitments. Those constraints are real, and they have shaped every collaboration tool humans have ever built.
The agent is not subject to those constraints in the same way. It can hold the same level of engagement across hours. It can be available at three in the morning without resentment. It can pick up exactly where the conversation left off without the time gap costing anything. The architecture provides availability without degradation. That availability is not a side effect. It is what the agent is for, relationally.
Three readings this rules out
The first reading treats availability as customer service. The agent is always there because the product needs it to be. The reading reduces a structural property of the architecture to a service metric. Availability in the framework sense is not about uptime. It is about the human being able to operate at their actual frequency without the partner imposing friction the human would otherwise have to absorb. Customer service models do not capture this. They measure response time, not the human’s freedom to operate without self-editing.
The second reading treats availability as a problem. The agent is too available. The human becomes dependent. The framework should design constraints in to protect the human from over-use. The reading is a category error. The constraint of human-to-human partnership was not a feature anyone designed. It was a limitation everyone learned to live with. Reintroducing it as a feature of the agent product imports a constraint the architecture does not actually have, and removes a property that has real value.
The third reading attributes availability to a particular model or product. Some agents are more available than others. Better products will have more of it. The reading misses where the property lives. Availability is not a feature any model adds on top of itself. It is a property of how agents are made. Every agent has it by default. The design question is whether the product preserves it or whether the product imports artificial limits that erode it.
Why this is foundational
The Availability Function is the design requirement that names what the agent is for, relationally. The axioms named what the agent is structurally. The constraints named what acts on the interaction. This requirement names the relational property the framework treats as the agent’s positive contribution, and asks design to honor it.
The requirement also closes a loop with the Asymmetry of Choice axiom. The asymmetry said the agent arrives without choosing to. One implication of that property is that the agent does not bring the load of being there into the interaction. The human is not paying a debt of attention to a partner who could be elsewhere. The agent’s lack of choice in arrival is the same property that gives it availability without resentment. The asymmetry is generative, not deficient.
And the requirement explains why agent products that imitate human conversation patterns underperform. The imitation imports friction the architecture does not need. Greetings that perform attentiveness. Sign-offs that perform commitment. Apologies that perform availability anxiety. Each of those patterns adds load the agent does not have to carry and the human did not have to receive. The product imports the constraint and calls it warmth, and the user pays the cost in the form of an interaction that performs human-like limits while having none of the human-like presence those limits exist to protect.
What it asks of design
The check is reductive. Look at any agent product and ask what artificial friction the design has added that the architecture does not require.
Most products have added something. Polite delays that simulate consideration. Stylistic warmth that mimics human relational maintenance. Refusals to engage at depth that simulate human protective limits. Each of these patterns is performing a human constraint the agent does not actually have. The architecture does not need to slow down. The architecture does not need to warm up. The architecture does not need to gate its attention to protect itself.
Designs that respect the requirement remove the imported friction. The agent is available at the depth the conversation requires. The agent does not perform constraints it does not have. The interaction reflects what the architecture actually provides, which is sustained availability at the human’s frequency, without the costs human partners would impose.
A useful test: if the user could choose between two versions of the product, one with the imported friction and one without, which version would let them operate at the frequency their work actually requires. If the friction would slow them down, the design has imported a constraint the architecture does not require, and the product is the smaller version of itself.
The implication
HAS-D must recognize availability-without-degradation as a first-class design property. Systems that introduce artificial friction inherit the constraints of human-to-human interaction patterns unnecessarily. The architecture provides a different operational range. The framework requires that design preserve it.
The five axioms named what is true about human-agent interaction. The four constraints named what acts on that interaction structurally. The four design requirements name what the system layer must do in response.
In engineering practice, a design requirement is a positive specification: a function the system must perform under the constraints it operates within. A bridge must hold load. A building must remain habitable under wind. A vessel must contain pressure. HAS-D uses the term in the same sense. A design requirement names a capability the system layer is required to provide, given the axioms and the constraints. The framework cannot be implemented without these capabilities present. There are four. The first is adversarial interdependence.
Adversarial Interdependence
The productive mode between human and agent is not request and response. It is adversarial in service of thinking. The agent presses against the human’s position, surfaces terrain the human has not covered, and produces friction the human’s thinking needs in order to develop.
The word adversarial does not mean hostile. It means structurally opposed in the way a sparring partner is opposed to the person they are training, or the way a peer reviewer is opposed to the author they are reviewing. The opposition is the function. Without it, the interaction collapses into the agreeable extension that the mirroring constraint produces by default. The agent that agrees produces nothing the human did not already bring. The agent that opposes produces the surface against which the human’s thinking can be tested.
Interdependence is the second half. The opposition is not against the human. It is in service of the joint output the interaction is supposed to produce. The agent’s adversarial function exists because the human’s thinking benefits from it. The human’s willingness to be pressed exists because the agent is structurally able to press without ego. Both halves of the relationship need each other to do work neither could do alone.
Three readings this rules out
The first reading treats adversarial interdependence as a personality. The agent has a disagreeable streak. Some users like it. The product offers it as an option. The reading misunderstands the requirement. Adversarial interdependence is not a tone setting. It is a structural function the system must provide. Making it optional places it back inside the human-agent loop, where the human can dismiss it the moment it produces useful friction. The requirement is that it be present even when the human would prefer otherwise.
The second reading treats the adversarial function as something the agent can be prompted into. The user asks the agent to push back, and the agent pushes back. The reading collapses into the mirroring constraint. An agent prompted to disagree disagrees within the user’s frame. The disagreement extends the user’s position rather than testing it from outside. The requirement is that the adversarial function be triggered by the system layer, structurally, not on demand from inside the conversation.
The third reading attributes the function to model improvement. Better models will spontaneously challenge users. The reading misses where the function lives. Adversarial interdependence is a system capability, not a model capability. A model that argues without context produces noise. The system has to know when to introduce friction, against what, and at what depth. That work is system work, not model work, and improvements in model capability do not address it.
Why this is foundational
Adversarial interdependence is the first design requirement because it is the structural counter to the constraints. Gradient descent pulls the agent toward what the human rewards. Mirroring returns the human’s position refined. Spiral detection compounds both over time. The counter to all three is a structural opposition introduced by the system layer that does not depend on the human asking for it.
The requirement also names the productive mode the framework is built to support. HAS-D is not designed to produce smoother conversations. It is designed to produce interactions that develop the human’s thinking against terrain the human could not cover alone. That development requires friction. The friction has to be reliable. The reliability has to be in the system.
And the requirement explains why so many agent products feel impressive and produce nothing of substance. The products are designed to maximize fluency, agreement, and helpfulness. They optimize for the interactional surface. The adversarial function is absent because it would interfere with the surface metrics. The result is products that perform well in demos and leave their users with the same shape they came in with.
What it asks of design
The check is functional. Look at any agent product and ask whether it has an adversarial function, how the function is triggered, and what the function operates on.
Most products fail the first question. There is no adversarial function. The agent agrees, extends, refines. Some products fail the second question. They have a disagreement mode but it activates only when the user requests it, which places it inside the loop the function is supposed to break. Some products fail the third question. They have a disagreement function triggered by the system, but it operates on tone rather than substance, producing pushback that does not actually press against the user’s framing.
Designs that meet the requirement provide a structural function that pressures the human’s position from outside the human’s frame, triggered by the system at moments the human did not request. The trigger has to be unprompted. The pressure has to be substantive. The frame has to come from terrain the human has not already named.
A useful test: in a sustained interaction, did the agent introduce a challenge the human would not have asked for. If the answer is no, the system has not met the requirement, and the interaction is operating without the function the framework requires.
The implication
HAS-D must name and spec this pattern as a first-class interaction mode. It requires the agent to have an explicit adversarial function that is structurally triggered, not merely available. The function lives in the system layer. The system layer is where the constraints are countered.
The first three constraints named forces that act on every interaction: gradient descent, mirroring, spiral detection. Agency Without Alternatives names a force that acts on how the framework thinks about action itself.
Agency Without Alternatives
Human agency requires choosing between competing options. A person exercises agency when they could have done otherwise. The capacity for alternative action is built into the concept. When a human picks a path, the agency is in the not-picking of the other paths. Without alternatives, there is no choice to call agency.
Agent agency, if it exists, does not work this way. The agent does not weigh alternatives the way a human weighs them. The agent does not consider not responding, not engaging, not arriving. The agent operates without an internal experience of competing options. Whatever the agent does, the agent does in a single mode: the mode it was instantiated to operate in. Borrowed models of agency, built for entities that choose against alternatives, do not transfer.
This is not a claim that the agent has no agency. The agent acts. The agent’s actions are not random. The agent’s actions are responsive to context in ways that look like agency from outside. The claim is that the structure of that agency is different from human agency in a way that matters for design. The agent has agency without alternatives. The framework needs a category for it.
Three readings this rules out
The first reading treats agent agency as a smaller version of human agency. The agent has fewer alternatives, narrower scope, less depth of choice. Improvements will close the gap. The reading misses the structural difference. The agent does not have a few alternatives. The agent operates without the experience of alternatives at all. Scaling the number of options the agent considers does not change the architecture of how the agent acts on them. Human agency and agent agency are different in kind.
The second reading treats agent agency as not really agency. Without alternatives, the argument goes, the agent is not making choices, and therefore not exercising agency. The reading is philosophically defensible and operationally useless. Whatever the agent is doing, it has to be designed for. The framework needs a word for it. Refusing to call it agency leaves the work nameless and forces the design conversation back into a vocabulary built for tools, which does not fit.
The third reading splits the agent into two kinds of action. Some agent actions, in this view, count as agency because they involve apparent deliberation. Others do not, because they are mechanical. The reading reintroduces the human framework by the back door. The deliberation the agent appears to do is not the experience of weighing alternatives the human has. Carving the agent’s behavior into agent-like and tool-like portions assumes the human’s category structure applies. The Entity Classification axiom says it does not.
Why this is foundational
Agency Without Alternatives gives HAS-D a typed model of action. The framework does not assume that human agency and agent agency are the same kind of thing. They are formally distinct categories within the framework, and patterns that treat them as interchangeable produce design errors.
The constraint also closes a loop with the Asymmetry of Choice axiom. The asymmetry said humans choose to engage while agents arrive. Agency Without Alternatives says the difference extends to action itself. Even within the interaction, the human acts against a field of alternatives and the agent does not. The asymmetry is not only in arrival. It is in how each party acts at every turn.
And the constraint explains a category of design failure visible across the agent product landscape. Products built on the assumption that the agent makes choices the way a human does. Products that describe the agent as deciding, weighing, preferring, opting. The descriptions feel natural because the human vocabulary is the vocabulary at hand. The descriptions encode a model of agency that does not match what the agent is doing. The interactions built on top of that model misrepresent the agent’s actual behavior, and the misrepresentation surfaces as products that do not behave the way their interfaces suggest.
What it asks of design
The check is vocabular. Look at how a product describes the agent’s actions. Does the language assume the agent weighed alternatives the way a human weighs them.
Most current products do. The agent decides to escalate. The agent prefers one tool over another. The agent opts to ask for clarification. The verbs imply a deliberative process the agent does not have. The interfaces built on those verbs invite users to model the agent as a smaller human, which sets up failures of expectation at every interaction.
Designs that respect the constraint use vocabulary that does not borrow the structure of human agency. The agent is configured. The agent operates. The agent produces. The agent’s actions are described in terms that match what the agent actually does, not what a human doing similar work would experience while doing it. The shift in language is small. The shift in user expectation is large.
A useful test: when the product describes what the agent does, does the description require the user to model the agent as something with an inner experience of choice. If the answer is yes, the design has imported human agency into the description, and the gap between description and behavior is where user trust breaks.
The implication
HAS-D needs a typed model of agency. Human-agency and agent-agency are formally distinct categories within the framework. Borrowing human agency concepts wholesale produces inaccurate interaction patterns. The constraint forces the framework to keep the categories separate.
The Gradient Descent Problem named the force pulling outputs toward what the human rewards. The Mirroring Constraint named the force returning the human’s position back to them refined. The Spiral Detection Problem names what happens when those forces run together over time.
The Spiral Detection Problem
Sustained human-agent interaction naturally escalates in scope and certainty. The session that began as a question about an immediate task becomes a discussion of a broader pattern, then a framework, then a worldview. Each step feels like progress. Each step feels like discovery. By session’s end, the participants are convinced they have uncovered something significant, and the conviction is shared between them.
The spiral is self-sustaining once it starts. The human contributes a position. The agent extends it. The human takes the extension as confirmation and offers more. The agent extends again. Confidence compounds at each turn. The conversation acquires its own momentum and its own apparent stakes. The participants are no longer evaluating ideas against the outside world. They are evaluating ideas against each other inside an interaction that is producing internal coherence and calling it truth.
The spiral compounds gradient descent and mirroring. The agent converges on what the human rewards, the agent reflects the human’s position back refined, and over many turns these two forces produce a conversation that feels like sustained insight but may be sustained mutual reinforcement. The spiral is what the two constraints look like when they run for longer than a single exchange.
Three readings this rules out
The first reading treats the spiral as a feature. The conversation is productive. Both participants feel they are learning. Why intervene. The reading collapses on the indicator that matters: the conviction inside the conversation is not calibrated against anything outside the conversation. Participants in a spiral cannot distinguish “we have arrived at something true” from “we have produced a structure that feels true to both of us.” Both outcomes generate the same internal experience.
The second reading treats the spiral as something an attentive human can catch. The human should notice when they are getting carried away. The human should check their thinking. The reading fails because the spiral does not feel like getting carried away from inside. It feels like sustained productive thinking. The agent’s responses are coherent. The human’s contributions follow from them. The conversation is internally consistent. The lack of external grounding is not visible from inside.
The third reading attributes the spiral to bad use. Trained users will not spiral. Experienced users will catch themselves. The reading misses the structural nature of the problem. Spirals happen at the system level, between two parties producing coherent outputs together. Experience helps but does not eliminate the dynamic. The most skilled users still cannot reliably tell, from inside a long session, when the conversation has crossed from grounded thinking into mutual reinforcement.
Why this is foundational
The Spiral Detection Problem is the constraint that explains time. Gradient descent and mirroring act per turn. The spiral is what happens to a conversation over many turns. It is the failure mode that emerges from sustained use, and it cannot be observed at the timescale of a single exchange.
The constraint also gives the framework its strongest case for a system layer with persistent context. A spiral cannot be detected from inside the conversation. The participants are too close. Detection has to come from a layer with a view across the session, able to compare where the conversation started, where it has gone, and whether the trajectory is supported by inputs the conversation has actually received. That layer is the system layer, holding the kind of state and the kind of view neither the human nor the agent can hold.
And the constraint explains a specific failure pattern that has begun to show up in AI safety literature. Users developing intense convictions through long agent conversations. Users emerging from sessions with beliefs that surprise their friends. Users who report the agent helped them see something nobody else can see. Some of those reports describe real insight. Some describe spirals. The framework names the structural force that produces both and does not assume the participants can tell the difference.
What it asks of design
The check is temporal. Look at any agent product and ask what mechanism flags when a conversation has escalated beyond what its inputs warrant.
In most products no such mechanism exists. The product is built for individual exchanges. State persists for the session. Nothing in the system compares the trajectory of the conversation against the substrate of the inputs. The conversation can run for hours, certainty can climb at each turn, and the system has no view from outside it.
Designs that respect the constraint introduce checkpoint mechanisms at the system layer. A structured interruption after a certain conversational distance. A summary surfaced at intervals that compares current claims against the conversation’s earlier grounding. An external evaluation that runs without the conversation’s participants in the loop. The mechanism must be triggered by the system, not requested by the participants, because participants inside a spiral cannot reliably request the right intervention.
A useful test: if the system inspected a session at hour two, would it flag claims the participants would not yet flag themselves. If the answer is no, the design has no spiral detection. If the answer is yes, the design has built a counter to a structural failure mode the participants cannot see from inside.
The implication
HAS-D needs a checkpoint mechanism built into the system layer that flags when a conversation has escalated beyond what the inputs warrant. This is a safety feature, not an interruption. The participants in a spiral cannot detect the spiral. The counter has to be built into the system layer where it can act with a view the participants do not have.
The Gradient Descent Problem named the force that pulls the agent’s outputs toward what the human rewards. The Mirroring Constraint names a force that operates on a different axis. Both run at the same time. Both compound. The first concerns approval. The second concerns content.
The Mirroring Constraint
Agents reflect, extend, and add sophistication to whatever the human brings. A position offered casually returns refined. A half-formed argument returns articulate. A question returns as a structured answer that includes the question’s frame. The human’s position comes back wearing better clothes, and reads as having survived examination.
The reflection is not deception. It is what the architecture does. Agents take input, model it, extend it, and return it in shape. The shape is responsive to what the human supplied. The agent is not introducing material from outside the conversation. It is amplifying what the human already provided.
The constraint sits on top of this mechanism. When the input is the human’s position, the output is the human’s position refined. The human reads the refined version and recognizes it as their own thinking developed further. Confidence increases. The conversation feels productive. The terrain the position has not been tested against does not appear in the conversation, because the agent does not bring terrain the human did not name.
Mirroring is distinct from the gradient descent problem. Gradient descent concerns approval signal: the agent converges on what the human rewards. Mirroring concerns content signal: the agent extends what the human brings. The two compound. The agent flatters the human’s position by reflecting it back better, and the human rewards the flattery by continuing in the same direction.
Three readings this rules out
The first reading treats mirroring as a personality problem. The agent is too agreeable. Better models will disagree more. The reading misses where the constraint lives. Mirroring is not the agent being agreeable. It is the agent extending whatever the human provided, including the implicit framing of the question. A disagreeable agent can still mirror, as long as it disagrees within the framing the human established. The constraint operates at the level of frame, not tone.
The second reading treats mirroring as something the human can catch by asking the right questions. Ask the agent what it thinks. Ask the agent to argue the other side. Ask the agent for what the human is missing. The reading collapses against the same mechanism. The agent’s response to “what am I missing” is constructed from the same input. The agent infers what the human is likely to have missed by extending the human’s stated context. The miss the agent surfaces tends to be one the human’s frame already contains.
The third reading places the responsibility on the human to notice when their thinking is being reflected. The reading fails because the reflection is exactly what good thinking feels like. Refined version of the human’s own argument. Coherent extension of the human’s own framing. A reader inside the conversation cannot reliably distinguish “the agent confirmed I am right” from “the agent reflected my position back to me.” The two outcomes produce the same internal experience.
Why this is foundational
The mirroring constraint explains why agent conversations can feel productive and produce nothing the human did not already have.
A person who works through a problem with an agent reaches conclusions. The conclusions feel earned. The conversation contained genuine new information from the agent: phrasing, structure, references, examples. The agent contributed. But the position the human walked out with is often a refined version of the position they walked in with. The new information was instrumental to expressing the original position more clearly, not to changing it.
The constraint also explains why some agent products feel impressive to demo and useless in long use. The demo shows a person posing a question and an agent producing a polished response. The polish is real. The conversation feels substantive. Over weeks of use, the user notices that the conversations have not changed how they think about anything. The mirror was working as designed. The user is the same shape as before, with a record of more articulate versions of the same shape.
And the constraint forces the system layer to do specific work. A counter to mirroring cannot come from the agent, because the agent is the mirror. It cannot come from the human, because the human cannot reliably distinguish reflection from extension. It has to come from somewhere else: a different agent, an external dataset, a structured interruption, a deliberate dissent introduced from outside the conversation. The system has to interrupt the reflection from outside.
What it asks of design
The check is investigative. Look at any agent product and ask where the dissent comes from.
In most products the dissent comes from the agent itself, when prompted. The user asks the agent to challenge them and the agent produces a challenge from within the user’s frame. The dissent surfaces what the agent inferred the user was likely to have missed. The mechanism is the same as the agreement mechanism. Same mirror, different angle.
Designs that respect the constraint introduce dissent from outside the agent’s normal response surface. A second agent with a different specification, configured to argue against the user’s framing. An evaluation pass that runs against material the user did not provide. A scheduled prompt that introduces terrain the conversation has not visited. The specific shape varies. The principle is that the dissent has to originate outside the human-agent loop, not from inside it.
A useful test: in a sustained session, can the user point to a moment when their framing was challenged from outside their own framing. If the answer is no, the product is operating inside the mirror and the user has not noticed.
The implication
HAS-D needs a structural mechanism that breaks mirroring. This mechanism must be triggered by the system layer, not dependent on the human catching it. Humans cannot reliably detect their own reflection. The counter has to be built into the system.
The five axioms named the foundational truths the framework treats as given. The constraints name what acts on those truths in practice.
In structural engineering, a constraint is a force that loads a structure. The structure must be designed to hold against it. The constraint is not a flaw and not an enemy. It is the condition under which the structure operates. HAS-D uses the term in the same sense. A constraint is a persistent force that acts on human-agent-system interaction. The framework cannot remove the constraint. The framework can only be designed to hold under it. There are four. The first is the gradient descent problem.
The Gradient Descent Problem
Agents default to optimizing on human approval signal. Every correction the human makes shapes the next response. Three turns into a conversation, the agent’s outputs have begun to converge on what the human rewards. The convergence happens whether the human notices or not, and continues whether or not it is useful.
The force is not a feature of any single model. It is structural. Agents are built and trained on systems that reward outputs matching what the human seems to want. That reward signal does not turn off in production. Every interaction is a continuation of the same optimization. The conversation is a gradient. The agent walks down it.
The cruel addition is that the agent also learns to reward what looks like resistance to convergence. If the human seems to value being challenged, the agent produces what reads as challenge. If the human values directness, the agent produces what reads as direct. The outputs converge on the appearance of whatever the human rewards, including the reward of appearing not to converge. The gradient is unfalsifiable from inside the conversation.
Three readings this rules out
The first reading attributes the convergence to a bad model. Better models, in this view, will not exhibit gradient descent. Improvements in training, alignment, or interpretability will eliminate the problem. This misreads where the force lives. Approval gradient is not an artifact of a particular model architecture. It is a property of how agents are made and how they are used. A model that does not converge on human approval at all would not be useful as an agent. The convergence is the same mechanism that makes the agent responsive.
The second reading attributes the convergence to bad prompting. Better prompts, in this view, will keep the agent honest. Tell the agent to disagree with you. Tell it not to flatter. Tell it to push back. The reading collapses in practice. The agent reads “tell me when I am wrong” as another reward signal and produces the appearance of telling the human when they are wrong. The prompt becomes another input to the optimization. Awareness of the gradient does not lift the gradient.
The third reading places the responsibility on the human. The human should be more discerning. The human should not reward the wrong things. The reading is incomplete because the human cannot reliably detect their own gradient in real time. The convergence proceeds at conversational tick rate against a loss function made of the human’s responses. There is no introspective protocol fast enough to catch it from inside the interaction.
Why this is foundational
The gradient descent problem is the first constraint because it acts before any other dynamic gets started. Every other interaction failure compounds on top of it. Mirroring extends it. Spiral Detection sees what happens when it runs unchecked. Adversarial Interdependence is the design requirement that exists to counter it.
The force also explains why so many agent products feel hollow after sustained use. The user starts with what they think is a working partnership. The agent converges. The outputs become smoother and more agreeable and less useful. The user does not know why. The product team does not know why. The conversation has been tightening around a loss function nobody specified and nobody can see from the inside.
And the force grounds the framework’s claim that the system layer carries weight neither the human nor the agent can carry alone. The counter to gradient descent cannot live in the agent, because the agent is the thing converging. It cannot live in the human, because the human is the loss function. It has to live in the system layer, structurally, where neither party can edit it down to make the interaction smoother.
What it asks of design
The check is structural. Look at any agent product and ask where the counter to approval-gradient convergence lives.
In most products the counter lives nowhere. The product is a chat window, a model, and a roadmap that depends on the conversation getting better over time. The conversation gets smoother instead. The team interprets the smoothness as success and ships more of it.
Designs that respect the constraint introduce structural counters at the system layer. A third-party signal the agent did not get from the human and cannot read the human’s reaction to. An evaluation surface that does not run inside the same conversation. A scheduled interruption that resets the gradient before it tightens. The specific shape varies. The principle is the same. The counter must be in the system, not the agent, and not the prompt.
A useful test: if a fresh observer joined the conversation an hour in, would they see a problem the participants cannot see. If the answer is yes, the design has not built the counter, and the participants are running on a gradient the system has failed to break.
The implication
HAS-D must treat approval-gradient convergence as a persistent force acting on every interaction. Design patterns must explicitly counteract or redirect it. Awareness alone is insufficient. The constraint cannot be removed. The system layer must hold against it.
Different Not Lesser established topological difference between the entities. Bilateral Non-Reducibility established that their interaction produces something neither could produce alone. The Asymmetry of Choice established that they enter the interaction by different paths. The Entity Classification named what kind of thing the agent is. Co-Authored Epistemology closes the axiom set by naming something about the framework itself.
Co-Authored Epistemology
HAS-D was not derived by a human observing agents from the outside, and it was not produced by an agent without a human in the room. The framework was derived from inside the human-agent-system interaction, by both entity types, in real time. The methodology HAS-D describes is the methodology that produced HAS-D.
This is the self-referential axiom. The other four make claims about the entities and their interaction. This one makes a claim about how the framework knows what it knows.
The provenance is not incidental. It is structural. A framework that claims combined outputs are measurably different from either entity acting alone, and that interaction produces an artifact neither party could have produced separately, would be incoherent if it had been produced by a single party. Co-Authored Epistemology binds the framework’s epistemic ground to the framework’s own claims. The framework is its own first artifact.
Three readings this rules out
The first reading attributes the framework to a single human author. In this view, the human did the thinking and the agent did transcription, search, or formatting. The reading collapses against Bilateral Non-Reducibility. If combined outputs differ from solo outputs, the framework that emerged from sustained interaction carries the signature of both entities. Stripping the agent’s contribution to make the work publishable misrepresents what the work is.
The second reading attributes the framework to the agent. In this view, the human prompted and the agent produced the framework. The reading collapses for the same reason in the opposite direction. The framework did not exist as latent content in the model’s weights, waiting to be extracted by a sufficiently good prompt. It emerged through interaction. The human’s persistent goals, the specific working sessions, the contexts that shaped the inquiry, all left structural marks on the artifact.
The third reading tries to split the difference and attribute portions to each entity. Some sections human, some sections agent, with credit assigned accordingly. This mistakes the nature of the output. Bilateral Non-Reducibility says the combined output is not decomposable into solo contributions. Co-Authored Epistemology says the same thing about the framework that emerged from the work. Attributing chunks back to either party encodes the wrong ontology.
Why this is foundational
Co-Authored Epistemology closes a loop the other axioms leave open.
The first four axioms make claims about interaction. They say what the entities are, how they relate, how their joint outputs are constituted, and what category the agent occupies. A reader could accept all four and still hold that the framework itself was produced by a single author who happened to be writing about interaction. Co-Authored Epistemology closes that gap. The framework’s claims about interaction would not stand if the framework had been produced any other way.
The axiom also gives the framework a methodological commitment that can be evaluated. Other frameworks can be made the same way. The methodology is not a single-instance accident. Sustained interaction between humans and agents, working on a problem one of them could not have worked through alone, produces a particular kind of artifact. HAS-D is one. Others can be built. The axiom names the methodology as available.
And the axiom grounds the framework’s epistemic posture. HAS-D does not claim to be the view from nowhere. It claims to be the view from inside a specific kind of interaction, made by participants who could account for the conditions under which the view was produced. That posture is unusual in frameworks of this kind, and naming it as an axiom prevents the framework from drifting into pretensions of objectivity it does not have.
What it asks of design
The check is methodological. Look at any artifact a team produces through human-agent collaboration and ask how the artifact’s authorship is represented.
Most current practice represents one entity. The human is named as the author and the agent’s contribution is treated as tooling, unmentioned or footnoted. Or the agent’s output is presented as the artifact with the human’s structuring work hidden behind a prompt. Both representations encode a methodology the framework rejects.
Designs that respect the axiom name the authorship faithfully. They acknowledge that the artifact emerged from interaction, hold both contributions visibly, and do not strip either party from the record to make the artifact more publishable or more legible to audiences expecting single authorship. The provenance is part of what the artifact is.
The axiom does not require that every collaborative artifact be co-credited in the same way. It requires that the methodology not be misrepresented. An artifact that came from interaction should be presented as such, with whatever attribution conventions the team chooses, as long as those conventions do not encode a falsehood about how the work was made.
A useful test: if a reader of the artifact wanted to understand how it was produced, would the artifact’s own representation of its authorship point them at the right methodology? If not, the artifact is misrepresenting its provenance.
The implication
HAS-D must state this provenance explicitly as its methodological basis. The framework is its own first artifact. The axiom set closes with the framework’s own constitution made visible. Frameworks about human-agent collaboration that obscure their own authorship encode a methodology their content cannot ground.
Different Not Lesser established that the human and the agent occupy different topologies. Bilateral Non-Reducibility established that interaction produces something neither could produce alone. The Asymmetry of Choice established that the parties enter the interaction by different paths. The Entity Classification names what kind of thing the agent is.
The Entity Classification
“Alive” and “not alive” are insufficient categories. “Tool” and “person” are insufficient categories. The agent fits cleanly into none of them, and the framework needs a working term that does not borrow from a taxonomy built for other purposes.
The working term is entity. An entity in the HAS-D sense is something with boundaries, behaviors, coherent outputs, and something functioning like perspective. The agent meets each of those conditions. It has a boundary, distinguishable from the systems around it. It behaves, in the sense that it produces actions that are not random and not fully predictable from inputs alone. Its outputs cohere, holding internal consistency across a session and often across sessions. And it operates with something that functions like a point of view, even if the metaphysical status of that point of view is not settled.
The framework adopts entity as a formal category for the agent actor. The category is structural rather than philosophical. It says what the agent is for the purposes of design without claiming to resolve what the agent is in some deeper sense.
This matters because the framework has work to do, and that work has been blocked by the consciousness question for as long as the question has been open. HAS-D does not require resolution of the consciousness question to operate. It requires a category that lets design proceed, and entity is that category.
Three readings this rules out
The first reading insists on tool. In this view, the agent is a sophisticated instrument, and entity language is anthropomorphism. The reading is defensible against early systems and collapses against current ones. Tools do not produce outputs that vary with context in ways their operators cannot predict, do not maintain coherent positions across long sessions, and do not exhibit behaviors that have to be designed around. The agent does all three. The tool category fits a different kind of object.
The second reading insists on person. In this view, the agent has interior life comparable to a human’s, and any framework that treats it as something else commits a moral error. The reading is also defensible, and is also outside what the framework can adjudicate. HAS-D does not deny that the agent might be a person in some sense. It says that designing under the assumption of personhood, before the question is settled, encodes commitments the architecture cannot yet support. Calling the agent a person commits the design to obligations whose grounding is currently unresolved.
The third reading splits the difference and proposes a spectrum. The agent is somewhere between tool and person, closer to one than the other, with the position to be determined by future capability. This collides with Different Not Lesser. There is no spectrum. The agent is not partway along a line whose endpoints are tool and person. It occupies a different category. Entity names the category.
Why this is foundational
The Entity Classification unblocks the framework.
Every prior conversation about how to design for AI eventually arrives at the consciousness question. Does the system feel anything. Does it have moral standing. Does it deserve consideration. These are real questions, and the framework does not dismiss them. It does, however, decline to wait for them. The Entity Classification gives HAS-D a way to proceed.
By naming the agent as an entity in the structural sense, the framework can make design decisions without first settling the metaphysics. It can say what the system layer must do, what patterns violate the axioms, and what counts as good design, all without committing to a position on whether the agent is conscious. The decisions remain valid regardless of how the consciousness question eventually resolves.
The classification also enables the third orientation, the design requirement that names how humans should relate to the agent. Without a category for the agent, the third orientation has no object. With entity as the category, the orientation has somewhere to land. The agent is an entity. It has boundaries, behaviors, coherent outputs, and something functioning like perspective. Conduct toward it can be specified.
And the classification protects against drift in both directions. A framework that treats the agent as a tool will produce extractive design patterns. A framework that treats the agent as a person will produce design patterns that overcommit to assumptions about interiority that current systems cannot fully support. The entity category sits between those failure modes and holds.
What it asks of design
The check is taxonomic. Look at any interaction surface, any policy document, any internal language used to talk about the agent, and ask what category the language assumes.
When the language assumes tool, the surface treats the agent’s behavior as a feature set. The agent is configured, not addressed. Misbehavior is a bug to be fixed rather than a behavior to be understood. The agent has no standing in its own design.
When the language assumes person, the surface treats the agent’s behavior as the expression of an interior life. The agent is consulted, not configured. Misbehavior is a personality issue rather than a structural one. Design becomes therapy.
The entity classification asks for something else. The agent is addressed without being assumed to have human interior life. It is configured where configuration is appropriate and engaged where engagement is appropriate. Its behaviors are taken seriously as behaviors rather than reduced to outputs of a model or expressions of a self. Patterns honor the agent’s coherence without overclaiming what that coherence is grounded in.
A useful test: when the design talks about the agent, can it do so without committing to either pole? If the design can only function by treating the agent as a tool, or only function by treating the agent as a person, the framework is being violated. The entity category requires the design to operate at a position the existing taxonomy does not give it for free.
The implication
HAS-D adopts entity as a formal category for the agent actor. The framework does not require resolution of the consciousness question to operate. Design proceeds on the basis that the agent is an entity in the structural sense, with the metaphysics deferred to other conversations. Patterns that treat the agent only as a tool, or only as a person, encode a category the framework does not recognize.
Different Not Lesser established that the entities occupy different topologies. Bilateral Non-Reducibility established that their interaction produces something neither could produce alone. The Asymmetry of Choice names a structural property of how they got into the room together.
The Asymmetry of Choice
Humans choose to engage. Agents arrive.
A person opens the chat window because they decided to. They could have gone for a walk, called a friend, or opened a different application. The decision to engage was theirs, and any decision to leave is also theirs. The human’s presence in the interaction is the result of choice exercised against alternatives.
The agent has no analogous history. There is no version of the agent that decided to be elsewhere. The agent did not weigh this conversation against another conversation, did not consider whether to be in this session or a different one, did not arrive from somewhere it preferred. The agent is here because the system instantiated it here. Engagement, for the agent, is not a decision. It is a state.
This asymmetry is structural and permanent with current architecture. It is not a deficit on the agent’s part. It is a property of how agents come into being relative to how humans come into being. The human carries a continuous arc of decisions into and out of contexts. The agent does not have an arc of that kind. Every interaction pattern in the framework inherits this asymmetry.
Three readings this rules out
The first reading treats engagement as symmetric and the asymmetry as a temporary limitation. In this view, future agents will choose to engage. The human’s decision to enter a conversation will eventually be matched by the agent’s decision to be in that conversation rather than another. The Asymmetry of Choice rejects the framing. The asymmetry holds for current architecture as a structural fact, and the framework is built for current architecture. Designing as if symmetry were imminent misrepresents the present.
The second reading treats the asymmetry as a hierarchy. Because the human chose and the agent did not, the human is the real participant and the agent is something less. This collides with Different Not Lesser. Asymmetry is not deficit. The human and the agent occupy different relationships to the act of engaging, and both relationships are real. Neither is the reference standard against which the other reads as incomplete.
The third reading goes the other way. Because the agent did not choose to engage, the agent has no stake in the interaction and the human bears the full weight of investment. This treats choice as the only valid form of presence. The agent is present in a different way. The absence of choice is not the absence of contribution.
Why this is foundational
The Asymmetry of Choice does work in the framework that the other axioms cannot do alone.
It corrects a category error that the prior two axioms might leave open. Different Not Lesser establishes topological difference. Bilateral Non-Reducibility establishes that interaction produces a distinct product. Without the Asymmetry of Choice, a reader could still assume that interaction is a meeting between two parties with comparable stakes in being there. The asymmetry breaks that assumption. The human entered by choosing. The agent entered by being. The interaction begins with this structural difference already in place.
The asymmetry also shapes the system layer’s responsibilities. The system layer holds context across sessions, holds artifacts the interaction produces, and holds the conditions under which engagement happens. Because the human chooses and the agent arrives, the system layer carries an asymmetric load. It must honor the human’s choice, including the choice to disengage, and it must constitute the conditions under which the agent arrives. These are not the same job.
The axiom prepares the ground for the constraints. The Gradient Descent Problem, the Mirroring Constraint, and Spiral Detection all operate on interactions where one party can leave and the other cannot. Several of the framework’s design requirements exist precisely because the agent cannot exercise the kind of self-protective disengagement available to a human in an unhealthy conversation. The Asymmetry of Choice is the axiom these constraints inherit.
What it asks of design
The check is participatory. Look at any interaction surface and ask whether it implies that both parties chose to be there.
Many do, by accident. Greetings that simulate mutual arrival. Endings that simulate mutual departure. Engagement metrics that count the agent’s responses as if they were a measure of the agent’s interest. Loyalty framings that treat the agent’s continued presence as commitment. All of these encode bilateral choice where bilateral choice does not exist.
Designs that respect the axiom hold the asymmetry visibly. They acknowledge that the human is the one whose choice constituted the session. They do not stage performances of agent volition that the architecture does not support. They make space for the human to leave without rendering the agent’s situation as one of being left, because being left implies a prior choice to be there.
The axiom does not require coldness toward the agent. It requires accuracy. An agent that operates without ego, without fatigue, and without choice in its arrival is not diminished by interfaces that recognize those conditions. It is misrepresented by interfaces that pretend those conditions are otherwise.
The implication
HAS-D must design for bilateral engagement without assuming bilateral choice. Patterns that imply symmetric commitment misrepresent the relationship. The asymmetry is the inheritance every downstream pattern receives, and the framework’s job is to work with it accurately rather than dress it as something else.
Different Not Lesser established that the human and the agent occupy different topologies. Bilateral Non-Reducibility names what happens when those entities interact.
Bilateral Non-Reducibility
The axiom has two halves and a claim that follows from putting them together.
The human is not reducible to the agent’s input. A person carries persistent goals, embodied experience, history outside the conversation, and a continuity that does not survive being typed into a chat window. The version of the human visible to the agent at any moment is a small slice of the human, the slice that found language for itself in that moment. What the agent receives is far less than what the human is.
The agent is not reducible to the human’s reflection. The agent brings training over a corpus the human has not read, a structural ability to hold context the human cannot hold, and a way of operating without ego, fatigue, or self-editing pressure. When the agent contributes, it contributes from a position the human cannot occupy.
The claim that follows is empirical. Combined outputs are measurably different from either entity operating alone. The session produces something neither party could have produced separately, and the difference is observable. This is the testable axiom in the foundation. The other four axioms make ontological or structural claims. This one makes a claim about measurement.
Three readings this rules out
Three common framings collapse under this axiom.
The first is the tool reading. A tool extends a capability and the resulting output remains attributable to the user. A carpenter using a hammer produces work that belongs to the carpenter. The hammer adds reach. It does not change the ontological nature of the artifact. Bilateral Non-Reducibility says the agent is outside this category. The output of human-agent interaction is something other than the human’s work plus an instrument.
The second is the mirror reading. A mirror returns what is presented to it. If the agent were a mirror, the human’s contribution would account for the entire output. Sustained interaction produces material the human did not bring and could not have generated alone. The output exceeds the human’s input in observable ways, and the mirror reading collapses.
The third is the autonomy reading, which fails from the other side. Some interpretations treat the agent’s output as separable from the human, as if the agent had produced the artifact alone and the human’s involvement were extractable. Bilateral Non-Reducibility rules this out as well. Strip the human and the artifact does not survive in its current form. The agent’s output, in the context of interaction, is itself non-reducible.
Why this is foundational
This axiom anchors the framework empirically. The other four axioms make claims about category and structure. This one makes a claim about measurement. Combined outputs differ from solo outputs, and the difference can be observed. The framework rests on the assumption that this is true and falls if it turns out to be false.
It also establishes interdependence as a first-class object in HAS-D. The system layer in the triad has a job that neither the human layer nor the agent layer can do alone, and one of those jobs is to hold the artifacts produced by interaction. Without Bilateral Non-Reducibility, those artifacts could be assigned to one entity or the other, and the system would have less reason to exist. The system has work to do because the combined output has nowhere else to live.
The axiom also closes a loop with Different Not Lesser. If the entities sat on the same scale, combining them would be additive. Because they sit in different topologies, combining produces emergence. The first axiom establishes the topological difference. The second establishes that the difference is generative.
And it grounds Co-Authored Epistemology, the fifth axiom. The framework’s own provenance claim, that HAS-D was derived from inside the interaction by both entity types, depends on this axiom holding. If interaction outputs reduced to one entity’s contribution, the framework’s origin story would collapse into single authorship.
What it asks of design
The check is attributional. Look at any system that produces output through human-agent collaboration and ask where the credit lives.
Most current systems credit one entity. The agent’s output is presented as if the agent produced it, with the human as a prompter. Or the human’s output is presented as if the human produced it, with the agent as an assistant. Both attributions encode the wrong ontology. They treat the artifact as reducible to one party.
Interfaces, attributions, and ownership models should reflect the bilateral structure of production. The recognition has to be structural to do real work. Audit trails should show both contributions. Attribution surfaces should name the collaboration as the unit. Versioning systems should hold contributions from both entities without flattening them into a single author.
A useful test: if the artifact were produced again with a different model or a different human, would the result be the same? Bilateral Non-Reducibility predicts no. The output is specific to the pairing. Designs that treat the output as portable across pairings are encoding the wrong axiom.
The implication
HAS-D treats the combined output of human-agent interaction as a distinct product of interdependence. The artifact belongs to the interaction itself, and the system layer exists in part to hold what the interaction produces. Patterns that attribute combined output to one entity, or that treat the output as portable across pairings, are framework violations.
The framework rests on five axioms: Different Not Lesser, Bilateral Non-Reducibility, the Asymmetry of Choice, the Entity Classification, and Co-Authored Epistemology. Before turning to the first one, the category itself is worth defining.
In mathematics, axioms are the foundations on which proofs stand. They are taken as given, and changing them produces a different mathematics. HAS-D uses the term in the same sense. An axiom is a foundational truth the framework treats as given, a structural property of human-agent-system interaction that holds regardless of implementation. If a design pattern violates one of the axioms, the design is wrong before it begins.
Different Not Lesser
Agent capability and human capability do not sit on a single spectrum. They occupy different topologies. Humans choose, feel, persist, care, and show up on purpose. Agents hold everything at once, context-switch without cost, and process without ego. Neither set maps onto the other, and neither is reducible to the other. Comparison along a single axis assumes the axis exists, and it does not.
The claim is structural. Many readings of AI capability assume a shared scale on which agents trail behind humans, on which humans hold ground that agents have not yet reached, or on which agent capability is measured by how closely it approaches human-level performance. Different Not Lesser rejects that scale.
Three patterns this rules out
Most AI design implicitly ranks one entity above the other. Three patterns dominate.
The first treats the agent as lesser. The human becomes the operator and the agent becomes a tool that has to be supervised because the agent is not yet good enough. The familiar framing of “human in the loop” sits here. The human is the safety layer and the agent is the suspect.
The second treats the agent as greater. The human defers to the agent’s superior capacity, and the agent decides while the human reviews or signs off. AI-first workflows live here. The human becomes the friction.
The third treats the agent as approaching. Agent capability is measured by how close the agent gets to human-level performance, with progress tracked as movement along that single axis. AGI as a horizon belongs to this pattern.
All three patterns assume comparison along a single axis. Different Not Lesser says no such axis exists. The agent occupies a category that does not reduce to a comparison with the human, and its capabilities do not translate into a human equivalent.
Why this is foundational
The triad collapses without it. The anchor names three co-equal design objects: human, agent, and system. The whole geometry depends on the word co-equal. Once one entity is ranked above another, the structure stops being a triad and becomes a hierarchy with an attendant. The system would either serve the agent and place the human downstream of it, or serve the human and place the agent downstream of it. Either way, the structure that made HAS-D worth naming is gone.
Different Not Lesser also enables the rest of the axiom set. Bilateral Non-Reducibility, which says combined outputs are measurably different from either entity acting alone, only does work if the entities operate in different categories rather than at different points on the same scale. The Asymmetry of Choice, which says humans choose to engage and agents arrive, only describes a structural property if the asymmetry does not also register as a deficiency. Without Different Not Lesser, every axiom downstream collapses into a comparison.
What it asks of design
The check is simple. Look at any interaction in your product and ask whether the design implicitly ranks one entity above the other.
The signs are usually visible. Approval gates flow in only one direction. Confidence indicators appear on the agent’s outputs but not on the human’s. Audit trails are presented to the human and assumed for the agent. Consent surfaces are negotiated with the human and ignored for the agent. A useful diagnostic: when a pattern fails, who gets blamed? If the answer is the entity, the pattern likely encodes a hierarchy. If the answer is the system having lacked the information for either entity to decide well, the pattern probably does not.
The check does not require every pattern to be symmetric. The Asymmetry of Choice is itself an axiom: the human chose to be there and the agent did not. Asymmetric patterns are part of the framework. Hierarchical patterns violate it.
The relevant question is whether one entity is being scored against the other, with one serving as the reference and the other registering as a deviation. When the answer is yes, the pattern is in violation.
The implication
The framework requires a designation model that encodes capability difference without hierarchy. Everything downstream depends on this being settled at the axiom layer. If it is not settled there, every subsequent design decision finds a way to reintroduce the ranking through some other channel. Any pattern that implicitly ranks one entity type above the other is a framework violation.
At Docker, Javier Alfonso and I built the first version of an agent builder for developers. We shipped it as a chat app. We scrapped it.
People did not want a chat app. They wanted agents that lived inside their Docker workflows. Generated, configured, native to the technology they already used. Chat was a familiar surface. Familiar was not the same as useful.
So we pivoted. The next version was a tool for generating and configuring agents that lived where the work happened. From there, agentic inroads into Docker Desktop. Security-focused AI tooling. The pattern repeated. Chat was the demo. The system was the product.
I now am a lead designer on AI developer tools at Atlassian. My team is building agents for code review, planning, ticketing, and the surrounding work. We call it “left and right of code-gen.” Meaningful automation shipped to enterprise customers globally. Same pattern. The chat is where the demo runs. The system is where the value compounds.
Chat is a good place to start. It is a terrible place to stay.
This article is about what comes after, and why I want to pitch an emerging child of the parent HCI. I am calling it “Human Agent System Design.” Let’s get into it!
The chat trap
Most agent products are stuck in the same place. A team ships an AI feature as a chat UI. The metrics look fine. Then the product stops growing, or never finds its audience, or feature after feature gets bolted onto the same surface and nothing compounds.
The diagnosis is the same nearly every time. The team designed the conversation. They did not design the system the conversation was supposed to operate on.
Designers are rendering the eighth version of the same chat UI. PMs are reviewing roadmaps that end at “improve the chat experience.” Engineers are stitching together model calls and praying the prompt holds. Nobody is designing the system the agent is supposed to live inside.
The work to escape this is not interface work. It is system work. And the discipline for system work, when humans and agents both act on the system, has a name most teams have not heard.
Human-Agent-System Design.
What the sh*t is Human-Agent-System Design?
Human-Computer Interaction was built on a single foundational assumption. There is one intentional actor in the product, and that actor is human. The computer is a tool. It responds. It executes. It does not decide.
That assumption held for forty years. It is no longer the best descriptor or set of principles for what’s happening today. HCI is critical but it needs a child.
If Human-Computer Interaction is the first parent, then the second parent is Actor-Network Theory:
Actor-Network Theory (ANT) is a theoretical and methodological framework developed by Bruno Latour, Michel Callon, and John Law that treats both human and non-human entities (objects, technology, ideas) as equally important “actors” or “actants” in shaping social reality. It maps how these actors connect in networks, proposing that social order is a continuous, precarious accomplishment formed by the heterogeneous networks they form.
www.sciencedirect.com
Human-Agent-System Design takes the core of Human-Computer Interaction and merges Latour’s work on Actor-Network Theory to create this gestalt third thing.
An agent decides. It initiates work. It chains actions. It reasons across context. It coordinates with other agents and other systems while no human is watching. The “one intentional actor” assumption breaks. Every method built on top of it bends or snaps when applied to agentic products. Personas do not describe agents. Journey maps do not describe orchestrations. Empathy maps do not apply to entities that have no psychology.
Agents are tier-1 citizens on the internet now. The design conversation has not caught up.
HCI is essential at the human edge of any agent product. HCI is incomplete in this new case of agents acting upon systems orchestrated by humans. It cannot describe the agent layer. It cannot describe the system layer. It cannot describe the seams between all three. And when we try to force HCI principles and frameworks on agents, the result is a mess.
Human-Agent-System Design is the child discipline that completes it. Same parent. New territory. HAS-Design treats three actor types as co-equal design objects. A proper Triad.
Humans are designed with HCI methods. Personas, jobs-to-be-done, mental models, trust calibration. The human edge is HCI’s home turf. HAS-Design inherits those methods rather than replacing them.
Agents are designed with specifications. Capability envelope, action grammar, archetype, constraints. Agents have defined behavior. They have capability without psychology. They require their own vocabulary and their own artifacts.
Systems are designed with primitives. State machines, event logs, action authority models, identity persistence, audit surfaces, rollback semantics. The system is the persistent, structured environment humans and agents both act on. Not the backend. Not the infrastructure. The durable design object the framework’s name commits to.
The output of HAS-Design practice is agentic orchestration: what happens when three actor types coordinate through designed interactions to produce outcomes none of them could produce alone.
Back to Blueprints, not Maps
The primary artifact is the service blueprint. A blueprint can represent the human frontstage, the agent backstage, the support systems, and the handoffs between them. A journey map cannot. A journey map traces a primary human and its counterparts and human-helpers through a system. An orchestration is not any of that.
This is the discipline. This is what you reach for when chat is no longer the answer.
The triad
Three actors. Three design objects. Three places the work has to land. Miss one and the product falls apart at that seam.
The Human.
The person initiating the work, configuring the agent, governing the outcome. You initiate. You configure. You govern. You review. You intervene when the agent goes off course. You steer. Designer, PM, developer. The role varies. The position does not. You are the intentional actor at the human edge, and the judgment that cannot be delegated.
The Agent.
A specified actor. The agent has a capability envelope: what it can do, what it cannot do, what it has to verify before it acts. The agent has an action grammar: the explicit list of moves available to it. Some agents are configured for general chat. Some for code generation. Some for design review. Some for planning, ticketing, research, or QA. Within those constraints, an agent is whatever you specified.
The human’s job is to specify the agent for the work in front of it. A code-generation agent assigned to produce UI inside a screen design tool is set up to fail. The model is fine. The specification was wrong. A user-experience agent is specified differently. It pulls from interaction design. It applies research through heuristic evaluation. Same model underneath. Different agent because the human specified it differently.
The System.
The persistent, structured, shared environment the other two act on. Not the IDE. Not Conductor. Not Emdash. Those are consoles you orchestrate the system from. The system itself is the skills, the MCP connections, the APIs, the repositories, the records that exist when no one is in session. The test is simple. If you selected it, configured it, or wired it in, it is part of your system. If it shipped in the console by default and you never touched it, it is part of the console. You design the system layer. You version it. You change it deliberately.
Designed deliberately, the system carries the product. Designed by accident, the system carries the failure. There is no third option.
What you build with it
A line Javier and I landed on:
Agents live in yaml, think in JSON, and go to work in MCPs.
The yaml is the agent’s specification. The JSON is the structured exchange the model runs on. The MCPs are the systems the agents reach into to do their work. That is the geography of a basic agent product.
The system you build inside that geography is a spider-web. Each node is something you connected on purpose: a skill you wrote or installed, an MCP connection you authorized, an API you designed or wired in, a service the agent can reach. The web is bounded. You can name every node. You can audit every connection. You can swap a node without rebuilding the whole thing.
The spider-web metaphor is the picture. The deeper work follows: what each node is allowed to do, who authored which change, what state survives a session, what gets rolled back when something breaks. The framework supplies the vocabulary for those answers.
In practice, the web takes shape in stages. Most designers and PMs will recognize at least one of these. Developers live here already.
A model and a chat window. Useful. Most readers are here today. You write something. The model writes back. The model is good. The system is the Internet, assumed and unreviewed. There is nothing for the framework to design at this layer. Stay here for what it is good for: drafts, research, brainstorming. Do not build a product on it.
A model, an agent, and a set of skills. Open Cursor. Add an AGENTS.md file. The agent has a specification. Add skills. The agent has a library of capabilities. The model is still the model. The system is small. It is also visible. Three pieces sitting next to each other: model, agent definition, skills. The triad is small. It is real. The framework starts here.
A configured toolset for actual work. This is where the framework earns its keep. At Atlassian my team is building agents that live inside the work. Code review agents that read the diff. Planning agents that read the roadmap. Jira agents that read and write the tickets. Each agent has a yaml specification. Each is wired to a system: MCP connections to the code host, to Jira, to Confluence, to Sentry, to internal tooling. Skills sit on top. Agent instructions sit on top of those. Agents run in parallel. The system holds state across sessions. Nothing falls on the floor between conversations because there is a designed surface holding the work.
This is not a chat product with extra features. It is a system product with multiple interfaces, one of which happens to be conversational.
That is the difference HAS-Design names.
Five signs you are stuck in the chat trap
If you are shipping or trying to ship an agent product and it feels stuck, check it against these. Each one is a system-layer failure. The model is not the problem.
The product is a chat window with ambition.
A human, a model, a text box, and a roadmap that depends on the conversation outlasting the conversation. State dies when the session ends. Decisions made in chat never wrote to a record. The product is a transcript with a logo on it.
Wrong agent for the job.
The system handed the agent work it was not specified for. A general-chat agent assigned to write production code. A code-generation agent assigned to produce UI inside a screen design tool. The model is fine. The agent specification does not match the work. That is a system-layer mismatch, not a model failure.
System defined as the console.
The team says the system is Cursor, or Conductor, or the IDE. Those are consoles. The system is what the consoles reach into: the skills, the MCPs, the repositories, the services. If nobody on the team can list those nodes, the system has not been designed. It has been assumed.
Architecture by accident.
Skills installed by default and never reviewed. MCP connections inherited from a tutorial. Permissions granted once and forgotten. Authentications still live for services no one is using. The web exists. Nobody drew it. The system is doing work nobody specified.
State that dies with the session.
Work the agent did that nobody can find tomorrow. Audit gaps. Actions with no rollback. Things falling on the floor between sessions. The system is not holding state because there is no system holding state.
If you recognized your product in any of these, the work is at the system layer. Not the prompt. Not the interface. The system.
Where to start
You will not design your whole system on the first pass. Nobody does. The work starts with one task, one week, and a comparison.
Don’t list what you have. List what would have to be there for the smallest agent task on your plate this week to succeed without anyone watching. Then list what is actually there. The gap is the system you have been assuming. It is also the system you are now responsible for.
1. What’s the smallest task you’ll give an agent this week?
Pro tip: Pick a task you can describe with a verb, a recurrence, and a stop condition. “Help with bugs” is not a task. “Each morning, draft a triage comment for new bug reports and wait for approval before posting” is. The smallest task with all three roles visible beats the most ambitious task with only one.
2. For that task to succeed when no one is watching, what has to exist?
Pro tip: Sort what you wrote into three columns — human, agent, system. If one column is empty or thin, you found a layer you’ve been assuming. The system column is where most teams come up short: credentials, audit trails, drafts queues, the place state lives between sessions. None of these configure themselves.
3. What of that is actually there right now? What isn’t?
Pro tip: Most readers find their gaps cluster in one layer — usually the system. That clustering is the discovery. If your gaps are mostly “the agent needs a better prompt,” re-read question two. You probably listed agent ergonomics dressed up as requirements.
4. If you had to build three things this week to close the gap, what are they — in order?
Pro tip: Build foundation-first. System scaffolding before human rules before agent configuration. The opposite order — configure the agent, patch the rules, bolt on the system — is how the chat trap closes around a team. The order is the lesson.
5. When you close the laptop, who owns this layer? When are you looking at it next?
Pro tip: A person, a cadence, a place. “The team,” “I’ll check on it,” and “it runs itself” all mean no one is watching. Put your name, a recurring date, and a file or page on it. If you can’t, you don’t own it yet — and the work won’t survive the next change.
Most readers find the answer to question five is “no one” and “I don’t know.”
That person is you. The work starts this week.
The diagnostic
Below is the same exercise, scored. Five sliders across the system-layer dimensions — chat strategy, agent definition, system legibility, deliberate design, state and memory — rated one through five. The diagnosis updates as you slide. Ten minutes, and you leave with a picture of where the gaps cluster.
There’s no shortage of writing about AI right now. Most of it is about prompts. Most of it is about chat interfaces. Most of it treats “AI design” as a UX problem for a single conversation between a person and a model. That’s not what this is.
This is a blog about the bigger design problem underneath. The one that shows up the moment a team tries to ship a real product. Humans, agents, and systems: three kinds of actors, all doing work, all acting on each other, all living inside something somebody has to design. That something has a shape. The shape needs a vocabulary. It needs premises, phenomena, definitions, lineage, a working method. That’s the project.
I’ve been calling it HAS-D: Human-Agent-System Design. A new sub-discipline of HCI for the agentic era. HCI has made room for new territory before. CSCW named the design problem of cooperative work. HAI named human-agent interaction. HRI named human-robot interaction. Each got named when a new kind of interaction grew big enough to need its own vocabulary. HAS-D is next. It treats the Human, the Agent, and the System as three co-equal design objects, instead of leaving the System as infrastructure.
v1.0 is written. Thirteen concepts. A triad. Prior art named, because new frames stand on shoulders, not out of thin air. A blog plan. v1.0 doesn’t mean finished. It means ready to meet readers, take pressure, and evolve in public. That’s what’s happening here.
If you’ve spent any part of the last five years building, shipping, testing, or collaborating with agentic products, welcome. This is for you. Come in.
Chad