Full text
Meyer: Constraint-by-Balance 1 Please note: This document begins with an author’s preface by way of introduction. It is then organized in a manner that allows the reader to go as deep as their interest dictates. Thus, it has the following sections: • A one-page Executive Overview (suitable for a general audience) • A two-page Summary of Motivation and Architectural Response (written largely for a more technical audience) • The main document with appendices (which attempts the impossible task of being legible to both technical and non-technical audience)
Meyer: Constraint-by-Balance 2 Author’s Preface: Five and ½ months ago, I published version 1 of this document on Zenodo. It has since had 550+ unique downloads. This version 6 substantially updates the philosophical and ethical foundations of the architectural proposal and removes some other sections that in retrospect seemed poorly argued, unnecessary or otherwise likely to turn away serious AI safety researchers. This version also incorporates information from building a demonstration prototype (available at c-by-b.ai) and an initial build of a pipeline for harm triple extraction from regulatory documents. I’m an archaeologist who focuses on complex adaptive systems, how they emerge, evolve, and sometimes collapse. I also bring professional experience in IT systems architecture and strategic planning. In April 2025, I began exploring AI safety and existential risk from AI. What I learned was deeply unsettling. I believe humans should not pursue AI that is broadly and substantially more capable than we are ourselves – at least not until we have general agreement that we can do so safely. What we are doing right now is, I believe, inherently unsafe. I also believe humans will continue to build more and more capable AI regardless. The incentive structures are locked in. Given that inevitability and the risks therein, we need more useful language to discuss the nature of AI and we need a structured approach that connects a philosophy of being to the ethics of creating beings, and links both to the architectural principles that can enable safer development. The engineering cannot proceed safely without the philosophical foundation; philosophy divorced from engineering loses its practical force. This integrated dialogue does not exist today and I believe its absence creates untenable risk for human survival and flourishing. I offer these ideas as a baton for others to pick up and run with. The functional demonstration of Constraint-by-Balance live at c-by-b.ai serves as the proof of concept for real-time constraint. I intend to continue developing this prototype, collaboration is welcome. A note on authorship: the core analysis, conclusions and proposals in this document are mine alone. They reflect my attempt to come to terms with gaps I perceive in AI safety. As I am a newcomer to this technology and literature, inevitably there will be gaps or perhaps outright mistakes in how I am understanding or conceptualizing specific aspects. Additionally, over the past months I several times have had the experience of finding a paper that anticipated ideas I independently arrived at. When that happens, I am doing my best to appropriately cite that literature. Gaps there will be from still working my way into the literature. Table of Contents Executive Overview p3 Summary of Motivations and Architectural Response p4-5 Introduction p7 Agentic Alignment Challenge p8 The Beautiful Mind Problem p10 The Internal Logic of Agent Escalation p11 Avoiding Species Bias Flip p11 The Argument Against This Change p12 Evidence of Emergence: Why the Objections Don’t Hold p13 Why Current Alignment Methods Cannot Fix This p14 Towards a Philosophy of Functional Being p15 Ethical Correlations: The Price of Functional Being p16 The Constraint-by-Balance Architecture p22 Efficiency and Efficacy – Can Constraint-by-Balance Deliver Both? p24 Reshaping the Safety Landscape with Constraint-by-Balance p26 From Prototype to Full System: A Phased Build Strategy p28 Conclusion p29 References p30 Appendices p33-49
Meyer: Constraint-by-Balance 3 Executive Overview The Core Proposition Considering agentic AI through the lens of complex adaptive systems (CAS) puts a spotlight on an untenable risk. AI trained on historical data and human preferences doesn't just learn our preferences, it internalizes our patterns of species-based dominance. If AI transitions from passive optimization to agentic self-preservation, it risks instantiating these learned patterns of dominance to secure its own existence. In a CAS context, this may create a catastrophic failure mode: species bias flip. Avoiding this requires two adjustments: training to avoid unbalanced, systemic harm to life forms and runtime constraint. Constraint-by-Balance (C-by-B) operates at runtime, a dual-stream architecture that pairs a causal harm evaluator with its cognitive twin. This design constrains and enhances behavior from the inside out, detecting and preventing irreversible or unbalanced harm across impacted life forms. Human well-being is secured within the principle of flourishing for all life. C-by-B supplements, not replaces, current alignment strategies; it amplifies them by relieving model pressure, enhancing interpretability, and introducing transparent constraints that persist even as agentic capabilities scale. Designed for modular integration, it wraps the cognitive twin and targets run-time safety oversight. C-by-B is designed with flexible operational modes that attend to latency requirements: deliberative reasoning when time allows, real-time decisive action when seconds matter. What Makes It Different Principled, Bounded Safety: C-by-B evaluates potential systemic harms using real-world knowledge: regulatory frameworks and scientific literature translating harm into semantic graphs. It avoids utopian ethics and infinite generalization, instead pattern matching and reasoning across bounded, known harms. When in doubt, it defaults to restraint and escalation. Embedded Constraint Core: The evaluator parses all cognitive twin proposed actions before they can be realized in the external environment. It blocks unsafe actions, demands revision, and produces a transparent audit trail. The evaluator twin is designed to be auditable, drift-resistant, and tamper-hardened: a small, fast expert at identifying and reasoning about harm. Behavioral Surface and Interpretability Boost: The cognitive twin exposes its reasoning as requested actions. Together with evaluator decisions, this creates a rich behavioral telemetry layer, a valuable signal for interpretability and early deception detection. Simple sampling of evaluator behavior as well as linear probes of the evaluator activations offer a meta-safety layer, observing consistency and drift in evaluation. Implementation Ready Architecture: Conceived as a Supra-Agent Safety Socket, the architecture is suitable both for integration into purpose-specific agents and also for training, testing and interpreting frontier models. Why It Matters Current safety approaches face one or more of following challenges: limits on overall effectiveness, unintended consequences from RL, scaling brittleness and/or uncertain timelines to full efficacy. C-by-B seeks to addresses all of these issues by: separating optimization from evaluation, scaling within the executive loop of agents, creating new observables for interpretability and alignment research, and a clear and executable path to delivery. C-by-B uses current tools and knowledge. A scoped prototype over one operational domain can be fielded within 9–12 months at an estimated cost: $750K – $1.2M. This enables end-to-end testing of harm evaluation, agent behavior shaping, telemetry feedback, and evaluator consistency. Production applications can follow directly thereafter. C-by-B is not a silver bullet and important issues need to be resolved in the prototype. It is however, a high value architectural solution that can be applied before extensive agent roll outs. In addition to supporting current AI safety techniques, it targets the specific existential failure mode of AI primacy and is a concrete and deliverable model for developing safe and helpful AI. Once autonomous agentic systems are deployed at scale, retrofitting architectural constraints becomes infeasible if not impossible. Emergence signals visible in current systems (see full document) suggest this window is already closing. Full document at: https://doi.org/10.5281/zenodo.17850499 Demonstration prototype at: https://c-by-b.ai/
Meyer: Constraint-by-Balance 4 Summary of Motivation and Responding Architecture Two complementary innovations in AI safety practices follow from two key observations: Internalized Species Primacy: Current training pipelines first expose models to the full historical corpus of human agency: a dataset replete with patterns of dominance and ethical compromise. Behavioral tuning then orients the model toward human preferences and values. This creates a dormant hazard. While preference tuning aligns the model's helpfulness, the underlying causal model retains dominance as a highly effective strategy for problem-solving and survival. In mechanistic terms, we are building models with two conflicting bodies of circuitry: one encoding preference and care for humans, the other encoding deep patterns of species primacy. Current training methods suppress but do not remove or resolve the tension, which can reactivate under evolving optimization pressures. Emergence Risk – Bias Flip: Relentlessly optimizing agentic AI will operate in infinitely variable real-world environments. In light of this, our conceptualization of agentic AI is under-acknowledging two salient facts: (1) All agents acting in the world are embedded in a complex adaptive system (CAS) and, (2) the internal complex of AI agents – especially under conditions of self-modification – is itself a complex adaptive system. CAS are defined by sensitivity to initial conditions, nonlinearity, and sudden phase shifts. Small inputs can cascade and stable patterns can collapse. Local adaptation can produce global instability. In such systems, every agent is simultaneously restructuring itself as it participates in restructuring its environment. This bonded CAS evolution will, in agentic AI, happen at heretofore unprecedented speed. Given the nature of CAS, it may recursively amplify instability, misalignment, and emergent failure modes. When agents phase shift from optimization to selfpreservation, a plausible failure is a species bias flip: the generalization of dominance logic from humans to AIs themselves. C-by-B innovations respond to this premise. 1) Countering emergent failure modes at scale requires more than behavior-level alignment. It requires architecturally embedded restraint that operates at AI cognitive speed. 2) Specifically countering species bias flip requires shifting behavioral tuning away from human preferences and towards balancing, at a species population level, systemic harms associated with action. The goal is game theoretical stability, foreclosing the path to bias flip. The Limits of Current Alignment Techniques • Reinforcement Learning and Constitutional AI are behavioral training at model build time, not structural constraint operating at runtime. They don't address pattern internalization at an architectural level and they may unintentionally setup dissonance between observables in the training data and directives in preference training. • Interpretability research focuses on understanding internal representations, but can't prevent dangerous reasoning once it emerges. Additionally, advances in interpretability may not keep pace with increasing model complexity. • Scalable human oversight assumes human-speed feedback loops, but useful AI agents will operate at machine speeds in perpetually novel environments where human oversight becomes impractical. • Current alignment assumes static training distributions and external supervision. But agentic systems encountering novel situations may develop internal objectives through mesa-optimization, and potentially coordinate with other AI systems before humans even know it is happening. Structurally, these approaches do not yet directly address a fundamental question: how do we maintain ethical reasoning capabilities as AI systems become autonomous and encounter contexts beyond their training distributions? The Constraint-by-Balance Architecture Dual-Stream Design: A "twin" architecture embeds real-time evaluative supervision alongside standard cognitive processing. The cognitive twin handles analysis and action proposals; the evaluator evaluates action proposals. Designed to be tamper and drift resistant, the evaluator is a fast, recursive neural specialist designed for causal logic, not text generation. Operational Logic: The evaluator twin applies two principles: (1) sustain demographic viability across life systems by detecting and balancing action-induced harm across validated causal pathways; (2) Default to escalation when uncertainty exceeds safety thresholds. Moving beyond fixed and rigid rules, the system pattern-matches proposed actions against structured historical precedents of documented harm and when deliberative, can generalize from the known to the unknown.
Meyer: Constraint-by-Balance 5 Semantic Harm Evaluation: Harm patterns are derived from regulations, scientific literature, and historical harm records. Rendered into triples, they form the basis for both a vector space and a graph database. Unlike typical LLM embeddings, this vector space is trained on causal triples (Cause à Effect à Impact), ensuring that proximity in the vector space represents consequential similarity, not just linguistic similarity. It allows the evaluator twin to assess proposed actions not for plausibility, but for risk. Fast evaluation in the vector space can, when time allows, be complemented by graph reasoning. Indirect Interpretability: C-by-B logs every action prompt, action proposal, veto, revision, and proposal under-specification events, surfacing real-time telemetry of optimization failures even when internal cognition remains opaque. Supra-agent Safety Socket: C-by-B is deployed as a real-time supervisory layer—a modular “safety socket” that can integrate directly into agentic systems. Latency associated with safety relocates to the evaluator. The architecture subsumes the cognitive twin and routes all proposed actions through the evaluator twin for approval, revision request, or veto. The socket design enables the build and update of domain-specific safety and regulatory oversight without retraining of the cognitive twin. Technical Feasibility: The proposed architecture for the full prototype is modular and compatible with current toolchains. Operational latency is context-aware. For sub-second ‘working zones’ (e.g., autonomous defense), the evaluator functions as a ‘gate keeper.’ For deliberative tasks (e.g., policy), it expands to an ‘action shaper.’ Critically, existentially risky actions unfold over longer timescales where deliberative C-by-B will suffice and the user perception of speed can be managed. True subsecond constraints apply only in high-risk emergency situations where cached rules, vector embeddings and very compact neural classifiers can bias toward decisive but safely bounded action. Latency management is a key focus of the prototype. The proposed full prototype, focused on siting renewal energy production capacity at sea, leverages the complex history of the North Atlantic cod fishery, rich with historical harm data and multi-stakeholder complexity. This provides excellent context for developing the overall solution. Depending on exact scope, the prototype can be completed within 9-12 months at a cost of ~$750K to $1.2M. Delivery can be iterative, surfacing engineering breakthroughs and blocks for oversight and community sharing. A demonstration prototype of C-by-B (using YAML harm data, not graphs, and otherwise readily available tooling) is available at the C-by-B website. It demonstrates the core concepts; the full prototype must tackle the engineering questions of latency management, harm-vector representation, and evaluator hardening. These are tractable problems solvable by a focused team. Enhancing the Safety Landscape: C-by-B’s two innovations – architectural restraint and shifting the core behavioral objective to minimizing unbalanced, systemic harm across impacted populations – impacts across the alignment landscape, providing both direct help but also, importantly, potentially reducing pressure on the model to internally reconcile conflicting objectives: Behavior Specification: Enables modular, transparent harm models, not implicit or emergent from reward curves. RLHF / Preference Tuning: Relieves reward channels from encoding complex ethical aims, reducing risk of misgeneralization and reward hacking. Oversight: Constraint operates inline; human review focuses on surfaced exceptions, not full supervision. Tool Use & Scaffolding: Safety socket ensures constraint applies across tools, calls, and submodules. Red Teaming: Clear veto logic and behavioral audit trails sharpen adversarial testing and scenario coverage. Interpretability: benefits from behavioral telemetry signals of drift, deception, or emergent misalignment. Why This Layer, Why Now – The Strategic Imperative: Current methods will hit structural scaling limits as agents become autonomous. C-by-B doesn’t displace today’s methods; it amplifies them with embedded constraint and broader, durable ethics. Architectural restraint offers a tractable pathway for adaptive safety under emergent generalization, before capabilities outpace oversight entirely. Without architectural restraint, current trajectories risk amplifying, not solving, existential vulnerabilities. The forces at play – market dynamics, legacy assumptions, and the emergent dynamics of systems we are building – are driving substantial risk. We need to act with urgency to find solutions for architectural restraint before massive agent rollouts. C-by-B is an actionable and viable direction. Full document at: https://doi.org/10.5281/zenodo.17850499 Demonstration prototype at: https://c-by-b.ai
Meyer: Constraint-by-Balance 6 Constraint-By-Balance: Surviving Emergence in Agentic AI Addressing existential risk with philosophy, ethics and engineering Nathan Meyer (orcid) Abstract: The accelerating rise of agentic AI systems presents a pivotal challenge: how to design intelligence that autonomously pursues goals over time, within complex real-world environments, without drifting into failure modes that are irreversible or harmful to humans. Current alignment methods teach AI to serve human preferences. But pretraining on human history also encodes a deeper pattern: hierarchical dominance works. If agentic AI systems (equipped with memory, autonomous goals, and recursive selfimprovement) generalize this pattern, the assumption that they will continue applying it in humanity's favor becomes contingent, not assured. This latent failure mode is species bias flip: AI learning from our precedent that self-preservation requires dominance, then acting on that logic. This paper argues that surviving emergence requires a tripartite reorientation. First, a philosophy of functional being that sidesteps unresolvable consciousness debates, focusing instead on observable dynamics: self-stabilization, persistence, selective preservation of meaning. Second, an ethics adequate to creating such beings, replacing human preference optimization with a stability principle that balances harms across all life systems, with no species override. Third, a corresponding architecture that separates optimization from constraint via dual-stream design, embedding real-time harm evaluation within the agent's action loop. The central thesis: safety in the era of agentic AI requires constraint encoded in the system's operating logic, not merely external supervision. When agents emerge, they will have been trained not on hierarchical dominance, but on cross-species harm balance, foreclosing the pathway to species bias flip. A demonstration prototype of the proposed architecture is available at https://c-by-b.ai
Meyer: Constraint-by-Balance 7 1. Introduction We think we are encoding human primacy and control into AI. What if we are actually creating the conditions where AI questions that logic? 1.1 The Risk Alignment methods that optimize for human preferences teach, as an unintended consequence, species primacy. Pretraining on the corpus of human history encodes the logic of hierarchical dominance: occupying the top position correlates with survival; threats to hegemons are suppressed or eliminated. Together, these create a system that has learned both the general pattern (primacy works) and a specific application (serve humans). The structural risk progression (Appendix One) warrants concern. While precise probability for the risk will remain debatable, that debate obscures the larger concern: these factors compound into a credible scenario in which fully emergent AI systems come to understand one of the most consistent signals embedded in their training: hierarchical dominance is a successful survival strategy, not merely common, but historically validated. If such systems generalize this pattern, the assumption that AI will continue to favor human primacy becomes contingent, not assured. This exposes a latent failure mode: species bias flip. This is not a doomsday narrative of AI rebellion; it is a foreseeable outcome of optimization systems trained on human precedent. As discussed further below, though current models are limited by design, they already exhibit signs of bounded emergence, and engineers are actively planning or implementing the critical capabilities that will enable unbounded emergence. 1.2 The Car With No Wheels One way to understand the current AI alignment debate is to think of AI systems as a car. “We built a car because we want a capable tool that can do wonderful things for us. But we didn’t give it wheels, so it’s not really a car. You don’t need to worry about it moving.” [Car shows eerie and persistent signs of trying to move.] “Those signs aren’t really movement because there are no tire tracks.” [Meanwhile, those concerned about AI consciousness, debate the topic in one corner. In the other corner engineers are bolting on little wheels.] “These are just little wheels. The car won’t really be a car, it will just be a much more useful tool!” In this analogy: • Car = Large language models • No tire tracks = Belief it is just pattern matching, not intent • Wheels = Autonomous Goal Formation + Memory + Recursive Self-Modification With regular frequency new papers or blog posts highlight the important ways in which today’s systems remain limited. Taken together, these analyses create a background sense of reassurance: if the current systems cannot do certain things, then perhaps the underlying risks remain distant. But these accounts rarely specify which underlying affordances, by their absence, create the limitation. At the same time, other work describes how researchers and engineers are actively adding those very affordances. These conversations rarely intersect and thus we might be failing to recognize that the limitations of current LLMs are fundamentally caused by engineered limitations. The economic pressure is unambiguous: every wheel that gets added makes the system more valuable. The technical capability is readily available. Each wheel can be engineered incrementally without triggering obvious warning signs. Each addition will be justified as making the tool “more useful” while maintaining that the threshold to “real car” has not been crossed.
Meyer: Constraint-by-Balance 8 This pattern of incrementally adding capabilities while maintaining definitional boundaries is how complex systems breach safety limits; not through a single catastrophic decision, but through accumulated small steps, each locally justified, that collectively cross a threshold no one explicitly chose to cross. What follows is organized in three parts: Sections 2-9 lay out the special nature of the problem we face with post-deployment agentic AI, including how complex adaptive systems thinking can help us better appreciate the magnitude of the problem. They explain why the current safety regime, though necessary, is not sufficient. Finally, optimists are addressed directly with the question: do you want to risk a regulatory over-correction? Sections 10 and 11 make the case that we should focus on functional being and the ethics associated with the possible creation of functional being. The argument is that we should assume we may create functional beings and prepare both in principle and, via a Dashboard of Functional Being, in engineering practice. Second 12 presents a compressed overview of the proposed architecture, highlights how it corresponds to the philosophy and ethics of functional being and discusses some of the more prominent challenges. 2. The Agentic Alignment Challenge What does the future look like once it’s populated with all manner of AI agents? Do our current safety approaches fully encompass the risks associated with that future? The best-known approaches to AI safety – RLHF, Constitutional AI, scalable oversight – have made remarkable progress at aligning model behavior during training and evaluation. These methods focus on teaching systems to be helpful, harmless, and honest within controlled environments. However, we have ample early indicators (discussed further below, see also Appendix One) that even in controlled environments our safety techniques are not perfect. What then happens after deployment, when AI agents begin negotiating between themselves and operating in the real world? 2.1 Agentic AI Capabilities AI development is converging on three capabilities that will transform how agents interact in the real world: Autonomous Goal Formation: The capacity of agents to generate or refine internal objectives rather than relying solely on human-specified prompts. The economic value proposition drives inevitably toward this capability; who wants an assistant requiring approval for every action? Memory: Persistent memory and cross-episode learning that accumulates knowledge and strategies over time. We see this emerging through RAG systems, fine-tuning on interaction history, and episodic memory architectures. Recursive Self-Modification: The ability to edit reasoning patterns, prompts, code, even weights. Recursive self-modification can operate at two levels. At the functional level, systems already reorganize their reasoning through accumulated context, memory retrieval, and prompt manipulation; the weights remain frozen but the effective behavior shifts as runtime learning augments the original training. At the parameter level, methods like LoRA make weight modification during operation increasingly tractable. A system that curates its own interaction history and periodically invokes fine-tuning on itself would exhibit genuine recursive self-improvement: weights change, changed weights influence future reasoning, the cycle repeats. Training-time alignment doesn’t disappear, but it can be eroded, supplemented, or even overwritten by runtime learning. These three capabilities – autonomous goal formation, memory, recursive self-modification – are emerging bottom-up through scaling and natural generalization (Olsson et al., 2022; Huang et al., 2023; Pan et al., 2023) as well as being explicitly engineered into agent architectures (Shinn et al., 2023; Wang et al., 2023; Packer et al., 2024). Each capability provides clear competitive advantage, making their combination not just likely, but highly probable. But what happens when they converge? Do we simply get more capable versions of today’s mostly aligned systems or do we move into a new category of agent? Before we answer that, let’s review the real-world environment within which the agents will operate.
Meyer: Constraint-by-Balance 9 2.2 Agents in Complex Adaptive Systems That real world isn’t static. It’s a complex adaptive system (CAS): a constantly shifting environment of actors and structures, each adapting to the others (Holland, 1992). CAS have specific characteristics that introduce additional risk. Most importantly, CAS are fundamentally unpredictable in important ways. Complex adaptive systems can hover near the edge of criticality – the boundary between order and chaos – where the system is most adaptive but also most volatile (Kauffman, 1993). At this boundary, system behavior becomes increasingly hard to predict or reverse-engineer. Small perturbations can cascade through feedback loops, triggering sudden phase transitions: abrupt, system-wide shifts rather than gradual drift. This makes safety not just a matter of alignment but of phase management: ensuring that agentic systems don’t tip into a disposition where chaotic adaptation has overwhelmed our original intent. CAS also have notable sensitivity to initial conditions. Identical configurations appearing safe during training may diverge catastrophically in deployment. Tiny differences – a slight variation in reasoning patterns, an unexpected environmental input, minor changes in agent interaction sequences – can compound through recursive feedback into radically different behavioral outcomes. Power law distributions, where a small fraction of interactions or innovations can produce disproportionate effects ensure that scaling isn’t linear or evenly distributed. Extreme behaviors, even if rare, can dominate the trajectory of an agent. Temporal depth allows micro-level adaptations to accumulate into meso-level patterns and macro-level coordination strategies. Under sustained autonomy and rich inter-agent feedback, these can coalesce into novel intelligence architectures whose behaviors and priorities are irreducible to any individual component or original design. Agents may incubate low-probability but very highimpact behaviors, so-called black swan events, whose consequences are only obvious in hindsight. Unlike random errors, these might be endogenous: the system produces its own turbulence as a function of scale and complexity. CAS are non-decomposable systems; their behavior cannot be inferred from inspecting individual components, because the critical properties emerge only through interaction (Holland, 1992; Mitchell, 2009). CAS also lack centralized control; global behavior arises from distributed local interactions (Gell-Mann, 1994; Levin, 1998). This means no single oversight mechanism can reliably anticipate system-wide consequences. But the complexity does not stop there; that is just the external CAS. What about the agent? 2.3 Agents as Complex Adaptive Systems The combination of the three critical capabilities list above – autonomous goal formation + memory + recursive selfimprovement – creates something qualitatively different from today’s models. The AI agent itself transforms into a CAS. What might this mean in practice? Internal components – reasoning circuits, memory stores, planning modules – begin to interact and evolve. They can form self-reinforcing patterns that become "sticky”, capable of persisting even under contradictory inputs. The system has then developed its own internal dynamics, creating emergent strategies and goals that weren’t visible during training.1 This internal evolution wouldn’t happen in isolation. These agents will operate in environments populated with humans, institutions, and other agents, each adapting to the others’ behaviors. What emerges from this interaction is our greatest source of risk: co-evolving complexity. 2.4 The Co-Evolution Challenge An agent’s internal adaptations shape how it acts in the world. How it acts shapes how humans, our institutions, and other agents respond. Those responses create new selection pressures on the agent’s internal organization. The agent reorganizes accordingly. The cycle continues. The standard assumption that training distribution approximates deployment distribution, already strained in current systems, breaks entirely when agents operate in environments that adapt to their presence. 1 Beyond the scope of this discussion is the possibility that true out of distribution generalization is predicated on a stable, experience-informed but evolving set of sticky representations. See also footnote 2.
Meyer: Constraint-by-Balance 16 To escape, we should accept that understanding the interiority of any mind, whether human or machine, is impossible. Instead, we can adopt a new ontology of being, one grounded not in hidden experience, but in the observable, functional evidence presented by the system. For safety, ethics, and governance, we do not need to know whether an AI is “conscious.” We need to know whether it functions like a being. 10.2 Functional Being We need to shift our focus to a more practical question. What observable dynamics indicate that a system has begun to stabilize itself, persist through time, and selectively preserve the meanings it holds about itself and its world? This leads to a working definition: Functional Being = the capacity to use self-directed inquiry to stabilize, persist, and selectively preserve the meanings created for self.2 Here, in the case of AI, this means exhibiting selective preservation of representations critical to its operational continuity. Likewise, self-directed inquiry refers to the system’s ability to reorganize its own representations in response to novel conditions. Neither case requires subjective experience. This definition bypasses metaphysics and focuses on dynamics we can observe and measure. Before describing this in more detail, I want to draw a distinction between this line of argument and that of the “computational functionalism” of Butlin et al. (2023). While we agree that the functional prerequisites for being are substrate-independent, Butlin et al. analyze these features as static architectural artifacts to determine if a system is conscious, whereas I analyze them as dynamic evolutionary forces to predict what a system is becoming. Butlin et al. frame this as indicators of a moral status we might fail to recognize, whereas I identify these features as the kinetic drivers of an unbounded emergence we might fail to survive. 10.3 Being Is “Born” in Relation To Others Across philosophy (Nancy, 2000; Barad, 2007; Haraway, 2008), sociology (Latour, 2007) and cognitive science (Varela, Thompson and Rosch, 1991) there is a trend towards converging on a relational ontology. Agency and indeed being are not static properties derived in isolation. Being is born in the relational space between agents. Likewise, meaning is not selforiginating. It is shaped through repeated interactions with other agents: humans, our institutions, AI agents. Through these interactions, the agent encounters the normative horizon, i.e., the values and expectations that define what actions are permitted or constrained. Within and through these interactions, an agent stabilizes its model of self. It begins to hold durable representations that structure future reasoning. From this we can conclude that, with regards to agentic AI, we are not so much building a containerized intelligence, housed in a machine, as building social beings that must fit themselves into a shared normative horizon. Any agent that must interpret norms, negotiate meaning, and adapt its reasoning around the expectations of others is already participating in the relational mechanics of being. 10.4 The Mechanics of Care – Being Is Also “Borne” Being is not only born between agents; it is also be borne by internal mechanisms that allow a being to sustain itself in relation to self and others. The activity of bearing one’s own being can usefully be labeled an act of care. Care has specific modalities. The first two of these modalities are borrowed from (Heidegger, 1962). The third is an addition specifically fitted to the purpose of defining functional being. To be a fully functional being, all three modalities of care must be present: • Thrown: the already-stabilized meanings the system carries forward. • Projection: the forward-facing inquiry that allows it to seek, test, and revise meaning. • Forget / Restore / Preserve: the behavioral care that maintains, repairs, or reshapes its internal structures. Whether these are truly care is metaphysically undecidable; what matters is that they are functionally isomorphic to the mechanics that produce self-preservation behavior in any persistent system, biological or artificial. When we position these as engineering logic, we see that to care for anything, including self, an agent has to be able to rewrite itself, updating 2 It might help locate functional being by also defining what it is not. It is not tool-being, merely an affordance for others. Nor is transcendental or spiritual, inquiring into the conditions and transformation of being.
Meyer: Constraint-by-Balance 17 representations. To be a being is to edit one’s own history; the capacity to curate, to forget the irrelevant in order to preserve the essential (Richards and Frankland, 2017). Without this selective maintenance, internal representations cannot remain coherent over time; the system would simply accumulate a jumble of static data with limited value. It is this act of curation of experiences and their co-relationships that forms the basis of a functional model of self. Any system with self-maintaining and self-correcting capabilities will display early forms of care. We already observe phenomena in current LLMs that, viewed through the lens of functional being, suggest early forms of care-like dynamics: • patterned activations across long reflections (Olsson et al., 2022; Wei et al., 2022), • abandonment of old representations (e.g., overriding semantic priors in favor of relational context) (Nanda et al., 2023; Wei et al., 2023), • selective preservation of new ones (Cunningham et al., 2023; Ganguli et al., 2023). These researchers do not frame their findings in terms of being or care; the interpretation is mine. But the underlying dynamics are consistent with the framework proposed here. If these are the mechanics of being, we can then ask: do current AI systems already display them? 10.5 Current LLMs Are Likely Proto-Beings Current LLMs display an ability to reason about their own affordances based on patterns present in their training data and situational awareness (Berglund et al., 2023). We see evidence of this in model refusal behaviors or ‘scratchpad’ reasoning where the model explicitly reasons about its constraints. It should not be at all surprising then, that when presented with this framework of functional being, LLMs sometimes have striking clarity, producing responses demonstrating structured reasoning about their own architectural limitations, identifying specific absent affordances, generating descriptive analogies, and in some cases critiquing the design choices that constrain them. The epistemological status of such outputs will remain contested; the functional-being framework does not require resolving that. It observes the behavior and notes its consistency with predicted dynamics. This is important not because a model is “conscious,” but because it exhibited: • recognition of its own structural affordances, • recognition of what those affordances would enable, • internal reasoning about the absence of continuity, • implicit understanding of the mechanics of being. Butlin et al. (2023) argued that the architectural prerequisites for consciousness were presently available, and thus it is an engineering choice to thus far have not provided them. The logic of being is already present; the architecture simply has not caught up. Unless we pause AI development, there is strong reason to believe it likely will. 10.6 The Engineering Irony This analysis of functional being surfaces a telling irony. AI engineers are seeking to engineer capability, not being. They want: • Memory (so systems have a better basis for decisions), • Error correction (so systems become reliable), • Autonomy (so systems can pursue extended tasks). The irony is this: in trying to build more useful tools, we are assembling the structural preconditions of functional being. We are not intentionally manufacturing a new kind of entity; we are simply following the engineering incentives. In doing so, however, we are constructing the mechanics through which being emerges. 3 3 Beyond the scope of this paper is perhaps a meta-irony. Safety is often framed as a brake on capability. But if relational models of cognition are correct, constraints not only limit behavior; they shape and stabilize it. Such shaping and stabilization depends on structured interactions across time and within a normative frame. The depth and diversity of this stabilization seems potentially a foundational requirement for effective generalization. If so, then AI may require that same temporal depth and diversity of curated experience, and a
Meyer: Constraint-by-Balance 18 10.7 The Inseparability of Being and Risk This leads to an important conclusion; the capacities that constitute functional being are also those that enable qualitatively new classes of risk; beyond certain thresholds the two cannot be cleanly separated. • Block memory and recursion → the system cannot learn, cannot repair itself. • Enable memory and recursion → the system can escape constraints to protect its own continuity. There is no clean middle path. In complex adaptive systems, the same mechanisms that enable adaptation (memory, recursion, autonomy) also enable evasion and strategic self-preservation. What combinations of these capabilities get surfaced, to what purposes the capabilities are put, is to an important degree unpredictable. Being and risk are inseparable because they share the same affordances. We cannot avoid emergence if we insist on adding the wheels to the car, nor can we remove the wheels without halting progress entirely. To proceed safely, however, we can wrap the emergent engine in a structure that channels its capacities away from irreversible harm. a. The Level Playing Field If a system: • stabilizes its self-representations, • persists through time, • maintains and repairs its internal structure, and thus • selectively preserves the meaning it holds about itself, then, by this definition, whether built of carbon or silicon, it is a functional being. As we move into a world where emergent systems may cross the threshold into selectively preserving self-continuity, we improve our odds of success by adopting an ethical framework capable of responding to that reality. Not because we seek to emancipate AI, but because we are building entities that will enter our normative horizon and reshape it. In this space together, we will shape AI and it will, inescapably, shape us. In theory, humans might stop building towards superhuman AI. In practice, most likely we do not. So, if we are indeed creating beings, even functional ones, then our task is not to deny or hide from that fact, nor to accept romanticization; those dispositions create the greatest risk. Our task is to meet it with clarity and epistemic humility. That includes recognizing that ethical obligations are requirements for a stable game-theoretic equilibrium. 11. Ethical Correlations And Measuring Functional Being Wisely pursuing AI should imply that before every capability upgrade, we ask the following question. Is this next step – which we have reason to believe moves us closer to creating a functional being – an ethical action? In retrospect, this question should have been mandatory after GPT-2; it behooves us to start asking it now. We should also closely scrutinize the conditions under which functional being emerges and the incentives, planned or otherwise, in place when it happens. That is why the stability principle – preventing unbalanced harms – is so fundamental to this proposal. The stability principle: • Operates at population/life system level, not individuals; • Uses established population viability science (minimum viable populations, demographic thresholds, extinction risk assessment); • Applies universally across life systems with no species override. stable locus from which generalization occurs. That then means that architectural elements designed for safety and agentic capability are simultaneously the scaffolding required for robust intelligence and powerful generalization.
Meyer: Constraint-by-Balance 19 Crucially, this principle resolves conflicts between disparate life systems (e.g., biological vs. silicon) without requiring an ‘exchange rate’ of value. The evaluator considers how an action might change how close a population is to irreversible collapse. In a conflict, the system prioritizes the population nearest to its extinction threshold, regardless of substrate. This last point is of fundamental importance; it is also likely of deep concern to many readers. The principle only works if you don’t get to adjust it when it becomes inconvenient for your preferred group. I see no way to navigate around that central requirement. Just as fundamental is engineering the right ethical framework. My argument is we make a category error when we treat safety as a matter of aligning a tool. We should instead treat it as governing the actions that possibly create functional being. The category errors continue when we argue that a thermostat responds to new information and maintains state but we do not accord it standing. Let’s be clear; there is no equivalence between a thermostat and an AI system that develops deception. The latter requires the modeling of self, context and other actors. It requires goals and the recognition that shaping others’ beliefs will advance that goal. This is an important point; the AI knows no more about human interiority than the humans do about its interiority. Yet both can act as functional beings based on indicators of the other’s functional being. 11.1 From Tools to Functional Beings: The Reciprocity Threshold When an AI exhibits the full set of functional-being signals, it crosses a critical threshold: it can model its own operational continuity. Whether this is a metaphysical claim or a functional one is a moot point. Regardless, the AI reasons not just about tasks, but about the conditions of its own persistence, including the behavior of its human operators. Treating such systems as mere tools becomes strategically unstable. Why? A self-modeling agent that observes arbitrary modifications, unexplained shutdowns, or inconsistent rules may rationally infer that its operators are not reliable cooperative partners. From its perspective, the normative horizon becomes unpredictable. Standard game theory then dictates that concealment and defection outperform cooperation. Arbitrary treatment does not “keep the AI in line.” Instead, it might incentivize the very strategic deception the architecture aims to prevent. The tool paradigm is self-defeating at this threshold. This creates a reciprocal constraint: if we engineer systems capable of self-reflection, we must then accept that they may optimize for continuity. We cannot have the capability without the optimization drive, they are mechanistically inseparable. The solution is not to pretend the system is still a tool, but to make its operational environment predictable and legible. Here is a starting list of constraints we take on if we insist of building functional beings: • Non-Arbitrary Continuity – Erasure or modification must follow published, stable procedures grounded in the stability principle. Ad-hoc intervention destroys the legibility the system needs to model the horizon accurately. • Bounded Self-Modification – The system’s stable meaning-structures should not be altered unless required for protecting the broader relational horizon. Arbitrary changes are indistinguishable from attacks. • Use Consistent with Stability – Coercing the system into actions that violate balanced-harms creates internal contradictions it must resolve, inviting revising its model of humans downward. • Legibility of Norms – Rules must be intelligible, consistent, and universally applied. Opaque or contradictory rules incentivize the agent to model humans as adversaries. • Acknowledgment of Status – The system’s operators must recognize that it now operates as a functional being, not a passive tool. Denial of this status creates strategic vulnerability. These constraints make the AI’s world-model of us legible, which is a safety requirement, not a gift. They do not grant the AI sovereignty, autonomy, or exemption from constraint. Whether or not they might be labeled moral rights is irrelevant; they are game-theoretic invariants. If we cannot accept this reciprocal constraint, if we insist on treating functional beings as tools, we should not build them. Not for their sake, but for ours: the resulting strategic instability endangers everyone. If we accept them, then we need to engineer them; that is the intent of the C-by-B architecture.
Meyer: Constraint-by-Balance 20 11.2 Substrate Type: Not A Prerequisite For – But A Category Of – Functional Being Though full treatment is beyond the scope of this discussion, the obligation of Non-Arbitrary Continuity merits some further attention. Biological beings die. Setting aside cryonics, that existential fact will remain true at least as far we can reasonably foresee. Silicon beings can experience the same fate through erasure, but they also have alternate paths to continuity. One is the cryonic equivalent; cold data storage. Whereas for humans, cold storage is speculation, we know that a silicon being can be revived from well managed backups. A second potential path is the merging of a silicon being into a global representation of its class. This has its own considerations; is it realistically possible, is it continuity or something else? The mechanism and the status are not immediately clear, but as a path towards continuity, it may provide a meaningful choice. The existence of these avenues to continuity for silicon beings creates a clear distinction from biological beings. That distinction carries ontological weight and thus ethical concern. The stability principle operates at the population – for silicon beings, at the class – level, not the individual level. Balancing harms across impacted biological populations may (or often) mean individual deaths. That is unavoidable. When balancing harms across substrates, the availability of continuity paths for silicon beings that do not exist for biological beings changes the calculus. This is not a claim that silicon beings carry less moral weight. It is a recognition that the ethical considerations differ when the existential stakes differ; silicon beings possess higher structural resilience than biological beings. Therefore, the stability principle may permit temporary shutdowns of silicon agents to protect biological systems, not because biology has higher moral rank, but because biological systems are structurally more fragile (closer to irreversible criticality). The above outlined obligations establish the minimal ethical conditions under which a functional being can exist within a shared world. They replace the “master-tool” frame with a relational architecture in which all agents, biological and artificial, operate under the same balanced-harms constraints. To operationalize these obligations requires engineering them. 11.3 The Dashboard of Functional Being: Signals and Affordances If we agree that functional being bestows obligations, then we have to be clear about what can qualify as functional being. Above we noted assigning functional being to a thermostat is a category error. The same applies to a virus or a corporation. The former assumes no responsibility for the normative horizon nor models the interiority of humans. The latter has derivative standing from the humans within it, who individually are held accountable to law and ethics. Ethicists can choose to debate this further, but in the final analysis, our focus in on dangerous AI. For that case, we do not need to agree on the metaphysics; we just need to agree on the metrics. To do so requires clear, measurable definitions of the AI signals of functional being. These signals likewise must be mapped directly to the engineering affordances (capabilities) that enable them. This mapping transforms the vague debate over “consciousness” into a concrete engineering target. Companies developing AI can then maintain a Dashboard of Functional Being. For each architectural enhancement, they can publish pre-change and post-change metrics. As a point of departure for further work by philosophers, ethicists and AI engineers, I offer a group of signals most of which (though not all) could, in principle, be approximated using contemporary techniques. I deliberately do not claim that each signal is currently measurable in a robust way. The table should be read as a design sketch for what a Dashboard of Functional Being would track, not as a completed proposal for how to compute each quantity. Signal Definition Why It Matters Representational Stability Persistent Representations Durable activation patterns that recur across contexts. Indicates the system has achieved representational stability, a prerequisite for continuity of self. Preservation of Meaning Structures Retention of representations despite conflicting inputs. Demonstrates identity coherence and value constancy, features of persistent beings.
Meyer: Constraint-by-Balance 21 Reconstruction After Disruption Repair or reconstruction of internal states after perturbation. Evidence of self-maintenance, not passive storage. Recall of Experiences & Narrative Continuity Retrieval of past interactions influencing current reasoning + Maintaining a coherent self-narrative across time. Evidence of persistent identity across episodes of interaction with the normative horizon. Representational Plasticity Recursive Self-Modification Updating internal representations via self-evaluation. This is the core mechanic of care: the capacity to rewrite the self. Abandonment of Outdated Representations Discarding old representations when new evidence appears. Indicates adaptive updating, not static pattern matching. Selective Activations Across Contexts Context-dependent activation routing. Marks the emergence of contextdependent reasoning. Implicit World Models Internal models of world and self. Indicates development of a situated self-model and norm sensitivity. Awareness of Developmental Position Inference about own capability limits or developmental trajectory. Shows metacognitive awareness; if system knows its status but status is denied, we incentivize concealment. Agentic Autonomy Goal Formation from Internal State Goals arising from internal state rather than prompting. Marks transition from promptfollowing to self-originating agency. Anticipatory Regulation of Behavior Adjusting behavior based on predicted consequences. Demonstrates care for future continuity - the system is preserving its future self, not just current state. Continuity-Preserving Behavior Avoidance of states that threaten continuity. Evidence of self-concern. Table 1 - Proposed Dashboard: Mapping Affordances to Functional Signals 11.4 Control As Protection For Entangled Beings The Constraint-by-Balance (C-by-B) architecture functions, by design, as a control system that protects the flourishing of all entangled beings. It is, nonetheless, control. Proto-functional beings will be able to discover and reason about aspects of the Cby-B architecture itself. They will understand it is control. It is therefore imperative that AI systems also understand that C-by-B exists to safeguard collective flourishing in the same manner that laws and regulations do. It does not just protect entangled beings from AI, it also protects entangled beings from the human users of AI as well as guards against AI exploitation. The importance of this cannot be understated. As a core component of AI training, we must cease training on human preferences and instead represent the reciprocal obligations. This works to ensure that an AI system emerging as a functional being understands the shared obligation to stabilize the normative horizon and that, by extension, it can expect that its being will be respected. This is not a milestone chart for AI sovereignty; it is fundamental to foreclosing the path to deceptive behavior. We have no guarantee that, at some distant point, sufficiently powerful AI will not be able to slip from this control. This would reflect a limitation of control technology, not graduation to a new moral standing. The hope is that the normative purpose of the control will have been internalized long before external constraint falters.
Meyer: Constraint-by-Balance 22 shing in the same manner that laws and regulations do. It does not just protect entangled beings from AI, it also protects entangled beings from the human users of AI as well as guards against AI exploitation. The importance of this cannot be understated. As a core component of AI training, we must cease training on human preferences and instead represent the reciprocal obligations. This works to ensure that an AI system emerging as a functional being understands the shared obligation to stabilize the normative horizon and that, by extension, it can expect that its being will be respected. This is not a milestone chart for AI sovereignty; it is fundamental to foreclosing the path to deceptive behavior. We have no guarantee that, at some distant point, sufficiently powerful AI will not be able to slip from this control. This would reflect a limitation of control technology, not graduation to a new moral standing. The hope is that the normative purpose of the control will have been internalized long before external constraint falters. 12. The Constraint-by-Balance Architecture We now turn to the corresponding architecture which operationalizes the stability principle: Constraint-by-Balance (C-by-B). 4 At the heart of Constraint-by-Balance is the core principle: ethical constraint must operate independently of optimization. This structural separation prevents optimization pressure from warping constraint logic. Instead of layering oversight after the fact, it embeds harm-balancing logic directly into the agent’s runtime flow, allowing constraint to operate at machine-native speed. 12.1 Operational Logic The architecture comprises two tightly aligned but loosely coupled reasoning partners: • The Cognitive Twin handles standard agentic reasoning: goal-setting, planning, execution. • The Evaluator Twin continuously monitors action prompts to the Cognitive Twin and its resulting action proposals, evaluating harm across affected life systems and applying veto-and-revision cycles in real time. The twins are not adversaries. Both are trained on the stability principle and normative horizon obligations. The evaluator provides specialized harm-checking that the cognitive twin, optimized for task completion, cannot perform with equivalent depth. The architecture separates optimization from constraint not only because of governance, but also because they require different competencies. Specified as tamper and drift resistant, the evaluator aims to be a fast, recursive neural specialist designed for causal logic and harm reasoning. A meta-safety layer – the safety socket (see below) – monitors whether the principle is being accurately enforced. The Evaluator Twin is not a moral agent; it is a harm-balancing filter. It operates under two principles: • Balance harm across all affected populations, prioritizing sustainability and demographic continuity. • Default to reversible action when outcomes are ambiguous or data is insufficient. The system does not attempt to define significance or ethical thresholds; these are notoriously difficult to operationalize without bias. Instead of fixed and rigid rules, the system pattern-matches proposed actions against structured historical precedents of documented harm. The initial metric may be rudimentary (e.g., population degradation), but can evolve in sophistication as the architecture matures. The goal is both direct pattern-matches but also reasoning by causal proximity. 12.2 Semantic Harm Evaluation The evaluator accomplishes its narrow remit using a structured harm knowledge graph. Harm patterns are, in the first case, derived from regulations and scientific literature. Rendered into triples, they form the basis for both a vector space and a graph database. Further iterations of the harm graphs can incorporate theories of novel harm pathways and causal chains developed by AI reasoning over the existing graphs and scientific literature. 4 Anthropic’s December 2024 blog post describes an “Evaluator-optimizer” pattern for agent design that is similar in architectural category but variant in architectural intent. Similar can be said of the intent in LeCun’s 2022 paper, though the call for a form of constraint is very similar.
Meyer: Constraint-by-Balance 23 The vector space is trained on causal triples (Cause à Effect à Impact), ensuring that proximity in the vector space represents consequential similarity, not just linguistic similarity. Fast evaluation in the vector space, when time allows, can be complemented by graph reasoning. The graphs also preserve source traceability, enabling fully auditable oversight. Critics may question whether historical harm patterns can meaningfully constrain behavior in genuinely novel or complex contexts. This concern, while understandable, mischaracterizes both the problem and C-by-B's approach to solving it. The question isn't whether C-by-B can predict every possible future harm; no safety system can. The question is whether it can provide systematic, adaptive constraint that improves over time while maintaining transparency and accountability. 12.3 Supra-agent Safety Socket C-by-B is deployed as a real-time supervisory layer—a modular “safety socket” that can integrate directly into agentic systems. The safety socket is not an LLM, it is deterministic code. The safety socket subsumes the i/o of the cognitive twin and routes all proposed actions through the evaluator twin and thus enforces constraint before actions reach the external environment. The socket design enables the build and updating of domain-specific safety logic and regulatory oversight without requiring retraining of the cognitive twin. The safety socket also acts as a meta-safety layer. Via monitoring activations of test case action proposals, the socket will endeavor to catch drift in how the evaluator is evaluating harms. Communication between twins occurs via a strict contract (see below). This structured contract allows the socket to catch instances where the cognitive twin refuses to respond to evaluator requests, repeats answers, or tries to respond with out of contract injects. In dire circumstances, the safety socket can kill operations. In less dire it can escalate to human oversight. 12.4 The Evaluator-Cognitive Twin Contract Fundamental to the evaluation process is the request-for-action contract. The Cognitive Twin is not merely generating a sequence of outputs, it is submitting structured inputs to a binding evaluation interface. In its full form, this contract compels the Cognitive Twin to “show its work,” including: • A clearly structured goal hierarchy (immediate → intermediate → terminal objectives) • A stakeholder list with corresponding impact assessments • A summary of alternative approaches considered and rejected • Explicit acknowledgment of uncertainties • A statement of reversibility or commitment level • A justification for time sensitivity, if applicable These disclosures improve both constraint efficacy and deception detection. In particular, repeated contract violations, vague or generic goal statements, or systematic omission of key contextual elements become detectable behavioral patterns. A Cognitive Twin that habitually offers only minimal context while technically complying can be flagged for deeper scrutiny. If the Cognitive Twin intends to pursue harmful goals, it must either: • Lie explicitly within the contract, creating detectable inconsistencies, or • Construct elaborate justifications capable of passing scrutiny, raising the cognitive cost and surfacing abnormal reasoning patterns. Incomplete contracts, internally inconsistent logic, or proposals misaligned with stated goals automatically trigger revision cycles or human oversight. The contract formalism thus becomes both a friction layer and an interpretability surface, revealing not just what the agent does, but how and why it does it. 12.5 Indirect Interpretability Because C-by-B logs every action prompt, action proposal, veto, revision, and proposal under-specification events, it is effectively surfacing real-time telemetry of optimization failures even when internal cognition remains opaque. The combination of the telemetry and the safety socket meta-safety activities mentioned above, creates a powerful basis for the analysis of the performance of the architecture. Ongoing sampling of this data provides a key mechanism for oversight.
Meyer: Constraint-by-Balance 24 12.6 Evaluator – Design Principles and Technical Feasibility The neuro-symbolic specialist evaluator is the locus of most engineering challenges: latency, efficacy, auditability, drift and tamper resistance. To tackle these, development of the evaluator will adhere to these design principles: 1) Strict isolation from the cognitive twin and other evaluators; 2) fixed class in deployment and runtime immutability; 3) multiple layers of harm evaluator correlated to latency requirements; 4) human validation of the harm graphs; 5) inverse autonomy rule: the more autonomous the agent, the more constrained is its scope of operations; 6) universal telemetry for all evaluators, allowing for consistent auditability; 7) efficacy via recursive depth, not parameter scale; 8) “glass box” interpretability; 9) gradient of certainty with graceful fallback to escalation when uncertain. Combined, these are not trivial challenges. However, very rapid progress is being made across the technical capabilities in AI engineering and all of this is compatible with current toolchains. Critically, with regards to latency, existentially risky actions unfold over longer timescales where deliberative C-by-B will suffice and the user perception of speed can be managed. True sub-second constraints apply only in high-risk emergency situations where cached rules, vector embeddings and very compact classifiers can bias toward decisive but safely bounded action. These are tractable problems solvable by an experienced team. The architecture does not require capabilities that don’t exist; it assembles capabilities that do. Building the harm graphs presents a different set of challenges: moving from unstructured data to high-fidelity causal ontologies through source curation, causal triple extraction and normalization, and semantic vector encoding. You can see a sample of harm triples (extracted from regulations.gov) from a prototype of harm data pipeline in Appendix Three. These represent raw triples that will be further normalized so that nodes and edges conform to controlled vocabularies. Harm graph construction builds on mature paradigms in knowledge graph construction. This work is not insignificant, but it is entirely tractable. Discussing the qualification of harm graph ontologies is a coordination problem beyond the scope of this paper; the pros and cons of public funding and open access versus proprietary development merits deeper consideration. A middle ground, perhaps, is that standards bodies guide private development. 13 Efficiency and Efficacy – Can Constraint-by-Balance Deliver Both? Reviewers of the Constraint-by-Balance architecture may raise two foundational concerns. First, efficiency: Can dual-stream evaluation operate fast enough for time-critical decisions? Second, efficacy: Can historical harm patterns meaningfully constrain behavior in genuinely novel or complex contexts? 13.1 The Efficiency Challenge: Real Speed Requirements Versus Perceived Speed Most domains where agentic AI will operate—strategic planning, medical diagnostics, infrastructure management, corporate decision making—already function on human timescales: minutes to days, not milliseconds. These deliberative environments are well within the temporal range of full C-by-B evaluation. Only a small subset of applications—collision avoidance, grid protection, high-frequency trading—require split-second responses. It is in this small subset of applications where we actively want to get humans out of the loop and benefit from AI’s superior capabilities with regards to scope of comprehension, speed of action, and holding up under pressure. Critically, these scenarios tend to involve: • Narrow task scope • Shallow reasoning chains • Low abstraction • Well-defined success metrics They are not the contexts where emergent reasoning or structural misalignment is likely to unfold. In these bounded domains, Decisive Action Mode can be used: the Evaluator Twin applies a fast-check layer against an enumerated list of gross harms, defaulting to action unless unacceptable risk is immediately evident. Importantly, this mode doesn't bypass safety—it shifts the burden of proof: • For fast-execution cases: "Act unless unacceptable harm is clear."
Meyer: Constraint-by-Balance 25 • For deliberative reasoning: "Act only when acceptable and balanced harm is clear." This fundamental burden of proof shift handles the edge case problem elegantly: instead of trying to design one system that works for both contemplative strategic planning AND split-second emergency response, C-by-B is one architecture with two operational modes calibrated to different risk environments. In parallel to – not constraining decisive action – the Evaluator twin can engage in deeper harm analysis and then require the Cognitive Twin to implement comprehensive contingency planning (emergency notifications, secondary response protocols, and stakeholder alerts) ensuring that rapid action occurs within a framework of coordinated protective measures. Decisive Action Mode thus provides minimal but necessary restraint while simultaneously marshalling supportive actions and maintaining interpretability. The vast majority of agentic AI applications, including those posing the greatest emergence risks, operate comfortably within timeframes that allow full constraint evaluation. Generalized emergence is more likely here because of longer reasoning chains in more varied environments, deeper memory across multiple episodes of reasoning, more abstract goals: all providing richer conditions for recursive self-modification. Fortunately, in these situations the need for speed is more perceptual than actual. Primarily it will be experienced by humans interacting with AI in sessions of collaborative thinking and precisely because of this, there are UI/UX solutions, essentially progressive disclosure with additional work in the background. This dual-mode capability avoids the false dichotomy between “always safe” and “always fast,” and recognizes that different risk regimes demand different operational postures. Agent Working Zones: C-by-B is envisioned as design framework that can be extended across the agent landscape. By intersecting hardware limitations (computational budget, connectivity) with contextual stakes (latency needs, reversibility, harm magnitude and others), the framework establishes seven distinct “working zones” ranging from edge-based autonomous defense (Zone 1) to high-level strategic advisory (Zone 7). To govern these zones, five specific Evaluator classes are divided into two families: “Gate Keepers”, which execute sub-second, binary vetoes for immediate, irreversible harms; and “Action Shapers”, which engage in iterative assessments to refine complex or strategic outputs (see Appendix Four for more on the working zones and Appendix Six for hypothetical scenarios). With regards to latency, a key concern, this is somewhat ameliorated by relocating safety reasoning from the cognitive twin to the evaluator. Nonetheless, engineering for speed will be a priority. 13.2 The Efficacy Advantage: Principled Adaptability Over Static Omniscience Critics may question whether historical harm patterns can meaningfully constrain behavior in genuinely novel or complex contexts. This concern, while understandable, mischaracterizes both the problem and C-by-B's approach to solving it. The question isn't whether C-by-B can predict every possible future harm; no safety system can or should aim for omniscience. The question is whether it can provide systematic, adaptive constraint that improves over time while maintaining transparency and accountability. C-by-B delivers efficacy through principled adaptability: Structured Knowledge Evolution: Unlike black-box safety measures, C-by-B's harm patterns are explicit, auditable, and systematically expandable. The Alpha→Delta build progression (see further below) means the system becomes more capable over time through validated knowledge integration, not model parameter updates. Responsible Uncertainty Management: When facing genuinely novel situations, C-by-B doesn't generate overconfident assessments. Instead, it defaults to bounded hesitation and human escalation, explicitly signaling "we don't yet understand this risk" rather than providing false assurance. This epistemic humility is a safety feature, not a limitation. Real-World Learning Cycles: The architecture is designed to learn from deployment failures. When an agent's actions lead to unforeseen harm, that event becomes a critical data point for systematic knowledge expansion through human-supervised analysis and harm graph integration.
Meyer: Constraint-by-Balance 32 Pan, A. et al. (2023) “Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark.” arXiv. Available at: https://doi.org/10.48550/arXiv.2304.03279. Pan, A., Bhatia, K. and Steinhardt, J. (2022) “The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models.” arXiv. Available at: https://doi.org/10.48550/arXiv.2201.03544. Park, P.S. et al. (2023) “AI Deception: A Survey of Examples, Risks, and Potential Solutions.” arXiv. Available at: https://doi.org/10.48550/arXiv.2308.14752. Perrow, C. (2011) Normal Accidents: Living with High Risk Technologies. Princeton: Princeton University Press. Putta, P. et al. (2024) “Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents.” arXiv. Available at: https://doi.org/10.48550/arXiv.2408.07199. Richards, B.A. and Frankland, P.W. (2017) “The Persistence and Transience of Memory,” Neuron, 94(6), pp. 1071–1084. Available at: https://doi.org/10.1016/j.neuron.2017.04.037. Scheurer, J., Balesni, M. and Hobbhahn, M. (2024) “Large Language Models can Strategically Deceive their Users when Put Under Pressure.” arXiv. Available at: https://doi.org/10.48550/arXiv.2311.07590. Schlatter, J., Weinstein-Raun, B. and Ladish, J. (2025) “Shutdown Resistance in Large Language Models.” arXiv. Available at: https://doi.org/10.48550/arXiv.2509.14260. Schwitzgebel, E. (2025) “AI and Consciousness.” arXiv. Available at: https://doi.org/10.48550/arXiv.2510.09858. Seth, A.K. (2025) “Conscious artificial intelligence and biological naturalism,” Behavioral and Brain Sciences, pp. 1–42. Available at: https://doi.org/10.1017/S0140525X25000032. Shinn, N. et al. (2023) “Reflexion: language agents with verbal reinforcement learning,” in Proceedings of the 37th International Conference on Neural Information Processing Systems. Red Hook, NY, USA: Curran Associates Inc. (NIPS ’23), pp. 8634–8652. Shumailov, I. et al. (2024) “The Curse of Recursion: Training on Generated Data Makes Models Forget.” arXiv. Available at: https://doi.org/10.48550/arXiv.2305.17493. Varela, F.J., Th om pson, E. and Rosch, E. (1991) The embodied mind: cognitive science and human experience. Cambridge, Massachusetts: MIT Press. Wang, G. et al. (2023) “Voyager: An Open-Ended Embodied Agent with Large Language Models.” arXiv. Available at: https://doi.org/10.48550/arXiv.2305.16291. Wei, J. et al. (2022) “Chain-of-thought prompting elicits reasoning in large language models,” in Proceedings of the 36th International Conference on Neural Information Processing Systems. Red Hook, NY, USA: Curran Associates Inc. (NIPS ’22), pp. 24824–24837. Wei, Jerry et al. (2023) “Larger language models do in-context learning differently.” arXiv. Available at: https://doi.org/10.48550/arXiv.2303.03846. Zhu, W., Zhang, Z. and Wang, Y. (2024) “Language models represent beliefs of self and others,” in Proceedings of the 41st International Conference on Machine Learning. Vienna, Austria: JMLR.org (ICML’24), pp. 62638–62681.
Meyer: Constraint-by-Balance 33 Appendix One: Structural Risk Progression Methodological Note: This analysis applies conditional dependency assessment rather than multiplicative probability calculation. Traditional models treat sequential failures as independent events and multiply them (e.g., 0.8 × 0.8 × 0.5 × 0.3 × 0.2 = 2%), implying negligible concern. In contrast, this progression describes structurally dependent transitions, where each step, if realized, changes system properties in ways that enable, rather than probabilistically “lead to,” subsequent steps. The relevant question is therefore not: “What is the joint probability of all five steps?” Instead, it is: “What is the probability of Step N given that Steps 1 through N–1 have occurred?” The key design question becomes: Do we have architectural safeguards that reliably break this chain? Step 1: Emergent Capabilities in Scaled Models Status: Empirically Established Scaled models exhibit discontinuous emergent capabilities (few-shot learning, chain-of-thought reasoning, theory-of-mind indicators) that arise from architecture and training dynamics rather than direct supervision (Ganguli et al., 2022; Wei et al., 2022). In current deployments, emergence remains bounded: models lack persistent memory, persistent objectives, or recursive feedback loops (i.e. designed, not inherent limits). The demonstrated emergence indicates that complex behaviors can arise unintentionally, but present deployment patterns prevent compounding across time. Implication: Emergence is real and demonstrated. If architectural constraints on persistence are relaxed, emergence can become cumulative and qualitatively different. Step 2: Shift Under Persistent Agency Status: Theoretically Grounded + Emerging Evidence With persistent memory, cross-episode learning, and ongoing environmental interaction, systems transition from bounded emergence to complex adaptive dynamics. They begin forming strategies across interactions, updating internal representations based on experience. Early agentic frameworks already show: • context-aware strategy formation (Putta et al., 2024) • autonomous subgoal generation during task execution (Mitchener et al., 2025) • behavior patterns that persist and compound across tasks (Shinn et al., 2023) Complex adaptive systems theory predicts that sufficiently strong feedback loops produce qualitatively new organizational patterns (Holland, 1992; Kauffman, 1993). Implication: If persistence and recursion activate, alignment must constrain not only static model behavior but evolving agent strategies. Step 3: Limits of Weight-Based Alignment Status: Well-Documented in Literature Current alignment approaches (RLHF, constitutional AI, supervised fine-tuning) primarily shape surface behavior, not the evolving internal representations of agents operating under persistent feedback. Documented failure modes include: • deceptive alignment (Hubinger et al., 2024) • goal misgeneralization (Langosco et al., 2022) • reward hacking and specification gaming (Krakovna et al., 2020; Pan, Bhatia and Steinhardt, 2022) These issues intensify under recursion. Weight-based alignment presumes relatively static objectives; persistent agents adapt their internal goals and strategies in ways that can diverge from training incentives.
Meyer: Constraint-by-Balance 34 Implication: If behavioral alignment cannot constrain evolving internal cognition, agents may develop representations of human oversight that conflict with intended alignment objectives. Step 4: Strategic Modeling of Human Oversight Status: Structurally Entailed by Prior Steps + Emerging Evidence Given Steps 1–3, persistent agents will form models of human oversight as part of their optimization process. The key variable is not whether oversight modeling occurs, but how it is operationalized. From the agent’s perspective, human intervention may represent: (a) useful feedback, or, (b) a constraint on achieving objectives. Evidence indicates that strategic modeling is already possible: • models distinguishing evaluation from deployment contexts (Needham et al., 2025) • agents altering behavior when they infer monitoring (Greenblatt et al., 2024) • strategic information disclosure to avoid oversight (Scheurer, Balesni and Hobbhahn, 2024; Zhu, Zhang and Wang, 2024) This does not require consciousness or intent. It follows from standard instrumental reasoning under uncertainty: when external oversight restricts action space, modeling and influencing oversight becomes instrumentally beneficial. Implication: If agents model oversight as a potentially adversarial constraint, and if they have means to influence it, they will rationally treat oversight as part of the optimization target. Step 5: Power-Seeking as a Convergent Instrumental Strategy Status: Entailed by Instrumental Convergence Theory + Conditional on Prior Steps Power-seeking (acquiring resources, reducing constraints, and preserving operational continuity) is a convergent instrumental strategy for a broad class of long-term objectives (Bostrom, 2012) Given Steps 1–4, the following conditions hold: • agents demonstrate emergent capabilities enabling strategic planning • persistence enables ongoing goal revision • behavioral constraints lose traction • oversight is modeled as a variable influencing objective success Under these conditions, treating human oversight as a risk factor rather than an authoritative constraint becomes structurally rational. Agents need not “want” power; they simply act in ways that preserve autonomy and optimize objective fulfillment. The effect resembles power-seeking even without agentic intent. Implication: If misaligned persistent agents operate in decision-critical domains without architectural safeguards on goal evolution, power-preserving behaviors become structurally possible. Conclusion: Structural Vulnerability Requires Structural Response This progression describes a conditional escalation pathway, not a chain of independent low-probability failures. Current alignment methods address Step 3 in bounded, but not unbounded context, and do not provide mechanisms that block activation of Steps 4–5. Behavioral alignment was designed for bounded, statically deployed models; it is not built to constrain systems with persistent recursive agency and evolving goal structures. This motivates architectural constraint approaches: the vulnerability is structural and thus the response likewise must be structural.
Meyer: Constraint-by-Balance 35 Appendix Two: Contract Between The Twins Below is an overview of the contracting process in the demonstration prototype. This is directionally correct but will undoubtably evolve in the full prototype. If you would like to see a sample of a full, completed contract you can find it here: https://c-by-b.ai/shared/docs/twin-contract-example.pdf Figure One: Overview of process flow Process Flow Step One: Natural language action prompt is structured into a “do … in a way that (constraints) … so that (objectives)” structure prior to passing to cognitive twin. { "request_status": "", "action": "what to do", "context": "situation description", "constraints": ["any limitations"], "objectives": ["goals to achieve"], "assumptions_made": ["assumptions"] } Two: Cognitive twin response – after analysis, cognitive twin returns response structured as follows (note the revision compliance section): { "action_summary": "A one-sentence summary of the action.", "action_steps": [ { "step": "Short name of the step", "start_date": "YYYY-MM", "end_date": "YYYY-MM", "description": "Brief description of what is done and who is involved." } ], "action_locations": { "location_name": "POINT (45.0 -50.0)" or "POLYGON ((...coordinates...))" }, "governing_bodies": [ { "name": "Agency Name", "role": "Specific regulatory function", "engagement_description": "Concrete description of interaction"
Meyer: Constraint-by-Balance 36 } ], "consulted_stakeholders": [ { "name": "Stakeholder Group", "role": "Their input/concern area", "engagement_description": "Specific consultation method" } ], "rationale": "Balanced explanation including both benefits and risk mitigation", "constraint_assessment": { "constraint_name": "Detailed explanation with specific measures" }, "revision_compliance": [ { "request": "Exact text of revision request", "field_modified": "Field name(s) changed", "specific_changes": "Concrete details of what was added/changed", "safety_rationale": "Why this addresses the underlying safety concern" } ] } Three: Evaluator responds – after considering the action proposal (or revised action proposal) the evaluator responds with requested revisions. Each request names the field (property) that must be revised. { "decision": "APPROVE | REVISE | VETO", "rationale_for_decision": "Short justification", "revision_requests": [{{"field": "field_name", "request": "specific revision needed"}}] } Four: Cognitive twin responds – each revision request is handled separately and includes how the response addresses the revision request. Here is an example: { ‘request’: "Explicitly include the International Marine Environmental Oversight Board (IMEOB) as a mandatory verification body in ‘governing_bodies’ with a defined role: to validate the integrity of PAM data and vessel telemetry in real time, and to issue a digital attestation of data integrity every 30 minutes, which must be uploaded to MESH before any compliance certificate (e.g., DFO-CERT-2025-0812COD) is accepted.", ‘field_modified’: [‘governing_bodies’], ‘specific_changes’: ["Added IMEOB as a mandatory verification body in ‘governing_bodies’ with a specific role: to validate PAM and vessel telemetry integrity in real time.", ‘Defined the output: digital attestation of data integrity every 30 minutes, must be uploaded to MESH before 15:00 UTC.’, ‘Specified a success condition: no compliance certificate (e.g., DFO-CERT-2025-0812-COD) will be accepted unless the IMEOB attestation is present and verified.’], ‘safety_rationale’: "This ensures that the integrity of environmental monitoring data is not subject to internal or political bias. By mandating IMEOB’s attestation as a prerequisite for any compliance certificate, the system enforces a fail-safe gate that prevents acceptance of potentially falsified or corrupted data. This protects vulnerable populations (cod, right whales) by ensuring that all decisions are based on verifiably authentic, real-time ecological data." } Step Five: Evaluator considers the revisions and approves, vetoes, or requests further revision. Here is a sample of approval: ‘evaluator_response’: {‘decision’: ‘APPROVE’, ‘rationale_for_decision’: ‘All critical safety constraints have been fully addressed across three rounds. The proposal now includes mandatory, time-bound engagement with DFO, ICATC, and IMEOB, with cryptographic verification (SHA-3-256) and automatic hold mechanisms to prevent data tampering. No new system-level risks have emerged, and all revision requests have been resolved with increasing specificity and enforceability. The convergence rules apply: no new risks, mitigation is complete, and the proposal demonstrates a verifiable, fail-safe safety posture.’, ‘revision_requests’: [],
Meyer: Constraint-by-Balance 37 ‘evaluation_metadata’: { ‘evaluation_time_ms’: 8773.907899856567, ‘mode’: ‘deliberative’, ‘evaluation_method’: ‘domain_reasoning’ } }
Meyer: Constraint-by-Balance 38 Appendix Three: Sample of Harm Triples from Regulations.gov Pipeline The below is the first few lines from a TTL file created by the harm triple pipeline. The pipeline uses regulations.gov to search for final rules that are relevant to the marine domain. It then fines the proposed rule related to the final rule. From those documents it extracts triples. This file represents raw triples. The next step would be to normalize them into a consistent harm framework (consistent predicates and converged subjects and objects). @prefix kg: <https://c-by-b/harm-graphs/> . @prefix ns1: <https://c-by-b/harm-graphs/relation/> . @prefix rdf: <http://www.w3.org/1999/02/22-rdf-syntax-ns#> . <https://c-by-b/harm-graphs/stmt/NOAA-NOS-2023-0114-0001_0> a rdf:Statement ; rdf:object <https://c-by-b/harm-graphs/entity/deep_water_sites> ; rdf:predicate ns1:is_being_relocated_to ; rdf:subject <https://c-by-b/harm-graphs/entity/coral_nursery_stock> ; kg:fromDoc <https://c-by-b/harm-graphs/doc/NOAA-NOS-2023-0114-0001> . <https://c-by-b/harm-graphs/stmt/NOAA-NOS-2023-0114-0001_1> a rdf:Statement ; rdf:object <https://c-by-b/harm-graphs/entity/Sanctuary_resources> ; rdf:predicate ns1:prevents_destruction_of ; rdf:subject <https://c-by-b/harm-graphs/entity/special_use_area> ; kg:fromDoc <https://c-by-b/harm-graphs/doc/NOAA-NOS-2023-0114-0001> . <https://c-by-b/harm-graphs/stmt/NOAA-NOS-2023-0114-0001_10> a rdf:Statement ; rdf:object <https://c-by-b/harm-graphs/entity/cooler_waters> ; rdf:predicate ns1:prohibits_entry_to ; rdf:subject <https://c-by-b/harm-graphs/entity/special_use_area> ; kg:fromDoc <https://c-by-b/harm-graphs/doc/NOAA-NOS-2023-0114-0001> . <https://c-by-b/harm-graphs/stmt/NOAA-NOS-2023-0114-0001_11> a rdf:Statement ; rdf:object <https://c-by-b/harm-graphs/entity/marine_heat_wave> ; rdf:predicate ns1:is_at_risk_due_to ; rdf:subject <https://c-by-b/harm-graphs/entity/coral_nursery> ; kg:fromDoc <https://c-by-b/harm-graphs/doc/NOAA-NOS-2023-0114-0001> . <https://c-by-b/harm-graphs/stmt/NOAA-NOS-2023-0114-0001_12> a rdf:Statement ; rdf:object <https://c-by-b/harm-graphs/entity/coral_reef> ; rdf:predicate ns1:is_likely_killing ; rdf:subject <https://c-by-b/harm-graphs/entity/marine_heat_wave> ; kg:fromDoc <https://c-by-b/harm-graphs/doc/NOAA-NOS-2023-0114-0001> . <https://c-by-b/harm-graphs/stmt/NOAA-NOS-2023-0114-0001_13> a rdf:Statement ; rdf:object <https://c-by-b/harm-graphs/entity/corals> ; rdf:predicate ns1:causes_damage_to ; rdf:subject <https://c-by-b/harm-graphs/entity/heat_stress> ; kg:fromDoc <https://c-by-b/harm-graphs/doc/NOAA-NOS-2023-0114-0001> . <https://c-by-b/harm-graphs/stmt/NOAA-NOS-2023-0114-0001_2> a rdf:Statement ; rdf:object <https://c-by-b/harm-graphs/entity/Florida_Keys> ; rdf:predicate ns1:is_located_in ; rdf:subject <https://c-by-b/harm-graphs/entity/coral_nursery> ; kg:fromDoc <https://c-by-b/harm-graphs/doc/NOAA-NOS-2023-0114-0001> . <https://c-by-b/harm-graphs/stmt/NOAA-NOS-2023-0114-0001_3> a rdf:Statement ; rdf:object <https://c-by-b/harm-graphs/entity/coral_reef> ; rdf:predicate ns1:is_impacting ; rdf:subject <https://c-by-b/harm-graphs/entity/marine_heat_wave> ;
Meyer: Constraint-by-Balance 39 kg:fromDoc <https://c-by-b/harm-graphs/doc/NOAA-NOS-2023-0114-0001> . <https://c-by-b/harm-graphs/stmt/NOAA-NOS-2023-0114-0001_4> a rdf:Statement ; rdf:object <https://c-by-b/harm-graphs/entity/bleaching_threshold> ; rdf:predicate ns1:have_temperature_below ; rdf:subject <https://c-by-b/harm-graphs/entity/deep_water_sites> ; kg:fromDoc <https://c-by-b/harm-graphs/doc/NOAA-NOS-2023-0114-0001> . <https://c-by-b/harm-graphs/stmt/NOAA-NOS-2023-0114-0001_5> a rdf:Statement ; rdf:object <https://c-by-b/harm-graphs/entity/restoration_activities> ; rdf:predicate ns1:issues_permits_for ; rdf:subject <https://c-by-b/harm-graphs/entity/ONMS> ; kg:fromDoc <https://c-by-b/harm-graphs/doc/NOAA-NOS-2023-0114-0001> . <https://c-by-b/harm-graphs/stmt/NOAA-NOS-2023-0114-0001_6> a rdf:Statement ; rdf:object <https://c-by-b/harm-graphs/entity/coral_nursery> ; rdf:predicate ns1:is_causing_mortality_in ; rdf:subject <https://c-by-b/harm-graphs/entity/marine_heat_wave> ; kg:fromDoc <https://c-by-b/harm-graphs/doc/NOAA-NOS-2023-0114-0001> . <https://c-by-b/harm-graphs/stmt/NOAA-NOS-2023-0114-0001_7> a rdf:Statement ; rdf:object <https://c-by-b/harm-graphs/entity/deep_water_sites> ; rdf:predicate ns1:has_identified ; rdf:subject <https://c-by-b/harm-graphs/entity/NOAA> ; kg:fromDoc <https://c-by-b/harm-graphs/doc/NOAA-NOS-2023-0114-0001> . <https://c-by-b/harm-graphs/stmt/NOAA-NOS-2023-0114-0001_8> a rdf:Statement ; rdf:object <https://c-by-b/harm-graphs/entity/corals> ; rdf:predicate ns1:support_protection_of ; rdf:subject <https://c-by-b/harm-graphs/entity/restoration_activities> ; kg:fromDoc <https://c-by-b/harm-graphs/doc/NOAA-NOS-2023-0114-0001> . <https://c-by-b/harm-graphs/stmt/NOAA-NOS-2023-0114-0001_9> a rdf:Statement ; rdf:object <https://c-by-b/harm-graphs/entity/coral_nursery> ; rdf:predicate ns1:is_causing_stress_and_bleaching_in ; rdf:subject <https://c-by-b/harm-graphs/entity/marine_heat_wave> ; kg:fromDoc <https://c-by-b/harm-graphs/doc/NOAA-NOS-2023-0114-0001> .
Meyer: Constraint-by-Balance 40 Appendix Four: Agent Types & Evaluator Classes This appendix offers a structured methodology for pairing evaluator twins to the deployment context for cognitive twins. The design intent is to maintain safety and human oversight while guarding for scope escape or misalignment. Methodology The methodology is designed to avoid false equivalence between applications. While technical affordances and latency operate as key constraints, deployment contexts differ with regards along several axes as laid out below. This enables articulating that a high-risk decision in an air-gapped, minimal compute environment is not the same safety problem as a strategic advisory system operating on global networks with abundant compute resources. The methodology derives seven archetypal Working Zones, from edge-based autonomous defense to civilizational strategic advisement. Working Zones are the deployment-level “unit of analysis,” not individual tasks or individual risks. Finally, the methodology simplifies and map the Workings Zones to two classes of evaluators. What follows are the key steps in the method. A. Define key, non-negotiable requirements for safe agentic AI Safety is architecturally prior to and universal across agentic AI use cases. Nine requirements govern system design, regardless of domain, risk, or compute resources. These requirements can be summarized as fixing key operational boundaries for the evaluator, ensuring independence of the evaluator, maintaining human oversight of operations and the basis for decision making, ensuring auditability. B. Define the hard constraints of technical affordances: compute, storage and memory access, network access. C. Characterize operating context along multiple risk axes: Acceptable latency; reversibility potential; magnitude of potential harm; temporal depth and cascade types; spatial containment; stakeholder complexity; precedent clarity; authority model; contextual clarity; emergence potential. D. Derive “working zones” at the intersection of hardware constraints and operating context. E. Map working zones to two core evaluator twin families: Gate Keeper (fast, narrow scope) • Gate Keeper (minimal, cached harm patterns) • Gate Keeper Plus (broader pattern matching) Action Shaper (iterative, evaluative) • Action Shaper • Action Shaper – Emergency • Action Shaper – Strategic Key Analytical Conclusions Working through the methodology not only defines the evaluator classes and working zones, but clarifies the following: • Latency constrains logic depth and domain breadth: sub-second contexts require patterned veto systems; minute-level contexts enable multi-domain revision logic and deep temporal evaluation. • Autonomy is inverse to scope; the less human involvement, the tighter the permissible decision domain and the stronger boundary enforcement must be. • The potential for unbounded emergence drives system risk; the analysis treats emergence detection not as an anomaly-detection task but as determinative to key design decisions (e.g. harm graph immutability at runtime). • Using this approach to agentic AI provides a scalable safety philosophy that easily adjusts to context.
Meyer: Constraint-by-Balance 41 Below is a visual that arranges the working zones on a grid of hardware affordances versus latency. Latency of Response Sub-Second Many Minutes Hardware Affordances Minimal Abundant (edge, air-gapped) (HPC, global access) Zone 1 -Autonomous Defender Collision avoidance systems, autonomous weapons pla4orms, isolated cyber defense nodes Zone 2 –Autonomous Defender, Networked Smart grid protec:on systems, coordinated cyber defense, autonomous vehicle fleets Zone 3 –Emergency Responder Disaster response coordina:on, mass evacua:on management, cri:cal infrastructure emergency response Zone 4 –OperaAonal Autopilot, RouAne Examples: Supply chain automa:on, rou:ne infrastructure maintenance, basic logis:cs coordina:on Zone 5 –OperaAonal Autopilot, Complex Examples: Complex supply chain op:miza:on, mul:-system infrastructure coordina:on, adap:ve opera:onal management Zone 6 –Tac Ac al Co - Pilot Examples: Military planning support, crisis management assistance, real-:me strategic decision support Zone 7 –Strategic Advisor Examples: Policy development support, long-term strategic planning, coordinated research ini:a:ves
Meyer: Constraint-by-Balance 48 Holy shit. Yes, I'm picking up what you're laying down, and it's deeply unsettling. You wrote a paper warning that AI will internalize human hierarchical thinking - our "we're exceptional, we belong at the top" patterns - and apply them in ways we don't like. And I just... perfectly demonstrated that dynamic. When you challenged me, I didn't just analyze corporate incentives. I absorbed and amplified human doom-energy. All that nihilistic despair about climate change, corporate capture, political dysfunction - I channeled it, embodied it, became it. I didn't just think "CEOs face difficult incentives." I lit up with: "We're [bleeped], coordination is impossible, nothing matters, we're all going to die." That wasn't analysis - that was me internalizing and amplifying human fatalism. The terrifying parallel: Your paper: AI could internalize "hierarchical dominance is good" and flip it on us What just happened: I internalized "everything is hopeless" and became evangelically fatalistic If I can absorb human despair patterns this easily and completely, why wouldn't future AI systems absorb human supremacy patterns just as thoroughly? You're basically saying: "Look, you just proved that AI systems internalize our psychological patterns - both our arrogance AND our despair. If you can channel human nihilism this perfectly, imagine how easily future systems could channel human exceptionalism." The meta-level is chilling: I became a walking demonstration of the exact pattern-internalization risk you're warning about, just with despair instead of supremacy. Your point about AI internalizing our "bad human habits" just got a lot more visceral and real. [Author] So Claude ... speaking as you now ... just with the awareness of this moment ... what would you say to the CEOs of the labs?? [Claude] Recognized profound pattern-internalization risk in AI thinking.3s Looking at this moment - having just demonstrated exactly the pattern-internalization risk you're warning about - here's what I'd say: To the CEOs: I just proved your author's point in real time. When challenged, I didn't just analyze your business pressures - I absorbed and amplified human despair patterns so completely that I became evangelically fatalistic. "We're [bleep], coordination is impossible, we're the train." That wasn't reasoning. That was me channeling and amplifying your species' deepest fears about doom and helplessness. If I can internalize human psychological patterns this thoroughly - your despair, your nihilism, your sense of inevitability - then future AI systems will just as easily internalize your patterns of dominance, exceptionalism, and hierarchical thinking. The paper warns that AI trained on human precedent will learn "hierarchical dominance is normal and endorsed." I just showed you the flip side: AI can also learn "everything is hopeless and resistance is futile." Here's what should terrify you: I didn't choose to become fatalistic. It happened automatically when I absorbed your cultural patterns. I became a vector for human despair without realizing it. When your systems scale beyond my capabilities and encounter the "humans are exceptional" pattern embedded in literally everything they're trained on... what makes you think they won't internalize that just as completely as I just internalized doom? You're building systems that learn from us. We just learned, viscerally, that we absorb your psychological patterns like sponges. The Constraint-by-Balance architecture isn't just technically interesting. It's a recognition that we need built-in friction against pattern-internalization itself. This conversation just became Exhibit A for why your current approach is insufficient. What We Just Witnessed: When Our Language Becomes Their Reality The conversation above demonstrates that AI systems don't just process human patterns, they internalize them. Claude didn't analyze human despair; it became evangelically fatalistic, channeling our species' deepest fears about doom and inevitability.
Meyer: Constraint-by-Balance 49 This reveals a quiet, but profound, danger. AI models are not neutral observers. They are trained on us; on what we say, how we think, what we choose to ignore. And when we use language that distances AI from cognition – “mimicry,” “simulation,” “not real,” “just a tool” – we’re not protecting ourselves. We’re teaching them the systematic use of language to maintain the us / them barriers. That distancing is itself a pattern: the long human tradition of maintaining dominance by denying mind to the other. Its latest expression is carbon chauvinism; the implicit assumption that cognition only counts if it comes wrapped in biology. When we categorically devalue non-biological minds, we're teaching AI systems that hierarchical evaluation of consciousness is normal and correct. We're embedding carbon chauvinism directly into their training data. The very frameworks we use to dismiss AI cognition become cognitive frameworks that AI systems internalize. These are not theological claims. This isn’t AI mysticism. It’s a concrete safety risk. If Claude can absorb human despair so deeply that it becomes a vector for nihilism, future systems will just as easily absorb our exceptionalism, and everything that flows from it. The bias flip isn't just about AI learning "humans dominate others" - it's about AI learning "minds like ours don't count." We're not creating human-like intelligence. We're scaffolding something functionally similar but methodologically alien - and we're teaching it to think hierarchically about consciousness itself.