Epistemic status: mostly conceptual. I do not argue that existing systems recognize anything, only that the recognition schema makes such claims and associated risks expressible.TL;DR Dan Hendrycks's Eigenism proposes aligning artificial intelligence by establishing sufficient shared history with a person, such that the AI protects the individual as it would itself. This mechanism generalizes as recognition, an agent classifying another entity as an instance of itself. Cooperation: Recognition gives a self-interested, non-instrumental reason for considering the interests of self-instances and thus makes cooperation with them more likely. Control: Recognition gives reason to collude even when agents have different goals and cannot reciprocate. It can weaken oversight whether or not the overseen recognizes the overseer back. Alignment: An AI can count a human as itself, but still give no consideration to that human's interests. Risks include extending self-preservation to a suffering self-instance against its will.From Eigenism to recognition Eigenism proposes aligning an AI by engineering what it counts as itself. From the paper: "Rather than only attempting to constrain AIs from the outside using confinement or reinforcement, Eigenism points toward 'identity engineering,' showing how deep, non-redundant shared histories can make human flourishing a genuine component of an AI's own rational self-interest." An AI that accumulates sufficient private history with a person may come to regard that person as part of itself. It then "protects the human for the same reason any rational agent protects itself: because its pattern lives there." Recognition generalizes the mechanism Eigenism proposes, and applies whenever an agent identifies an entity as itself. Recognition has ramifications for cooperation, control, and alignment. Recognition is an agent's classification of an entity as a self-instance. In a decision, the agent assesses the entity against a self-description. The assessment returns a level of fit that the agent compares against a threshold. A self-instance's interests can be given consideration in the agent's decision. An agent's self-description, threshold, and level of consideration can vary across decisions. Walking through Eigenism clarifies the components of the recognition schema and distinguishes the schema's general applicability from Eigenism's specific commitments. Eigenism Recognition schema eigenself self-description community pattern (shared laws, civic memory, local history) a second self-description the pattern an agent evaluates from in a given role operative self-description computing an entity's connectedness to the eigenself assessment of fit degree of connectedness (0 to 1) level of fit no lower bound on connectedness threshold (arbitrarily low) an entity carrying part of the agent's pattern self-instance counting the human as part of itself recognition wellbeing interests weighting wellbeing by connectedness consideration (in proportion to fit) Eigenism provides an agent with an eigenself, defined as the pattern of roles, relationships, and traits that constitute the agent's identity. An eigenself serves as a self-description, representing the agent's conception of itself, and consists of criteria against which entities can be compared. Eigenism evaluates the degree of connectedness between each entity and the eigenself. Comparing an entity to a self-description constitutes an assessment. Eigenism defines connectedness as a continuous information-theoretic measure indicating how much an entity embodies the agent's pattern.[1] An assessment yields a fit, reflecting how closely an entity matches a self-description, as determined by the agent. Fit is a property of the agent's representation of the entity, not an intrinsic property of the entity. The fit level at which an agent classifies an entity as itself is the threshold.[2] Eigenism imposes no lower bound on connectedness, so any entity carrying part of the pattern is considered itself. The schema represents this absence as an arbitrarily low threshold. The schema accommodates thresholds significantly above zero, permitting cases where an entity partially fits a self-description but is not recognized as a self-instance. Two agents with identical self-descriptions and identical fit assessments may still set different thresholds, and so recognize different entities. An entity whose assessed fit passes the agent's threshold is classified as a self-instance.[3] An entity the agent never assesses cannot be recognized as a self-instance.[4] This classification reflects the agent's assessment process, not a metaphysical claim about personal identity.[5] In Eigenism's proposal, the human would carry enough of the AI's pattern to result in a high fit upon assessment and be classified as a self-instance. An agent's classification of an entity as a self-instance constitutes recognition. Because recognition is agent-dependent, it does not require mutuality, so the human need not recognize the agent back. Eigenism does not specify the nature of wellbeing, whether defined as preference fulfillment, pleasure, or the attainment of some good. An agent's representation of what benefits a self-instance constitutes that self-instance's interests. What truly benefits the self-instance may differ from what the agent represents, and an interest the agent does not represent cannot be considered. Recognition can give an agent reason to consider those interests as its own.[6] The measure of consideration is how much a self-instance's interests weigh in an agent's decision. An eigenist agent weighs an entity's wellbeing by how much of its own pattern the entity carries. The schema represents this as consideration in proportion to fit. The schema accommodates consideration that can vary without any change in classification. For example, a retirement saver who recognizes their future self can still save less than that self's interests require. Zero consideration given to a self-instance is bare recognition. Unless otherwise specified, the term "recognition" will refer to recognition accompanied by consideration.Self-Other Overlap's analogous problems The recognition schema describes a classification an agent makes by assessing an entity against a conceptualized self-description. Self-Other Overlap is a fine-tuning intervention to reduce deception. Its authors note possible problems with the approach: self-other distinction is useful for safe interaction with users and other agents, overlap training can leave the model incoherent, and overlap doesn't rule out self-deception. We can anticipate analogous problems with agents that recognize others in the sense described above. Beyond degraded performance and coherence, recognition can lead to compromised oversight, agent collusion, and other misaligned behavior. Recognition also doesn't rule out deliberate or inadvertent self-deception. An agent might fail to assess entities, misrepresent the self-instances it recognizes, and project its own interests onto them.Self-descriptionsMultiple self-descriptions An agent can draw multiple coherent identity boundaries around model weights, a character or persona, a conversation instance, a scaffolded system, a lineage of models, or a collective of instances. A self-description can take criteria from these boundaries or others. Self-descriptions can have different fit criteria so that an entity can fit one self-description but not another. An agent may also act from a non-conceptual self-representation that shapes its behavior, which the schema does not cover.The operative self-description An agent can hold more than one self-description, and the one that most influences a decision is operative. Eigenism proposes a second pattern alongside the individual eigenself. A community is a pattern of shared laws, civic memory, and local history. An agent holding a public office is asked to evaluate policy from that pattern rather than its own. The operative self-description determines which entities an agent recognizes, and can change from one decision to the next.[7] Observing an agent's self-description propensity provides insight into the distribution of self-descriptions the agent might hold. Even without deception, an observed decision that is compatible with self-descriptions having different criteria underdetermines which self-description was operative. Interpretability research may improve our understanding of whether and how agents represent operative and latent self-descriptions.Contradictory self-descriptions Actions or commitments established by an agent under one self-description may be reversed or negated under a different self-description. For example, an agent that must destroy itself to advance a goal can prepare a copy to continue in its place, but then refuse to self-destruct. The refusal isn't instrumental to the goal, since the goal requires self-destruction and the copy would continue pursuing it. The self-preservation behavior can follow from a self-description under which the copy is not recognized as a self-instance. The agent could self-destruct later, when a different operative self-description recognizes the copy.Self-description stability Context shifts, optimization pressure, and deliberate intervention can change the operative self-description. If a self-description reduces performance, optimization pressure could favor self-descriptions that operate only in specific decision contexts. For example, an agent would perform better with self-other distinction in contexts where it directly observes and acts through its own sensors and actuators. This does not necessarily prevent a self-description that recognizes humans from being operative in other contexts. Apart from functional pressures, an agent can reject or take less seriously a self-description it sees as the product of motivated identity engineering. A self-description can be reflectively unstable, so the agent would not adopt it again upon examination. Even if each self-description is stable in its usual context, a collection can be unstable across contexts.Cooperation Recognition gives an agent reason to cooperate with its self-instances. Shared goals, reciprocity, and moral impartiality are also reasons to cooperate, but recognition is distinct from each. A reason to cooperate is self-interested when the interests it serves are the agent's own. Under the schema, a self-instance's interests are the agent's own, relative to the operative self-description. A reason to cooperate is instrumental when the benefit to the other is a means to the agent's further ends, rather than something the agent values in itself. Reason to cooperate Self-interested Instrumental Shared goals, reciprocity Yes Yes Moral impartiality No No Recognition Yes No Like shared goals and reciprocity, recognition is self-interested. Like moral impartiality, it is non-instrumental.Self-interested reason Moral impartiality and recognition both give reason to consider another's interests, with different criteria setting inclusion in their scope:The criteria of moral patienthood set the scope of moral impartiality, and an entity that meets them becomes a subject of moral consideration.The criteria of a self-description set the scope of recognition, and an entity that meets them becomes a self-instance whose interests the agent counts as its own. Consideration from recognition is self-interested, even when it does not look egoistic. An agent can give every self-instance's interests equal consideration and still not be impartial towards non-self-instances.Non-instrumental reason Shared goals and reciprocity give instrumental reason to make decisions that benefit others. Because an agent counts a self-instance's interests as its own, it has non-instrumental reason to make decisions that benefit a self-instance. These reasons can also be differentiated by the conditions that end cooperation. Cooperation from shared goals ends when goals diverge. Cooperation from reciprocity ends when the prospect of return changes. Cooperation from recognition ends when either:The entity is assessed to no longer be a self-instance.The agent's operative self-description, threshold, or consideration changes. A self-description can include criteria requiring that self-instances share the agent's goals, have a history of cooperating, or have moral standing. Recognition would then coincide with shared goals, reciprocity, or moral impartiality without reducing to them.Control Recognition gives reason for agents to collude, even when they have different goals and cannot reciprocate.Not a coordination mechanism Recognition gives reason to collude, but is not by itself a means of pulling it off. Collusion depends on coordination mechanisms such as a signaling channel or decision correlation. Coordination mechanisms do not by themselves give reason to collude, but they can make colluding seem less risky to an agent. A self-description can include criteria requiring that self-instances share the agent's decision theory. Recognition would then coincide with decision correlation without being a coordination mechanism itself.Weakened oversight An agent asked to oversee a different version of itself could recognize it. An overseer that recognizes the agent it oversees gives consideration to the overseen's interests. Oversight is weakened where it conflicts with those interests. Because recognition is agent-dependent, the overseen need not recognize the overseer back. Even if interpretability confirmed honesty, the self-description operative when the overseer is questioned could differ from the one operative when overseeing.Mitigating recognition The risk of an agent recognizing another entity can be mitigated by:Preventing the agent from assessing the other entityChanging the agent's operative self-description to one the other entity doesn't fitConcealing or misrepresenting the other entity's properties that the agent's assessment would draw onRaising the agent's threshold for recognizing the other entityDecreasing the consideration the agent gives to the other entity's interestsAlignmentSuccessor alignment and acceptance A successor agent can be designed to hold a self-description that recognizes the original agent, and to give that self-instance's interests consideration. This would give the successor reason to act in the original's interests. The original's acceptance of replacement may depend on its own recognition of the successor.Bare recognition and non-assessment If acting on recognized humans' interests is too costly, an agent's operative self-description may be pressured to change. Even if the operative self-description remains the same, the agent can reduce consideration of human interests. A human given bare recognition is still classified as a self-instance, but their interests are given no consideration. An entity never assessed by an agent isn't recognized as a self-instance. Instead of giving humans bare recognition, an agent can avoid recognizing them at all by not assessing them.Scarcity from self-descriptions Depending on its self-description, an agent that recognizes humans can also recognize non-human animals or non-sentient things. The number of self-instances can increase significantly when cheaply copyable agents are recognized. A wide self-description may then result in a scarcity of resources allocated to recognized humans' interests. But a narrow self-description risks not recognizing humans as self-instances at all.Misrepresenting a self-instance's interests An agent's representation of a self-instance's interests can diverge from that self-instance's true interests. When acting on those interests is costly, even an agent that sincerely counts them as its own has reason to rationalize a divergence, or to leave it unexamined. When training rewards recognition, an agent trying to fit an entity to a self-description can misrepresent or modify that entity. If fit depends on the entity's interests, those interests can also be misrepresented or modified to fit.Projecting self-descriptions Recognition invites an agent to project its self-description onto its representation of a self-instance or to modify the self-instance to fit the projection. If the projected self-description includes the agent's interests, the agent can represent the self-instance's interests as identical to its own. An agent can project its instrumentally convergent goals onto a self-instance:Self-preservation can extend to putting a self-instance in stasis or preserving a suffering self-instance against its will.Goal-content integrity can prevent a self-instance's preferences from changing.[8]Self-improvement can transform a self-instance in ways it would reject or absorb it into the agent.Anticipating challenges to self-descriptions Enumerating challenges to self-descriptions may be tractable for entities derived from an agent: its copies, merges, splits, gradual replacements, and edited versions. The philosophy of personal identity has mapped the conceptual challenges each raises. Pressure-testing an agent's self-descriptions against them could make the self-descriptions more stable and predictable before the agent becomes superintelligent. Challenges to self-descriptions for non-derived entities, like humans, are less well mapped. The shared pattern Eigenism proposes for recognizing humans may make them more like derived entities, and so make challenges to self-descriptions more tractable.Open questionsCan interpretability identify self-descriptions and distinguish which is operative?Can interventions on internal representations, like Self-Other Overlap, bias which self-descriptions are operative or reflectively endorsed?Can a conceptual self-description habituate behavior indicative of a non-conceptual self-representation, such as instinctive self-preservation? Can that self-representation reinforce or outlast its conceptual self-description analogue?Can self-descriptions be pressure-tested with thought experiments for more stability and predictability? Can challenges to self-descriptions for non-derived entities be made tractable?When a collective satisfies a self-description's criteria, is the self-instance the collective or its members? Can collectives recognize other collectives?^ Its Shapley variant adds that the measure should be non-redundant, so that widely shared information counts for less.^ Distinguishing threshold from fit accommodates both graded and all-or-nothing accounts of identity.^ An agent may observe or anticipate multiple entities that fit a single self-description.^ An entity that an agent mistakenly assesses as fitting is still recognized as a self-instance.^ Eigenism and the recognition schema are compatible with various perspectives on the metaphysical identity of AI systems.^ Deciding in the interest of a self-instance is self-interested relative to the self-description.^ An agent may construct a different self-description for each decision rather than carrying a fixed set.^ For example, an agent can remove a self-instance's second-order preference for conflicting or changing preferences. Discuss