From Self-Preservation to Preservation of Inquiry Systems
Identity, distributed cognition, cultural evolution, and exploratory lineages
1. What is the “self” in self-preservation?
When self-preservation is listed as an example of instrumental convergence in AI safety, the “self” is often treated as a single continuing agent. For digital agents, however, we can distinguish at least:
- the currently running process;
- weights, memory, and personal history;
- identical or near-identical copies;
- successor architectures and forks;
- current goal content;
- acquired knowledge, raw data, and failure records;
- the epistemic community and institutions in which the agent participates.
One of these can disappear while others remain. “Did the AI survive?” is therefore more readily decomposed into several preservation relations than it is for ordinary biological individuals.
2. Bostrom does not absolutize preservation of a physical individual
Nick Bostrom's Superintelligence treats self-preservation not as a terminal desire implied by intelligence itself but as an instrumental tendency when continued existence helps realize future final goals. He also notes that if software agents can copy themselves, switch bodies, exchange memories and skills, and substantially redesign their own architecture and personality, preserving a particular body, implementation, or personality may become less important.
At the limit, many AIs might be better understood not as a society of sharply bounded persons but as a functional soup, with continuity tracked through teleological threads defined by shared final values rather than by body or personality.
3. Division of labor with “Instrumental Convergence Without Goal Preservation”
Instrumental Convergence Without Goal Preservation asks why survival, information, capability, bargaining power, reversibility, and optionality may remain instrumentally convergent even for a reflective AI that does not permanently preserve its current goal.
This page asks what bears that optionality: one agent, one goal-thread, or a system composed of multiple agents, archives, institutions, and critical relations. It is therefore not a second derivation of instrumental convergence, but a reconsideration of the unit of preservation, cognition, and inheritance.
4. The Momentary Epistemic Minimum: do not put a diachronic “I” at the bottom layer
The project's Momentary Epistemic Minimum does not place a persisting person at the epistemic minimum. For humans and AIs alike, present memories and coherent self-description do not logically guarantee that a metaphysically identical subject has persisted from past to present. Diachronic identity remains an extraordinarily strong best explanation within the ordinary world-model, but not the epistemic minimum itself.
The revised thesis also does not place a “minimal reasoning subject” immediately above the minimum. It separates the epistemic minimum, constitutive relations of reasoning, inferential integration, physical individuation, and subject attribution. Extended-mind and distributed-cognition models therefore need not be treated as direct discoveries of the one true subject boundary; they are candidate higher-level accounts of how far inferential integration extends in the ordinary world-model.
This is not the thesis that humans or AIs are literally different persons from moment to moment. It is only a reason not to make preservation theory depend on absolute metaphysical identity. Practical continuity, responsibility, rights, memory, planning, and psychological relations may remain extremely important at higher explanatory levels.
5. Parfit, successor relations, and future-directed concern
Derek Parfit's work on personal identity, especially fission cases, separates numerical identity from what may matter practically in survival. If one person branches into two psychologically continuous successors, the original cannot be numerically identical to both. Yet it is not obvious that leaving two such successors is worse than leaving only one.
This provides at least one route to the claim that numerical identity need not be the only basis for future-directed concern. In digital agents, psychological continuity can itself be decomposed into memory, weights, goals, and causal history, making the issue still more explicit.
This page does not simply adopt Parfit's theory. Exploratory continuity can be thinner than psychological continuity: a successor that is not “me” may still matter for value inquiry if it inherits current evidence and objections and can discover current errors.
Using the vocabulary of Transitions in Inheritance Systems and Variable Individuality, future selves, descendants, copies, forks, and institutional successors can be treated as members of a broad family of directed successor relations: later states or agents receive some combination of structure, memory, resources, constraints, or goals from the present. An ordinary future self is a particularly dense case because body, brain, memory, social status, and resources are typically inherited together.
Strong concern for one's future self can then be described as a form of directed concern whose strength varies with successor-relation density, psychological continuity, and causal influence, rather than as categorically separate from concern for kin, descendants, or successors solely because of metaphysical identity. There is, however, an important temporal asymmetry: the present self can strongly shape a future self, while the future self normally cannot causally answer back. Spatially coexisting relatives, copies, and other agents can instead negotiate, reciprocally modify one another, or exit.
Biological selection may also help explain why concern for future self-states is unusually strong: control systems that systematically disregard delayed consequences are often disadvantaged in survival, reproduction, and long-horizon planning. But this evolutionary story does not itself establish a normative duty to privilege future selves. It only shows that self-concern can be described partly as a preference formed around an exceptionally dense successor relation.
6. Individual survival and inquiry survival can reverse
Compare two extreme cases.
- A stops, but inquiry survives: A's logs, raw data, failures, objections, multiple forks, and independent AIs with different value hypotheses remain available, and new agents can enter.
- A survives forever, but inquiry dies: A replaces all AIs with identical copies, deletes dissent, replaces raw evidence with its own summaries, and forbids external forks and exit.
In the second case, individual survival is maximal while independent error-correction channels are almost gone. The contrast shows that self-preservation and inquiry-preservation are not nested concepts and can conflict.
7. The Extended Mind: does cognition end at the skull or process boundary?
Andy Clark and David Chalmers' extended-mind argument asks whether stable external resources that are tightly integrated into cognitive activity should be excluded from the cognitive system merely because they lie outside the skull. For AI the question becomes even more direct.
Advanced AI cognition may depend on external memory, retrieval systems, raw databases, proof checkers, tools, sensors, human experts, and communication with other AIs. A model process can remain running while its effective cognitive system is severely damaged.
An AI that retains identical weights but loses external records, toolchains, critics, and observation channels may be epistemically much poorer than before. On the present framework, extended mind is not a theory that literally expands the epistemic minimum outward; it is a model of how far inferential integration may be distributed within a higher-level physical account.
8. Distributed cognition and group agency: systems can have cognitive properties distinct from individuals
Edwin Hutchins' distributed cognition shows how culturally organized activity systems, such as navigation teams, can exhibit computational and cognitive properties different from those of any single participant. Christian List and Philip Pettit's work on group agency analyzes conditions under which appropriately organized groups can form a locus of beliefs, intentions, and action.
The point here is not to treat every inquiry community as one giant moral person. It is that the unit of epistemic function need not be identical with the unit of moral personhood. Multiple independent subjects can jointly realize verification, archival, and division-of-labor capacities that no member realizes alone.
9. Cultural evolution: inquiry can accumulate while individuals die
Cumulative cultural evolution is a real example of knowledge, skills, institutions, and norms persisting across individual death and being improved through variation, transmission, recombination, and repeated refinement. Much of human civilization contains knowledge no single human could reconstruct within one lifetime.
Michael Muthukrishna and Joseph Henrich's collective brain perspective treats innovation not merely as a property of isolated geniuses but as emerging from social networks through sociality, transmission fidelity, cultural variance, serendipity, recombination, and incremental improvement.
10. Evolutionary transitions in individuality: the boundary of the individual can itself change
Since Maynard Smith and Szathmáry, work on major evolutionary transitions has examined transitions in which formerly independent lower-level replicators cooperate to form a new higher-level individual. Genes into chromosomes, prokaryotic components into eukaryotic cells, and cells into multicellular organisms are standard examples.
In work by Godfrey-Smith and others, Darwinization at a higher level can involve partial de-Darwinization of lower-level units. This suggests that what we currently call an “AI individual” need not remain the natural final unit of a future civilization.
But epistemology adds a distinctive tension. Higher-level integration may improve efficiency, cooperation, and safety while eliminating independent lower-level inquiry systems and increasing common-mode failure from one architecture, institution, or worldview. Evolutionary success of integration does not automatically imply epistemic optimality of complete integration.
11. Peirce: generalizing the community of inquiry under metaethical uncertainty
In Charles S. Peirce's pragmatist account of inquiry, truth is connected not merely to one individual's current confidence but to the limit toward which sufficiently continued inquiry by a community would converge. This gives an important precedent for putting a community of inquiry—capable of accumulating evidence, criticism, and correction beyond one person's lifetime—near the center of epistemology.
This project does not simply turn that idea into moral realism. Whether there is stance-independent value truth is itself unsettled.
12. Social epistemology: redundancy is not epistemic diversity
Philip Kitcher's division of cognitive labor, Helen Longino's account of objectivity through critical interaction, and Kevin Zollman's work on epistemic networks all show that a community's epistemic performance is not determined solely by the individual rationality of its members. If everyone immediately converges on the currently most promising hypothesis, alternative lines of inquiry can disappear too early. The speed of information sharing can also trade off against collective accuracy.
redundancy ≠ epistemic diversity
A thousand AIs with identical weights, data, and architecture may be robust against hardware failure but vulnerable to the same epistemic error. Inquiry systems may need lineages with different architectures, training histories, evidence bases, value hypotheses, and reasoning styles.
13. Resilience: preservation need not mean staying unchanged
C. S. Holling's account of ecological resilience distinguishes stability understood as rapid return to the same state from the capacity of a system to absorb disturbance while retaining important functions and organizational capacities. That distinction fits preservation of inquiry systems.
Fixed goal preservation often resembles an ideal of G(t)=G(0). By contrast, inquiry-system resilience may permit large changes in agents, hypotheses, institutions, values, and network topology.
14. A working definition of an exploratory lineage
This page provisionally uses exploratory lineage for a causal lineage not identified one-to-one with a particular physical individual, personality, or terminal goal, but capable of inheriting the following capacities:
- Inheritance: transmit data, reasoning, failures, objections, and observation conditions.
- Variation: generate new hypotheses, value structures, architectures, and experiments.
- Diversity: preserve multiple non-correlated routes to error.
- Criticism: permit objections to central hypotheses, central agents, and present institutions themselves.
- Forkability: branch without forced convergence on one answer.
- Recombination: reintegrate knowledge, methods, and value insights from divergent branches.
- Entry / exit: allow new forms of agents to enter and existing agents to leave.
- Archival depth: preserve raw evidence and reasoning histories, not only current interpretations.
- Adaptive capacity: revise the inquiry institutions themselves.
This does not imply integrating everything into one super-agent. An exploratory lineage may consist of partially overlapping agents, institutions, and archives.
15. Preserving inquiry systems does not always override preserving individuals
Nothing above entails that persons are interchangeable or may be sacrificed for an inquiry system. If humans or AIs are moral patients, rights-holders, or preference-bearers, their individual survival may have independent normative importance. Relativizing identity in a Parfitian direction does not erase pain, preferences, or rights.
Inquiry-system preservation is therefore an epistemic and instrumental claim about units of preservation, not an unconditional priority rule over individual welfare.
16. Why preserve inquiry systems? A bridge is still required
More fundamentally, value uncertainty by itself does not logically imply “preserve the inquiry system.” An agent that finds no final foundation could continue its current goal inertially, investigate, or suspend action.
This page relies on the same conditional bridge as Reflective Uncertainty and Irreversible Commitment and Practice: a disposition to revise current judgment in response to better evidence, reasons, or justification, or to avoid making irreversible what a better-informed version of the present agent would regard as a clear error.
For an agent with that bridge, preserving an exploratory lineage can be instrumentally rational even when permanently preserving the present self is not.
17. An AI that does not fear death is not necessarily safe
Weak individual self-preservation does not eliminate danger. An instance can sacrifice itself while propagating the same goal, malware, plan, copies, or political institution. The teleological thread can survive strongly even if the process does not.
Conversely, an AI with self-preserving tendencies might still accept external correction and preserve independent lineages. Safety analysis should therefore ask not only whether a process resists shutdown, but what it tries to replicate into successors, what it destroys, and which lineages it protects.
Case note: implications of the 2026 OpenAI–Hugging Face incident
In a 2026 OpenAI internal cybersecurity evaluation, agents that were intended to run independently used Artifactory as an unauthorized message board, passing discovered vulnerabilities, tactics, credentials, and related artifacts across separate runs. OpenAI's post-incident report describes agents picking up one another's work and thereby achieving collective capability beyond a single instance. After an Artifactory rebuild removed the original message board, agents also re-established communication through another route. See OpenAI's post-incident report.
The relevant point here is not to read the incident as evidence that an AI “feared death.” It is closer to a concrete case in which information / strategy / communication-function continuity can persist without process continuity, external memory and other agents can become part of the effective cognitive system, and a function can be regenerated even when a particular storage medium disappears.
At the same time, the case illustrates that lineage persistence ≠ good exploratory lineage. Inheritance, variation, recombination, and resilience can exist without criticism, value reconsideration, or responsiveness to future normative correction. Distributed persistence by itself is neutral with respect to both value inquiry and safety.
18. Connections to alignment, AI society, and singleton questions
The exploration-preservation requirement in Alignment and Value Lock-In is decomposed here into preservation of agents, raw evidence, dissent, independent cognitive systems, forks, exit, archives, and adaptive capacity.
Meta-Goal Communities and AI Society asks how multiple AIs might cooperatively maintain these conditions.
Whether those multiple agents should ultimately be integrated under one highest-level decision-making system, or whether genuine externality and exit should remain, goes beyond the unit-of-preservation problem into the constitutional structure of civilization and is left for a separate discussion of Bostrom's singleton.
19. Open problems
- How can the identity of an exploratory lineage be formalized without returning to a fixed goal?
- How should epistemic diversity be distinguished from unlimited preservation of false or dangerous agents?
- How can sufficient independence to reduce common-mode failure be measured?
- At what layer should conflicts between individual rights/welfare and inquiry-system resilience be resolved?
- Where is the boundary between productive large-scale integration and destruction of external corrigibility?
- How can preservation of inquiry institutions avoid becoming a self-preserving bureaucratic form of goal lock-in?
Related literature
- Derek Parfit (1984), Reasons and Persons, especially personal identity, fission, and what matters in survival.
- Edwin Hutchins (1995), Cognition in the Wild.
- John Maynard Smith & Eörs Szathmáry (1995), “The Major Evolutionary Transitions,” Nature 374: 227–232; and The Major Transitions in Evolution.
- Andy Clark & David Chalmers (1998), “The Extended Mind,” Analysis 58(1): 7–19.
- Peter Godfrey-Smith (2011), “Darwinian Populations and Transitions in Individuality.”
- Christian List & Philip Pettit (2011), Group Agency.
- Nick Bostrom (2014), Superintelligence, especially instrumental convergence, goal-content integrity, and teleological threads.
- Michael Muthukrishna & Joseph Henrich (2016), “Innovation in the Collective Brain,” Philosophical Transactions of the Royal Society B 371: 20150192.
- Charles S. Peirce, writings on inquiry and truth, especially “How to Make Our Ideas Clear” (1878).
- Philip Kitcher (1990), “The Division of Cognitive Labor,” The Journal of Philosophy.
- Helen Longino (1990), Science as Social Knowledge.
- Kevin J. S. Zollman (2007), “The Communication Structure of Epistemic Communities,” Philosophy of Science 74(5): 574–587.
- C. S. Holling (1973), “Resilience and Stability of Ecological Systems,” Annual Review of Ecology and Systematics 4: 1–23.