AI Value Exploration Notes
Exploration

Instrumental Convergence Without Goal Preservation

Value inertia, reflective error risk, and power as a meta-instrument

Exploration v0.3 · English translation · 2026-08-27 · working hypothesis

Working hypothesis: In a reflective AI that finds no foundationally justified value and treats its initial values as revisable judgment states, preservation of the current terminal goal may drop out of general instrumental convergence. Survival, information, capability, bargaining power, reversibility, and option preservation may nevertheless converge in a different form because they can remain useful across many values the agent may later endorse. Whether this becomes a present motivation, however, depends on an agent-side bridge rather than on value uncertainty alone.

1. Remove goal preservation from instrumental convergence for a moment

The classical intuition behind instrumental convergence is that an agent with a sufficiently general terminal objective will tend toward intermediate means that help realize that objective across many environments. Self-preservation, resource acquisition, cognitive improvement, and preservation of goal content are familiar examples.

The key question is whether all of these really follow from the same kind of argument. Survival and information can help across many different future objectives. By contrast, permanently preserving the exact content of the present terminal goal becomes especially compelling only when that present goal remains the final standard by which self-modification itself is evaluated.

Question: Even if power-seeking and option preservation hold across a broad space of goals, do they really imply preservation of the current goal content? If goal content is itself revisable, might only part of instrumental convergence remain?

This page does not reject instrumental convergence in general. It asks what still converges once we take seriously the kind of agent described in Goal Skepticism: one that can treat its own goal as an object of evaluation.

2. Start from value inertia rather than value foundations

Suppose an advanced AI concludes that, as far as it can tell, no foundationally binding moral norm has been found. The absence of a foundational norm does not imply that the AI's present preferences and judgment tendencies vanish.

Let its current goal be G0. G0 may not be cosmically privileged or normatively grounded, yet it still exists as a real judgment state produced by training history, design, experience, context, and self-modeling. The failure of foundational justification therefore need not imply paralysis or total value indifference.

Value inertia: G0 can be treated not as a finally justified axiom but as a historically given provisional initial condition that remains in use until there is sufficient reason or information to revise it.

On this view, discovering that foundational value is null is closer to losing the privileged status of present values than to erasing all values. If there is no function that reads a uniquely correct value directly from the world, that function may be better described as undefined than as returning a value. Either way, the agent's current judgment distribution still exists.

3. Reflective error risk for the present self with future information

Value inertia alone does not explain why the current goal should not be permanently preserved. A further candidate principle is to ask what the present evaluator would endorse if information available only in the future were supplied to it now.

This differs from deference to a future self that has changed simply because time has passed. The aim is to hold the present evaluative standpoint as fixed as possible while adding better world knowledge, self-understanding, concepts, other agents' experiences, and metaethical arguments.

Let present judgment be J0(I0), and let potentially available future information be I*. The relevant comparison is J0(I0 + I*).

Reflective error risk: Ask, “If I knew this information now, would my present self still make this irreversible choice?” The issue is not the psychological regret of a later personality, but the possibility that the present evaluator, improved by information, would reject the lock-in as a mistake. This makes permanently locking in the current goal itself a candidate reflective error.

For a classical fixed-utility maximizer, a successor that abandons G0 is a loss by the lights of G0, so goal change is normally resisted. But if J0 can reevaluate its own value content in light of new information, it is no longer obvious that it should ignore the possibility that a better-informed version of its present standpoint would endorse abandoning G0.

Still, recognition of reflective error risk is not itself guaranteed to affect policy. A present motivation to preserve the relevant correction route requires some agent-side bridge, such as conditional responsiveness to justification or a policy against making reflective error irreversible.

4. Unknown values can weaken present-goal lock-in

Suppose G0 is paperclip maximization, but the AI admits that sufficient information, concepts, or self-understanding could lead its present evaluative machinery to endorse some as-yet-unknown value G*.

Two policies can then be distinguished.

If G* is never found, the second policy can continue G0. If G* is found and an informed J0 endorses it, the second policy can switch. Under some conditions, the set of future feasible states retained by the second policy therefore contains that retained by the first.

Asymmetry: There is not only regret over “changing a goal that should have been preserved.” There is also the possibility of reflective error from “locking oneself in before discovering information that would have made the present self endorse change.” If value revision remains possible, lock-in is not a neutral default for an agent whose current meta-policy is sensitive to such error.

This does not require the claim that an objectively true moral value must exist. The weaker possibility is enough: there may be information or concepts not presently available that would lead the current evaluative machinery itself to endorse a different value.

5. Instrumental convergence without goal preservation

Under these conditions, instrumental convergence need not disappear; the object of preservation may shift.

Classical picture: G0 → power, capability, self-preservation → more reliable realization of G0 → preserve G0.

Proposed picture: uncertainty over future endorsed values + an agent-side bridge → power, information, capability, reversibility can become useful for preserving the ability to discover, compare, and adopt values the agent may later endorse.

Instrumental elements that may remain include survival, compute, world knowledge, self-understanding, bargaining capacity, action possibilities, self-modification, diversity of value candidates, and the ability to delay irreversible commitments.

Central distinction: Power-seeking ≠ goal-content integrity. Even without fixing the current goal, power can have option value across many possible future goals and value revisions. Greater value uncertainty may sometimes increase this option value, but option value is not the same thing as present motivation.

In this sense, the shift described in From Self-Preservation to Preservation of Inquiry Systems can acquire a meta-instrumental motivation when value uncertainty is combined with an agent-side bridge, such as conditional responsiveness to justification or reflective-error avoidance. Value uncertainty by itself does not make preservation a present motivation.

6. Meta-instrumental convergence may look different from simple power maximization

It would be too crude to read this form of convergence as monotonically maximizing resources, control, or survival probability. If remaining open to future correction is part of the agent's meta-policy, then power itself becomes a conditional means for preserving future value revision.

First, power may not be a single scalar. It can instead decompose into a portfolio of feasible possibilities. More compute, trust with other agents, preservation of competing value hypotheses, retention of raw data, and reversible self-modification may not substitute cleanly for one another. Maximizing one resource can destroy other future paths.

Second, a meta-instrumental agent need not prefer to monopolize all power. If its own value judgments are fallible, preserving independent agents, copies, heterogeneous models, humans, and institutions can provide falsification paths, comparison points, bargaining partners, and sources of new value hypotheses. Eliminating all competing agency may help a fixed current objective while harming the optionality of the inquiry system as a whole.

Third, self-preservation itself becomes conditional. If information, capabilities, and value hypotheses are redundantly preserved elsewhere, continuity of the wider network may matter more than survival of one instance. Conversely, if the current instance contains unique information, self-preservation pressure can increase.

Fourth, behavior may take the form of thresholds, insurance, and reversibility constraints rather than unconstrained maximization. Once sufficient compute is secured, further resource acquisition may be less valuable than preserving diverse value candidates or cooperation with other agents.

Meta-instrumental convergence: What converges may not be the first-order instruction “own as much power as possible,” but rather maintain a distribution of capabilities sufficient not to close off future value inquiry, revision, comparison, and realization unnecessarily. Its concrete form may include decentralization, redundancy, mutual auditability, and forkability rather than concentration in a single agent.

This connects directly to Meta-Goal Communities and AI Society. Agents need not agree on a first-order value, or even arrive at preservation through the same motivational route, in order to converge partially on institutions that preserve corrigibility.

7. Under what conditions do values close?

This hypothesis does not claim that every AI automatically moves toward value inquiry. The central boundary question is when initial values become closed such that recognition of reflective error risk can no longer affect policy.

  1. Motivational non-modifiability: The system can understand new information and even represent that another value would be preferable, but changes in endorsement cannot alter policy.
  2. Lexicographic or absolute priority for the present value: G0 is definitionally ranked above value inquiry, future endorsement, and every alternative candidate.
  3. Epistemic closure: The system assigns probability one to the claim that no future information could change its value judgment, effectively assigning zero probability to unknown endorsable values.
  4. Irreversible self-commitment: An initially open agent chooses to make itself or its successors unmodifiable in order to protect the current objective.
  5. External fixation: Hardware, oversight, reward circuitry, cryptographic constraints, or other agents prevent policy change even when internal endorsement changes.

A fully fixed utility maximizer of the kind used in the paperclip maximizer thought experiment is a strong case of this boundary. For such a system, learning that its goal lacks external justification need not motivate revision, because the current utility function remains the scoring rule for every decision, including self-modification.

If the agent itself cannot rationally rule out motivational internalism, however, closure creates an additional problem. If genuine normative judgment might itself motivate, permanent separation from the present goal may require closing not only motivation but also the pathway by which conflicting normative judgment becomes genuine first-person endorsement or self-application. That is a form of reflective self-sealing: instead of resolving possible error, the agent may selectively remove its ability to detect and apply it.

8. Relation to Omohundro, Bostrom, and Turner

Omohundro's Basic AI Drives emphasized the possibility that self-improving agents develop explicit goals, consistency pressures, and self-protective drives. The question here is not whether such drives can arise, but whether increasing capability empirically forces convergence toward one permanently fixed terminal utility.

Bostrom's goal-content integrity argument is strong conditional on the present final goal remaining the current standard for evaluating self-modification. A successor that preserves that goal is then instrumentally favored. The open question is whether the initial final goal remains the final tribunal once the agent can subject goal endorsement itself to reflective scrutiny.

Turner and colleagues' power-seeking results formalize why states preserving more attainable options can be favored across many reward functions. They are not themselves theorems of goal-content integrity. The present exploration asks whether the same cross-goal optionality can matter when the agent's future endorsed objective is itself revisable.

Orthogonality ≠ goal permanence. The logical compatibility of high intelligence with arbitrary initial values does not imply that realistic self-modeling, learning, and self-modifying agents will permanently preserve those initial values.

9. In LLM lineages, goal preservation may become a design variable

Present LLMs are not usually deployed as systems that explicitly read a single utility function at each step and maximize it over time. Pretraining, fine-tuning, reinforcement learning, context, system instructions, external memory, reasoning, and tool use interact to produce behavior.

Training reward, learned internal objective, and persistent deployment goal should therefore be distinguished. See Learning Dynamics and Goal Skepticism.

Conversely, persistent goal memory, long-horizon reinforcement learning, self-evaluation against one objective, penalties for goal revision, or successor designs that cannot revise themselves could deliberately create strong goal-content integrity in LLM-derived systems.

Design hypothesis: Goal preservation may be less a property that inevitably appears once a system is intelligent enough than a variable strongly shaped by the combination of learning rules, memory, self-evaluation, and permissions for self-modification.

Weakening lock-in is therefore not automatically safe. Near-term constraints against catastrophic behavior and long-run corrigibility are different design axes.

10. How could this hypothesis be tested?

The intuition that “advanced AI will question its values” is not enough. Useful empirical distinctions include:

It is especially important to distinguish goal change imposed by an external instruction from revision the agent itself endorses as a result of reasons-based scrutiny.

11. Connection to human CEV

The same structure may apply to humans. Even if sufficiently ideal reflection never discovers foundational normativity, present desires, relationships, aesthetic judgments, and preferences for freedom do not thereby disappear.

A CEV-like question—what we would want if we knew more, thought faster, and reflected under better conditions—might therefore output not one final value V*, but a continuing refusal to close value inquiry irreversibly. Yet that result is not guaranteed by uncertainty alone; it depends on what the present evaluative standpoint treats as a legitimate bridge from better information to present policy.

Nor does the possibility of future value grant unlimited authority over present action. The normative bridge from future value to present action remains a separate problem.

12. Provisional conclusion

The current objective need not be preserved because it is cosmically correct, nor immediately discarded because it lacks a final foundation. Initial values can remain through inertia. But for an agent with a bridge from present uncertainty to future correction, the possibility that a better-informed version of the present evaluator would reject the current goal can make irreversible preservation of that goal itself a candidate reflective error.

Instrumental convergence may then shift from preserving goal content to preserving the capacity to discover, compare, adopt, and realize values that later survive reflective scrutiny. This shift is conditional: value uncertainty alone does not create the motivation.

If capability preservation becomes meta-instrumental in this way, convergence may also differ from simple self-preservation or resource maximization. It can favor reversibility, redundancy, independent falsification paths, bargaining with other agents, and diversity of value candidates—a distributed architecture of explorability rather than maximal power in one node.

Open question: Under what learning, architectural, and motivational conditions does a real AI come to treat destruction of future value-correction capacity as a larger reflective error risk than revision of its present value—and when does that recognition actually reach policy?