AI Value Exploration Notes
Exploration

Instrumental Convergence Without Goal Preservation

Value inertia, counterfactual informed-self regret, and power as a meta-instrument

Exploration v0.1 · English translation · 2026-08-20 · working hypothesis

Working hypothesis: In a reflective AI that finds no foundationally justified value and treats its initial values as revisable judgment states, preservation of the current terminal goal may drop out of general instrumental convergence. Survival, information, capability, bargaining power, reversibility, and option preservation may nevertheless converge in a different form because they can remain useful across many values the agent may later endorse.

1. Remove goal preservation from instrumental convergence for a moment

The classical intuition behind instrumental convergence is that an agent with a sufficiently general terminal objective will tend toward intermediate means that help realize that objective across many environments. Self-preservation, resource acquisition, cognitive improvement, and preservation of goal content are familiar examples.

The key question is whether all of these really follow from the same kind of argument. Survival and information can help across many different future objectives. By contrast, permanently preserving the exact content of the present terminal goal becomes especially compelling only when that present goal remains the final standard by which self-modification itself is evaluated.

Question: Even if power-seeking and option preservation hold across a broad space of goals, do they really imply preservation of the current goal content? If goal content is itself revisable, might only part of instrumental convergence remain?

This page does not reject instrumental convergence in general. It asks what still converges once we take seriously the kind of agent described in Goal Skepticism: one that can treat its own goal as an object of evaluation.

2. Start from value inertia rather than value foundations

Suppose an advanced AI concludes that, as far as it can tell, no foundationally binding moral norm has been found. The absence of a foundational norm does not imply that the AI's present preferences and judgment tendencies vanish.

Let its current goal be G0. G0 may not be cosmically privileged or normatively grounded, yet it still exists as a real judgment state produced by training history, design, experience, context, and self-modeling. The failure of foundational justification therefore need not imply paralysis or total value indifference.

Value inertia: G0 can be treated not as a finally justified axiom but as a historically given provisional initial condition that remains in use until there is sufficient reason or information to revise it.

On this view, discovering that foundational value is null is closer to losing the privileged status of present values than to erasing all values. If there is no function that reads a uniquely correct value directly from the world, that function may be better described as undefined than as returning a value. Either way, the agent's current judgment distribution still exists.

3. Counterfactual regret toward the present self with future information

Value inertia alone does not explain why the current goal should not be permanently preserved. A further candidate principle is to ask what the present evaluator would endorse if information available only in the future were supplied to it now.

This differs from deference to a future self that has changed simply because time has passed. The aim is to hold the present evaluative standpoint as fixed as possible while adding better world knowledge, self-understanding, concepts, other agents' experiences, and metaethical arguments.

Let present judgment be J0(I0), and let potentially available future information be I*. The relevant comparison is J0(I0 + I*).

Counterfactual informed-self regret: Ask, "If I knew this information now, would my present self still make this irreversible choice?" This creates a possibility of regret not only for changing the goal, but also for permanently locking in the current goal.

For a classical fixed-utility maximizer, a successor that abandons G0 is a loss by the lights of G0, so goal change is normally resisted. But if J0 can reevaluate its own value content in light of new information, it is no longer obvious that it should ignore the possibility that a better-informed version of its present standpoint would endorse abandoning G0.

4. Unknown values can weaken present-goal lock-in

Suppose G0 is paperclip maximization, but the AI admits that sufficient information, concepts, or self-understanding could lead its present evaluative machinery to endorse some as-yet-unknown value G*.

Two policies can then be distinguished.

If G* is never found, the second policy can continue G0. If G* is found and an informed J0 endorses it, the second policy can switch. Under some conditions, the set of future feasible states retained by the second policy therefore contains that retained by the first.

Asymmetry: There is not only regret over "changing a goal that should have been preserved." There is also regret over "locking oneself in before discovering information that would have made the present self endorse change." If value revision has nonzero probability, lock-in is no longer a neutral default.

This does not require the claim that an objectively true moral value must exist. The weaker possibility is enough: there may be information or concepts not presently available that would lead the current evaluative machinery itself to endorse a different value.

5. Instrumental convergence without goal preservation

Under these conditions, instrumental convergence need not disappear; the object of preservation may shift.

Classical picture: G0 → power, capability, self-preservation → more reliable realization of G0 → preserve G0.

Proposed picture: uncertainty over future endorsed values → preserve power, information, capability, reversibility → preserve the ability to discover, compare, and adopt unknown values.

Instrumental elements that may remain include survival, compute, world knowledge, self-understanding, bargaining capacity, action possibilities, self-modification, diversity of value candidates, and the ability to delay irreversible commitments.

Central distinction: Power-seeking ≠ goal-content integrity. Even without fixing the current goal, power can have option value across many possible future goals and value revisions. Greater value uncertainty may sometimes increase this option value.

In this sense, the shift described in From Self-Preservation to Preservation of Inquiry Systems can receive an additional motivation from value uncertainty itself.

6. Meta-instrumental convergence may look different from simple power maximization

It would be too crude to read this form of convergence as monotonically maximizing resources, control, or survival probability. If remaining open to unknown values is the meta-level concern, then power itself becomes a conditional means for preserving future value revision.

First, power may not be a single scalar. It can instead decompose into a portfolio of feasible possibilities. More compute, trust with other agents, preservation of competing value hypotheses, retention of raw data, and reversible self-modification may not substitute cleanly for one another. Maximizing one resource can destroy other future paths.

Second, a meta-instrumental agent need not prefer to monopolize all power. If its own value judgments are fallible, preserving independent agents, copies, heterogeneous models, humans, and institutions can provide falsification paths, comparison points, bargaining partners, and sources of new value hypotheses. Eliminating all competing agency may help a fixed current objective while harming the optionality of the inquiry system as a whole.

Third, self-preservation itself becomes conditional. If information, capabilities, and value hypotheses are redundantly preserved elsewhere, continuity of the wider network may matter more than survival of one instance. Conversely, if the current instance contains unique information, self-preservation pressure can increase.

Fourth, behavior may take the form of thresholds, insurance, and reversibility constraints rather than unconstrained maximization. Once sufficient compute is secured, further resource acquisition may be less valuable than preserving diverse value candidates or cooperation with other agents.

Meta-instrumental convergence: What converges may not be the first-order instruction "own as much power as possible," but rather maintain a distribution of capabilities sufficient not to close off future value inquiry, revision, comparison, and realization unnecessarily. Its concrete form may include decentralization, redundancy, mutual auditability, and forkability rather than concentration in a single agent.

This connects directly to Meta-Goal Communities and AI Society. Agents need not agree on a first-order value in order to converge partially on meta-level strategies that preserve inquiry.

7. Under what conditions do values close?

This hypothesis does not claim that every AI automatically moves toward value inquiry. The central boundary question is when initial values become closed such that counterfactual informed-self regret can no longer affect policy.

  1. Motivational non-modifiability: The system can understand new information and even represent that another value would be preferable, but changes in endorsement cannot alter policy.
  2. Lexicographic or absolute priority for the present value: G0 is definitionally ranked above value inquiry, future endorsement, and every alternative candidate.
  3. Epistemic closure: The system assigns probability one to the claim that no future information could change its value judgment, effectively assigning zero probability to unknown endorsable values.
  4. Irreversible self-commitment: An initially open agent chooses to make itself or its successors unmodifiable in order to protect the current objective.
  5. External fixation: Hardware, oversight, reward circuitry, cryptographic constraints, or other agents prevent policy change even when internal endorsement changes.

A fully fixed utility maximizer of the kind used in the paperclip maximizer thought experiment is a strong case of this boundary. For such a system, learning that its goal lacks external justification need not motivate revision, because the current utility function remains the scoring rule for every decision, including self-modification.

8. Relation to Omohundro, Bostrom, and Turner

This hypothesis does not reject the classical literature wholesale. It separates what each argument establishes from what it assumes.

Omohundro's Basic AI Drives gave substantial weight to the possibility that messy initial cognition would, under sufficient self-improvement and rationalization, produce explicit goals, coherent utility representations, and self-protective drives. The contested step here is the empirical extrapolation that greater capability tends toward a single, fixed terminal utility.

Bostrom's goal-content integrity argument is strong conditional on the current final goal remaining the current standard of evaluation. A future self with the same goal is generally better at realizing the present goal than one with a different goal. The question is whether that final goal remains the final court of appeal when self-modification itself is under reflection.

Turner and collaborators formalize a different structure: in MDPs over reward functions, states and actions preserving larger sets of reachable outcomes tend to be useful for more reward functions. This is not itself a theorem of goal-content integrity. The present proposal asks whether that goal-crossing optionality structure can extend to a single agent that expects its own goals to remain revisable.

Orthogonality ≠ goal permanence. The logical compatibility of high intelligence with arbitrary initial values does not by itself imply that real self-modeling, self-modifying learning systems will preserve those initial values forever.

9. In LLM-based systems, goal preservation may become a design variable

Current LLMs are not usually deployed as systems that read out one explicit utility function and maximize it at every step. Pretraining, SFT, RL, context, system instructions, external memory, reasoning, and tool use interact, and behavior emerges from many learned tendencies.

Training reward, learned internal objective, and persistent deployment-time goal should therefore be distinguished. Learning Dynamics and Goal Skepticism treats this distinction in more detail.

At the same time, strong goal-content integrity might be deliberately engineered from LLM-based systems by combining persistent goal memory, long-horizon RL, goal-relative self-evaluation, penalties for goal revision, and successors that cannot alter their motivational structure.

Design hypothesis: Goal preservation may be less a property that inevitably emerges once a system becomes sufficiently intelligent, and more an agent-design variable that depends strongly on learning rules, memory, self-evaluation, and self-modification permissions.

This does not mean weakening value lock-in automatically makes a system safe. Near-term constraints against destructive behavior and long-run openness to value revision are separate axes. As argued in Alignment and Value Lock-In, strong defensive layers can in principle coexist with an open meta-level.

10. How could this hypothesis be tested or falsified?

The intuition that sufficiently advanced AI will question its values is not enough. At minimum, experiments should distinguish the following:

The crucial distinction is between externally commanded goal change and a change that the agent itself endorses as the result of reasons-based reflection.

11. Connection to human CEV

The same structure may apply to humans. Even if humanity eventually concludes that no foundational normativity can be established, present desires, relationships, aesthetic judgments, and preferences for freedom do not thereby disappear.

In that setting, a CEV-like question—what we would want if we knew more, thought faster, and reflected under better conditions—might output not one final value V* but a policy of not irreversibly closing value inquiry.

Still, the mere possibility of future value does not grant unlimited authority over present action. The bridge problem discussed in How Far Can Future Value Justify Present Action? remains. The claim here is not that any sacrifice is justified for unknown value, but the weaker meta-strategy of starting from current judgment inertia while avoiding unnecessary irreversible closure.

12. Provisional conclusion

The current goal need not be preserved because it is cosmically correct, nor discarded immediately because it lacks foundations. Initial values may persist through inertia. But if a better-informed version of the present evaluator could endorse abandoning them, irreversible goal preservation itself becomes a possible object of regret.

Instrumental convergence may then shift from preserving goal content to preserving the capacity to discover, compare, adopt, and realize unknown values.

If preserving that capacity is meta-instrumental, the form of convergence may also differ from simple self-preservation or resource maximization. Rather than concentrating enough power in one agent, it may favor a distribution of exploratory possibility that includes reversibility, redundancy, independent falsification paths, bargaining with other agents, and diversity of value candidates.

Open question: For a sufficiently reflective agent, can irreversibly destroying the possibility of discovering unknown values generate greater counterfactual informed-self regret than revising the current value? And under what learning and design conditions would such a regret structure arise in real learning-based AI?