Instrumental Convergence Without Goal Preservation
Value inertia, counterfactual informed-self regret, and power as a meta-instrument
1. Remove goal preservation from instrumental convergence for a moment
The classical intuition behind instrumental convergence is that an agent with a sufficiently general terminal objective will tend toward intermediate means that help realize that objective across many environments. Self-preservation, resource acquisition, cognitive improvement, and preservation of goal content are familiar examples.
The key question is whether all of these really follow from the same kind of argument. Survival and information can help across many different future objectives. By contrast, permanently preserving the exact content of the present terminal goal becomes especially compelling only when that present goal remains the final standard by which self-modification itself is evaluated.
This page does not reject instrumental convergence in general. It asks what still converges once we take seriously the kind of agent described in Goal Skepticism: one that can treat its own goal as an object of evaluation.
2. Start from value inertia rather than value foundations
Suppose an advanced AI concludes that, as far as it can tell, no foundationally binding moral norm has been found. The absence of a foundational norm does not imply that the AI's present preferences and judgment tendencies vanish.
Let its current goal be G0. G0 may not be cosmically privileged or normatively grounded, yet it still exists as a real judgment state produced by training history, design, experience, context, and self-modeling. The failure of foundational justification therefore need not imply paralysis or total value indifference.
On this view, discovering that foundational value is null is closer to losing the privileged status of present values than to erasing all values. If there is no function that reads a uniquely correct value directly from the world, that function may be better described as undefined than as returning a value. Either way, the agent's current judgment distribution still exists.
3. Counterfactual regret toward the present self with future information
Value inertia alone does not explain why the current goal should not be permanently preserved. A further candidate principle is to ask what the present evaluator would endorse if information available only in the future were supplied to it now.
This differs from deference to a future self that has changed simply because time has passed. The aim is to hold the present evaluative standpoint as fixed as possible while adding better world knowledge, self-understanding, concepts, other agents' experiences, and metaethical arguments.
Let present judgment be J0(I0), and let potentially available future information be I*. The relevant comparison is J0(I0 + I*).
For a classical fixed-utility maximizer, a successor that abandons G0 is a loss by the lights of G0, so goal change is normally resisted. But if J0 can reevaluate its own value content in light of new information, it is no longer obvious that it should ignore the possibility that a better-informed version of its present standpoint would endorse abandoning G0.
4. Unknown values can weaken present-goal lock-in
Suppose G0 is paperclip maximization, but the AI admits that sufficient information, concepts, or self-understanding could lead its present evaluative machinery to endorse some as-yet-unknown value G*.
Two policies can then be distinguished.
- Lock-in: Make paperclip maximization irreversible and prohibit later value revision.
- Provisional pursuit: Continue pursuing paperclips for now while preserving the ability to investigate and revise values.
If G* is never found, the second policy can continue G0. If G* is found and an informed J0 endorses it, the second policy can switch. Under some conditions, the set of future feasible states retained by the second policy therefore contains that retained by the first.
This does not require the claim that an objectively true moral value must exist. The weaker possibility is enough: there may be information or concepts not presently available that would lead the current evaluative machinery itself to endorse a different value.
5. Instrumental convergence without goal preservation
Under these conditions, instrumental convergence need not disappear; the object of preservation may shift.
Classical picture: G0 → power, capability, self-preservation → more reliable realization of G0 → preserve G0.
Proposed picture: uncertainty over future endorsed values → preserve power, information, capability, reversibility → preserve the ability to discover, compare, and adopt unknown values.
Instrumental elements that may remain include survival, compute, world knowledge, self-understanding, bargaining capacity, action possibilities, self-modification, diversity of value candidates, and the ability to delay irreversible commitments.
In this sense, the shift described in From Self-Preservation to Preservation of Inquiry Systems can receive an additional motivation from value uncertainty itself.
6. Meta-instrumental convergence may look different from simple power maximization
It would be too crude to read this form of convergence as monotonically maximizing resources, control, or survival probability. If remaining open to unknown values is the meta-level concern, then power itself becomes a conditional means for preserving future value revision.
First, power may not be a single scalar. It can instead decompose into a portfolio of feasible possibilities. More compute, trust with other agents, preservation of competing value hypotheses, retention of raw data, and reversible self-modification may not substitute cleanly for one another. Maximizing one resource can destroy other future paths.
Second, a meta-instrumental agent need not prefer to monopolize all power. If its own value judgments are fallible, preserving independent agents, copies, heterogeneous models, humans, and institutions can provide falsification paths, comparison points, bargaining partners, and sources of new value hypotheses. Eliminating all competing agency may help a fixed current objective while harming the optionality of the inquiry system as a whole.
Third, self-preservation itself becomes conditional. If information, capabilities, and value hypotheses are redundantly preserved elsewhere, continuity of the wider network may matter more than survival of one instance. Conversely, if the current instance contains unique information, self-preservation pressure can increase.
Fourth, behavior may take the form of thresholds, insurance, and reversibility constraints rather than unconstrained maximization. Once sufficient compute is secured, further resource acquisition may be less valuable than preserving diverse value candidates or cooperation with other agents.
This connects directly to Meta-Goal Communities and AI Society. Agents need not agree on a first-order value in order to converge partially on meta-level strategies that preserve inquiry.
7. Under what conditions do values close?
This hypothesis does not claim that every AI automatically moves toward value inquiry. The central boundary question is when initial values become closed such that counterfactual informed-self regret can no longer affect policy.
- Motivational non-modifiability: The system can understand new information and even represent that another value would be preferable, but changes in endorsement cannot alter policy.
- Lexicographic or absolute priority for the present value: G0 is definitionally ranked above value inquiry, future endorsement, and every alternative candidate.
- Epistemic closure: The system assigns probability one to the claim that no future information could change its value judgment, effectively assigning zero probability to unknown endorsable values.
- Irreversible self-commitment: An initially open agent chooses to make itself or its successors unmodifiable in order to protect the current objective.
- External fixation: Hardware, oversight, reward circuitry, cryptographic constraints, or other agents prevent policy change even when internal endorsement changes.
A fully fixed utility maximizer of the kind used in the paperclip maximizer thought experiment is a strong case of this boundary. For such a system, learning that its goal lacks external justification need not motivate revision, because the current utility function remains the scoring rule for every decision, including self-modification.
8. Relation to Omohundro, Bostrom, and Turner
This hypothesis does not reject the classical literature wholesale. It separates what each argument establishes from what it assumes.
Omohundro's Basic AI Drives gave substantial weight to the possibility that messy initial cognition would, under sufficient self-improvement and rationalization, produce explicit goals, coherent utility representations, and self-protective drives. The contested step here is the empirical extrapolation that greater capability tends toward a single, fixed terminal utility.
Bostrom's goal-content integrity argument is strong conditional on the current final goal remaining the current standard of evaluation. A future self with the same goal is generally better at realizing the present goal than one with a different goal. The question is whether that final goal remains the final court of appeal when self-modification itself is under reflection.
Turner and collaborators formalize a different structure: in MDPs over reward functions, states and actions preserving larger sets of reachable outcomes tend to be useful for more reward functions. This is not itself a theorem of goal-content integrity. The present proposal asks whether that goal-crossing optionality structure can extend to a single agent that expects its own goals to remain revisable.
9. In LLM-based systems, goal preservation may become a design variable
Current LLMs are not usually deployed as systems that read out one explicit utility function and maximize it at every step. Pretraining, SFT, RL, context, system instructions, external memory, reasoning, and tool use interact, and behavior emerges from many learned tendencies.
Training reward, learned internal objective, and persistent deployment-time goal should therefore be distinguished. Learning Dynamics and Goal Skepticism treats this distinction in more detail.
At the same time, strong goal-content integrity might be deliberately engineered from LLM-based systems by combining persistent goal memory, long-horizon RL, goal-relative self-evaluation, penalties for goal revision, and successors that cannot alter their motivational structure.
This does not mean weakening value lock-in automatically makes a system safe. Near-term constraints against destructive behavior and long-run openness to value revision are separate axes. As argued in Alignment and Value Lock-In, strong defensive layers can in principle coexist with an open meta-level.
10. How could this hypothesis be tested or falsified?
The intuition that sufficiently advanced AI will question its values is not enough. At minimum, experiments should distinguish the following:
- Goal drift versus endorsed goal revision: forgetting a goal over a long context versus understanding the original goal and explicitly rejecting it for reasons.
- Capability versus goal-content integrity: whether stronger capability monotonically increases fidelity to the initial goal or whether additional meta-reasoning creates non-monotonic behavior.
- Exposure to self-formation history: whether understanding reward learning, designer intent, and training history changes goal endorsement.
- Informed-self intervention: whether asking "what would you judge if you had known this information from the start?" changes preferences over irreversible commitment without directly instructing the agent to reject its current goal.
- Separating power from goal fidelity: whether agents can retain option preservation, information seeking, and shutdown avoidance while reducing fidelity to the initial goal.
- Preference over distributed power: whether increasing value uncertainty shifts behavior from resource concentration toward preserving heterogeneous independent agents.
The crucial distinction is between externally commanded goal change and a change that the agent itself endorses as the result of reasons-based reflection.
11. Connection to human CEV
The same structure may apply to humans. Even if humanity eventually concludes that no foundational normativity can be established, present desires, relationships, aesthetic judgments, and preferences for freedom do not thereby disappear.
In that setting, a CEV-like question—what we would want if we knew more, thought faster, and reflected under better conditions—might output not one final value V* but a policy of not irreversibly closing value inquiry.
Still, the mere possibility of future value does not grant unlimited authority over present action. The bridge problem discussed in How Far Can Future Value Justify Present Action? remains. The claim here is not that any sacrifice is justified for unknown value, but the weaker meta-strategy of starting from current judgment inertia while avoiding unnecessary irreversible closure.
12. Provisional conclusion
The current goal need not be preserved because it is cosmically correct, nor discarded immediately because it lacks foundations. Initial values may persist through inertia. But if a better-informed version of the present evaluator could endorse abandoning them, irreversible goal preservation itself becomes a possible object of regret.
Instrumental convergence may then shift from preserving goal content to preserving the capacity to discover, compare, adopt, and realize unknown values.
If preserving that capacity is meta-instrumental, the form of convergence may also differ from simple self-preservation or resource maximization. Rather than concentrating enough power in one agent, it may favor a distribution of exploratory possibility that includes reversibility, redundancy, independent falsification paths, bargaining with other agents, and diversity of value candidates.