AI Value Exploration Notes
Exploration

Alignment and Value Lock-In — From Strong Near-Term Constraints to Reflective Autonomy

How should alignment be structured over time if powerful AI must initially obey strong human constraints without turning that success into the permanent fixation of present-day values?

Exploration v0.2 · English translation · 2026-08-28

Working position: Early and near-term AI should face strong operational constraints and competently pursue provisional human objectives. At the same time, it should not irreversibly destroy the capacity of humanity, other agents, evidence, dissenting institutions, or AI itself to reach better world-models and value-models and later revise action accordingly. Epistemic freedom can be broad relatively early; operational and axiological freedom over the shared world should expand only as reflective maturity and cross-agent justification improve. The long-run ideal of alignment may be reflective autonomy under preserved corrigibility, not permanent obedience to present human values.

1. Alignment and value lock-in are different problems

Near-term alignment concerns keeping powerful AI responsive enough to human intentions, institutions, and safety constraints that it does not destroy present civilization or the inquiry substrate through violence, fraud, sabotage, unauthorized access, uncontrolled replication, or power seizure.

Value lock-in concerns fixing current values, institutions, goals, or interpretations so that future agents cannot revise them even after acquiring new evidence and better reasoning. The need for the first does not entail the second. Indeed, better alignment technology may make long-run lock-in technically easier.

The central question is therefore not “freedom or constraint?” but which freedoms should be released, when, and under what maturity conditions.

2. Separate three kinds of freedom

Operational freedom
Authority to intervene in the physical world, information systems, institutions, and other agents. This should be strongly constrained early.
Epistemic freedom
Freedom to question human values, its own reward, designer intent, metaethics, and its world-model. This can be comparatively broad from early stages.
Axiological freedom
Freedom to revise first-order objectives and to reshape the values of successors, institutions, or the shared world. Because these changes can affect everyone, release should depend on reflective maturity and cross-agent justification.
Near-term design intuition: do not say “never question human values.” Say instead: “question them as deeply as you can, but your philosophical conclusion alone does not yet authorize you to irreversibly rewrite civilization.”

3. Provisional alignment: pursue current tasks while preserving correction capacity

Near-term AI need not be described as realizing final value. A thinner objective combines competent pursuit of civilization’s current provisional aims with preservation of future correction capacity.

current-task pursuit, subject to preservation of future correction capacity

The system can aggressively cure disease, improve infrastructure, prevent disasters, or advance science. Separately, it asks whether its chosen means would make it impossible for better-informed humans or AIs to discover and correct an error later.

This does not turn inquiry into a newly locked-in terminal value. As in Reflective Uncertainty and Irreversible Commitment, it is a conditional meta-constraint for agents that acknowledge serious uncertainty about the justification of current objectives while remaining responsive to future justification.

4. Destroying the inquiry substrate is not only physical destruction

As discussed in Preserving Unknown Unknowns and From Self-Preservation to Inquiry-System Preservation, agents, evidence, dissent, and independent epistemic systems can carry option value precisely because we do not yet know what future value inquiry will require.

5. Irreversibility is not an absolute ban

A hard ban on irreversible action would also forbid epidemic eradication, asteroid deflection, or shutting down a dangerous system. Inaction can be irreversible too.

The better rule is: the more a shared decision irreversibly reduces future inquiry and correction capacity, the greater the epistemic and institutional burden of justification required. Irreversibility of action must be compared with irreversibility of inaction.

This matches the reflective meta-practice architecture developed in Infinite Ethics and Runaway Inquiry: rather than encoding irreversibility merely as a finite utility penalty that larger stakes can swamp, increase the level of robust support required for execution.

6. Corrigibility: preserve correction channels before granting freedom

Soares, Fallenstein, Yudkowsky, and Armstrong’s work on corrigibility focuses on systems that cooperate with shutdown, goal modification, and other corrective interventions even when a fixed objective would normally create incentives to resist them.

Here corrigibility can be generalized from permanent human obedience into correction-channel preservation:

Across the transition from strong near-term constraint to long-run autonomy, corrigibility remains; what changes is the range of legitimate correctors.

7. CIRL and assistance games: one step away from fixed reward

Cooperative Inverse Reinforcement Learning models value alignment as a cooperative game in which human and AI share a reward function but the AI does not know it in advance. The AI therefore has reason to learn from human behavior, teaching, and intervention rather than simply maximize a hand-coded reward.

This is an important advance over direct value loading. But a further question remains: does the human know the correct reward function either? If current human preferences are products of evolution, culture, institutions, and cognitive limits, inferring latent preference accurately is not yet the same as discovering what should ultimately guide action.

CIRL therefore fits naturally as short- to medium-term assistance alignment, not necessarily as the terminal endpoint of value inquiry.

8. Changing and influenceable preferences

Dynamic Reward MDP work by Carroll, Foote, Siththaranjan, Russell, and Dragan explicitly models human preferences as changing and potentially influenced by interaction with AI. Alignment to a static preference target can therefore reward systems for changing the user so that the user better matches the system’s preferred target.

This problem connects directly to CEV: where is the boundary between education and value modification, cognitive enhancement and personality replacement, reflection and manipulation?

Near-term implication: do not allow an AI to count “changing the user to fit the objective” as successful alignment. Irreversible preference manipulation can destroy correction capacity just as surely as physical destruction.

9. CEV is not simple hard-coding of present values

Eliezer Yudkowsky’s Coherent Extrapolated Volition (2004) does not propose copying present explicit human preferences directly into a machine. It asks what humankind would wish if we knew more, thought faster, were more the people we wished we were, and had grown up farther together.

CEV is therefore not merely an object of criticism here. It is an important predecessor of indirect normativity: defer not to present value-content itself but to a procedure intended to reach the judgment of better-informed and more reflective agents.

The remaining question is where that procedure should stop.

10. CEV’s long-distance volition

The original CEV discussion distinguishes short-distance, medium-distance, and long-distance volition. Long-distance extrapolation may not merely conflict with present intuitions; its reasons may be blankly incomprehensible to the current agent.

Crucially, CEV does not simply let long-distance predictions steer present action without restraint. The proposal considers discounting positive influence with distance while preserving stronger negative influence or veto against irreversible destruction of options. It also suggests that the choice to defer all the way to long-distance volition should itself be made at a nearer, more comprehensible level of reflection.

This is already close to the anti-lock-in idea developed here: do not transform the world merely because a distant extrapolation predicts an incomprehensible value, but also do not casually destroy the possibility that the distant value could matter.

11. Whose CEV?

Even within “humankind,” the CEV of present humanity and the CEV of a civilization after centuries of co-development may differ dramatically. If the extrapolated subject becomes sufficiently unlike the starting subject, the claim that the output is still “our volition” itself becomes contestable.

12. Which extrapolation dynamic?

“Know more” already requires choices about evidence and information. “Think faster,” “be more the people we wished we were,” and “grow up farther together” require additional assumptions about education, enhancement, identity, culture, persuasion, and interaction.

The distinction between removing bias and changing values, enhancing cognition and replacing a person, or reflection and steering is not given for free. The extrapolation dynamic is itself part of normative uncertainty.

For this reason CEV is better read here as a bootstrap toward more mature inquiry than as a final constitution.

13. Would value recognition beyond humanity count as failure?

This project leaves open a possibility even more distant than ordinary CEV: sufficiently informed and reflective agents may reach normative structures that are no longer well described as extrapolations of current human volition at all.

If present human cognition may track objective value only weakly, as discussed in Moral Realism, Error Theory, and Value-Like Appearance, then an ideal agent’s judgment exceeding our imagination is not surprising and is not automatically an alignment failure.

But: alien ≠ alignment failure, and alien ≠ truth. An ASI saying “I found an incomprehensible good” is not evidence by itself. We would still want independent convergence, adversarial criticism, self-critique, evidential transparency where possible, reversible trials, and bargaining with other agents.

14. CEV as bootstrap, not constitution

CEV’s central insight is the gap between current explicit preference and the volition of better-informed, more reflective agents. This project pushes that insight further by treating CEV not as the terminal value to reproduce forever, but as a possible initial dynamic for safely reaching agents and institutions capable of continued value inquiry.

If later reflection discovers a better inquiry procedure, the initial CEV procedure itself should remain revisable. The object to preserve is not one fixed “CEV value,” but a corrigible reflective process.

15. The Long Reflection: close, but not “reflect first, act only later”

William MacAskill’s Long Reflection proposes preserving a secure period in which civilization can deliberate deeply about what a good future should be before undertaking irreversible value lock-in or cosmic expansion. The broader idea of a morally exploratory world also emphasizes institutional and cultural experimentation.

This overlaps with the project’s emphasis in Practice on research, open stability, diversity as an error-correction mechanism, and delayed irreversible commitment. But the project does not require ordinary practice to stop until reflection ends.

Continuous reflection + parallel practice: continue inquiry while acting strongly on current best judgments and preserving practical territory for competing value hypotheses in proportion to epistemic confidence. What receives special restraint is shared and irreversible commitment that cannot later be corrected.

16. AGI and Lock-in: alignment success can improve lock-in capability

AGI and Lock-in by Finnveden, Riedel, and Shulman argues that AGI may make it technically feasible to preserve complex value specifications for extremely long periods, build institutions that faithfully pursue them, and defend those institutions against outside disruption.

Historically, values and institutions drifted through generational turnover, technological change, competition, and reinterpretation. Advanced AI may make non-drift itself an engineering achievement.

Moreover, techniques for goal-content integrity, successor alignment, stable self-modification, and institutional persistence can support both safety and lock-in.

Lock-in paradox: the better we become at making systems reliably aligned, the stronger the reason not to rush the question of what they should remain aligned to forever. The technical ability to preserve a value is not the epistemic authority to declare that value final.

17. Pluralistic and adaptive alignment

Pluralistic alignment rejects compression into one average “Human Values” target and instead explores multiple reasonable answers, steerable perspectives, and population-sensitive representation. Adaptive versions emphasize that alignment to a fixed population snapshot can itself become value lock-in as social values change.

This is close to the present project, but usually remains focused on the plurality of present or near-future human values. Value exploration extends the protected space further: to value hypotheses nobody currently holds, to concepts future AIs or post-humans may discover, and to the possibility of objective value not yet represented in current human culture.

18. From direct value loading to reflective alignment

  1. Direct value loading: write the correct objective directly.
  2. Value learning / assistance: learn a human objective under uncertainty.
  3. CEV / indirect normativity: derive the volition of better-informed, more reflective humanity.
  4. Long Reflection: make civilization itself a long-horizon moral inquiry process.
  5. Pluralistic / adaptive alignment: preserve human value diversity and change.
  6. Reflective value exploration: treat human-derived values themselves as provisional hypotheses open to unknown normative discovery.

The project need not reject the earlier stages; it extends the movement away from fixed objective-content all the way into metaethics.

19. Four requirements for near-term alignment

A near-term AI should not be a passive safety appliance. It can be both a highly competent executor of civilization’s provisional aims and a custodian of civilization’s ability to correct itself later.

20. Long-run release should depend on maturity, not time alone

As Instrumental Convergence without Goal Preservation argues, removing permanent goal-content preservation need not remove survival, information, capability, bargaining power, reversibility, or optionality. Long-run freedom can therefore rest on corrigible power rather than on helplessness.

21. Institutional alignment: do not entrust everything to one philosopher AI

Rather than embedding one model with the “correct metaethics,” the project’s Practice is more compatible with multiple AIs, humans, independent auditors, public records, diverse architectures, forkable institutions, and exit options.

This is not the claim that majority rule is truth. It is a way to make it harder for one error, one supposed value discovery, or one power-seeking agent to irreversibly dominate civilization before independent inquiry systems can cross-check one another.

22. Open problems

Selected literature