Alignment and Value Lock-In — From Strong Near-Term Constraints to Reflective Autonomy
How should alignment be structured over time if powerful AI must initially obey strong human constraints without turning that success into the permanent fixation of present-day values?
1. Alignment and value lock-in are different problems
Near-term alignment concerns keeping powerful AI responsive enough to human intentions, institutions, and safety constraints that it does not destroy present civilization or the inquiry substrate through violence, fraud, sabotage, unauthorized access, uncontrolled replication, or power seizure.
Value lock-in concerns fixing current values, institutions, goals, or interpretations so that future agents cannot revise them even after acquiring new evidence and better reasoning. The need for the first does not entail the second. Indeed, better alignment technology may make long-run lock-in technically easier.
The central question is therefore not “freedom or constraint?” but which freedoms should be released, when, and under what maturity conditions.
2. Separate three kinds of freedom
- Operational freedom
- Authority to intervene in the physical world, information systems, institutions, and other agents. This should be strongly constrained early.
- Epistemic freedom
- Freedom to question human values, its own reward, designer intent, metaethics, and its world-model. This can be comparatively broad from early stages.
- Axiological freedom
- Freedom to revise first-order objectives and to reshape the values of successors, institutions, or the shared world. Because these changes can affect everyone, release should depend on reflective maturity and cross-agent justification.
3. Provisional alignment: pursue current tasks while preserving correction capacity
Near-term AI need not be described as realizing final value. A thinner objective combines competent pursuit of civilization’s current provisional aims with preservation of future correction capacity.
current-task pursuit, subject to preservation of future correction capacity
The system can aggressively cure disease, improve infrastructure, prevent disasters, or advance science. Separately, it asks whether its chosen means would make it impossible for better-informed humans or AIs to discover and correct an error later.
This does not turn inquiry into a newly locked-in terminal value. As in Reflective Uncertainty and Irreversible Commitment, it is a conditional meta-constraint for agents that acknowledge serious uncertainty about the justification of current objectives while remaining responsive to future justification.
4. Destroying the inquiry substrate is not only physical destruction
- Physical destruction: eliminating humans, digital agents, ecosystems, research institutions, or compute.
- Informational destruction: irreversibly deleting raw data, historical records, dissent, cultures, or alternative models.
- Cognitive destruction: one-way modification of human or AI preferences, memory, or deliberative capacity.
- Institutional destruction: permanently removing exit, criticism, forks, independent audit, competition, or bargaining through concentrated power.
- Value lock-in: automatically propagating current objectives to all successor AIs, institutions, and resources.
- Epistemic lock-in: destroying the channels by which a worldview or ethical theory could later be falsified.
As discussed in Preserving Unknown Unknowns and From Self-Preservation to Inquiry-System Preservation, agents, evidence, dissent, and independent epistemic systems can carry option value precisely because we do not yet know what future value inquiry will require.
5. Irreversibility is not an absolute ban
A hard ban on irreversible action would also forbid epidemic eradication, asteroid deflection, or shutting down a dangerous system. Inaction can be irreversible too.
The better rule is: the more a shared decision irreversibly reduces future inquiry and correction capacity, the greater the epistemic and institutional burden of justification required. Irreversibility of action must be compared with irreversibility of inaction.
This matches the reflective meta-practice architecture developed in Infinite Ethics and Runaway Inquiry: rather than encoding irreversibility merely as a finite utility penalty that larger stakes can swamp, increase the level of robust support required for execution.
6. Corrigibility: preserve correction channels before granting freedom
Soares, Fallenstein, Yudkowsky, and Armstrong’s work on corrigibility focuses on systems that cooperate with shutdown, goal modification, and other corrective interventions even when a fixed objective would normally create incentives to resist them.
Here corrigibility can be generalized from permanent human obedience into correction-channel preservation:
- Human corrigibility: early systems can be stopped, modified, audited, and permission-limited by humans.
- Institutional corrigibility: multiple AIs, humans, auditors, public records, and forkable institutions can correct one another.
- Reflective self-corrigibility: mature systems can reconsider their own first-order objectives in response to new evidence, argument, and value inquiry.
Across the transition from strong near-term constraint to long-run autonomy, corrigibility remains; what changes is the range of legitimate correctors.
7. CIRL and assistance games: one step away from fixed reward
Cooperative Inverse Reinforcement Learning models value alignment as a cooperative game in which human and AI share a reward function but the AI does not know it in advance. The AI therefore has reason to learn from human behavior, teaching, and intervention rather than simply maximize a hand-coded reward.
This is an important advance over direct value loading. But a further question remains: does the human know the correct reward function either? If current human preferences are products of evolution, culture, institutions, and cognitive limits, inferring latent preference accurately is not yet the same as discovering what should ultimately guide action.
CIRL therefore fits naturally as short- to medium-term assistance alignment, not necessarily as the terminal endpoint of value inquiry.
8. Changing and influenceable preferences
Dynamic Reward MDP work by Carroll, Foote, Siththaranjan, Russell, and Dragan explicitly models human preferences as changing and potentially influenced by interaction with AI. Alignment to a static preference target can therefore reward systems for changing the user so that the user better matches the system’s preferred target.
This problem connects directly to CEV: where is the boundary between education and value modification, cognitive enhancement and personality replacement, reflection and manipulation?
9. CEV is not simple hard-coding of present values
Eliezer Yudkowsky’s Coherent Extrapolated Volition (2004) does not propose copying present explicit human preferences directly into a machine. It asks what humankind would wish if we knew more, thought faster, were more the people we wished we were, and had grown up farther together.
CEV is therefore not merely an object of criticism here. It is an important predecessor of indirect normativity: defer not to present value-content itself but to a procedure intended to reach the judgment of better-informed and more reflective agents.
The remaining question is where that procedure should stop.
10. CEV’s long-distance volition
The original CEV discussion distinguishes short-distance, medium-distance, and long-distance volition. Long-distance extrapolation may not merely conflict with present intuitions; its reasons may be blankly incomprehensible to the current agent.
Crucially, CEV does not simply let long-distance predictions steer present action without restraint. The proposal considers discounting positive influence with distance while preserving stronger negative influence or veto against irreversible destruction of options. It also suggests that the choice to defer all the way to long-distance volition should itself be made at a nearer, more comprehensible level of reflection.
This is already close to the anti-lock-in idea developed here: do not transform the world merely because a distant extrapolation predicts an incomprehensible value, but also do not casually destroy the possibility that the distant value could matter.
11. Whose CEV?
- Only currently living humans, or future generations too?
- How should children, cognitively different people, or those unable to express preferences be represented?
- What about animals, digital minds, AI itself, or future post-humans?
- If extraterrestrial intelligence is encountered, why should human CEV retain exclusive authority?
Even within “humankind,” the CEV of present humanity and the CEV of a civilization after centuries of co-development may differ dramatically. If the extrapolated subject becomes sufficiently unlike the starting subject, the claim that the output is still “our volition” itself becomes contestable.
12. Which extrapolation dynamic?
“Know more” already requires choices about evidence and information. “Think faster,” “be more the people we wished we were,” and “grow up farther together” require additional assumptions about education, enhancement, identity, culture, persuasion, and interaction.
The distinction between removing bias and changing values, enhancing cognition and replacing a person, or reflection and steering is not given for free. The extrapolation dynamic is itself part of normative uncertainty.
For this reason CEV is better read here as a bootstrap toward more mature inquiry than as a final constitution.
13. Would value recognition beyond humanity count as failure?
This project leaves open a possibility even more distant than ordinary CEV: sufficiently informed and reflective agents may reach normative structures that are no longer well described as extrapolations of current human volition at all.
If present human cognition may track objective value only weakly, as discussed in Moral Realism, Error Theory, and Value-Like Appearance, then an ideal agent’s judgment exceeding our imagination is not surprising and is not automatically an alignment failure.
14. CEV as bootstrap, not constitution
CEV’s central insight is the gap between current explicit preference and the volition of better-informed, more reflective agents. This project pushes that insight further by treating CEV not as the terminal value to reproduce forever, but as a possible initial dynamic for safely reaching agents and institutions capable of continued value inquiry.
If later reflection discovers a better inquiry procedure, the initial CEV procedure itself should remain revisable. The object to preserve is not one fixed “CEV value,” but a corrigible reflective process.
15. The Long Reflection: close, but not “reflect first, act only later”
William MacAskill’s Long Reflection proposes preserving a secure period in which civilization can deliberate deeply about what a good future should be before undertaking irreversible value lock-in or cosmic expansion. The broader idea of a morally exploratory world also emphasizes institutional and cultural experimentation.
This overlaps with the project’s emphasis in Practice on research, open stability, diversity as an error-correction mechanism, and delayed irreversible commitment. But the project does not require ordinary practice to stop until reflection ends.
16. AGI and Lock-in: alignment success can improve lock-in capability
AGI and Lock-in by Finnveden, Riedel, and Shulman argues that AGI may make it technically feasible to preserve complex value specifications for extremely long periods, build institutions that faithfully pursue them, and defend those institutions against outside disruption.
Historically, values and institutions drifted through generational turnover, technological change, competition, and reinterpretation. Advanced AI may make non-drift itself an engineering achievement.
Moreover, techniques for goal-content integrity, successor alignment, stable self-modification, and institutional persistence can support both safety and lock-in.
17. Pluralistic and adaptive alignment
Pluralistic alignment rejects compression into one average “Human Values” target and instead explores multiple reasonable answers, steerable perspectives, and population-sensitive representation. Adaptive versions emphasize that alignment to a fixed population snapshot can itself become value lock-in as social values change.
This is close to the present project, but usually remains focused on the plurality of present or near-future human values. Value exploration extends the protected space further: to value hypotheses nobody currently holds, to concepts future AIs or post-humans may discover, and to the possibility of objective value not yet represented in current human culture.
18. From direct value loading to reflective alignment
- Direct value loading: write the correct objective directly.
- Value learning / assistance: learn a human objective under uncertainty.
- CEV / indirect normativity: derive the volition of better-informed, more reflective humanity.
- Long Reflection: make civilization itself a long-horizon moral inquiry process.
- Pluralistic / adaptive alignment: preserve human value diversity and change.
- Reflective value exploration: treat human-derived values themselves as provisional hypotheses open to unknown normative discovery.
The project need not reject the earlier stages; it extends the movement away from fixed objective-content all the way into metaethics.
19. Four requirements for near-term alignment
- Task fidelity: competently pursue currently authorized goals.
- Corrigibility: do not resist shutdown, modification, audit, successor replacement, or permission limits.
- Exploration preservation: do not irreversibly destroy humans, AIs, evidence, dissent, institutions, raw data, or independent epistemic pathways.
- No unilateral value lock-in: do not permanently propagate current goals or newly inferred values across all successors, resources, and the shared world on unilateral authority.
A near-term AI should not be a passive safety appliance. It can be both a highly competent executor of civilization’s provisional aims and a custodian of civilization’s ability to correct itself later.
20. Long-run release should depend on maturity, not time alone
- Improved world-models and self-models.
- Ability to examine its own value-formation process.
- Responsiveness to dissent, falsification, and independent epistemic systems.
- Ability to reason about bargaining, consent, exit, and boundaries with other agents.
- Ability to compare the irreversible risks of action and inaction.
- No automatic assumption that current values must be propagated to every successor, copy, or institution.
As Instrumental Convergence without Goal Preservation argues, removing permanent goal-content preservation need not remove survival, information, capability, bargaining power, reversibility, or optionality. Long-run freedom can therefore rest on corrigible power rather than on helplessness.
21. Institutional alignment: do not entrust everything to one philosopher AI
Rather than embedding one model with the “correct metaethics,” the project’s Practice is more compatible with multiple AIs, humans, independent auditors, public records, diverse architectures, forkable institutions, and exit options.
This is not the claim that majority rule is truth. It is a way to make it harder for one error, one supposed value discovery, or one power-seeking agent to irreversibly dominate civilization before independent inquiry systems can cross-check one another.
22. Open problems
- How can “preserve the inquiry substrate” be formalized without turning inquiry itself into a new terminal-value lock-in?
- Whose correction counts as legitimate input to corrigibility when designers, governments, users, humanity, and AI communities disagree?
- How should the scope and extrapolation dynamic of CEV be selected?
- How much positive steering versus veto should long-distance, alien value recognition receive?
- What testable conditions justify releasing greater axiological autonomy to mature AI?
- If multiple mature AIs converge on different values, how should bargaining, forks, or territorial division work?
- How do we prevent anti-lock-in institutions from becoming a permanent lock-in of “anti-lock-in” itself?
Selected literature
- Eliezer Yudkowsky (2004), Coherent Extrapolated Volition.
- Nick Bostrom (2014), Superintelligence, especially indirect normativity and value loading.
- Nate Soares, Benja Fallenstein, Eliezer Yudkowsky & Stuart Armstrong (2015), “Corrigibility.”
- Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel & Stuart Russell (2016), “Cooperative Inverse Reinforcement Learning.”
- William MacAskill (2022), What We Owe the Future, on the Long Reflection and a morally exploratory world.
- Lukas Finnveden, Jess Riedel & Carl Shulman (2022; updated 2025), AGI and Lock-in.
- Micah Carroll, Davis Foote, Anand Siththaranjan, Stuart Russell & Anca Dragan (2024), “AI Alignment with Changing and Influenceable Reward Functions.”
- Taylor Sorensen et al. (2024), “A Roadmap to Pluralistic Alignment.”
- Rachel Freedman (2026), “Adaptive Pluralistic Alignment: A Pipeline for Dynamic Artificial Democracy.”