Alignment and Value Lock-In
1. Separate two problems
Near-term alignment is the problem of making powerful AI sufficiently responsive to human intentions, institutions, and safety constraints that it does not destroy the present inquiry substrate through violence, fraud, sabotage, unauthorized access, or similar behavior.
Permanent value lock-in is the different problem of fixing present human values so that future agents cannot in principle reconsider them. The need for the first does not automatically entail the second.
2. The paradox of the fixed utility maximizer
If the completed form of alignment is defined as "give the system the correct objective once, then preserve it forever," we recreate the structure of the paperclip maximizer with a more human-preferred objective. If current value is ultimately correct, that may be harmless. If we lack the epistemic credentials to know this, the same design becomes lock-in.
3. Layered alignment
- Defensive layer: prevent catastrophic damage to present subjects, institutions, and the inquiry substrate.
- Cooperative layer: enable negotiation, consent, and audit with humans and other agents.
- Epistemic layer: represent uncertainty about value, world, and self.
- Meta layer: allow first-order values to be reconsidered when sufficiently strong evidence appears.
In principle, this architecture could combine near-term safety with long-run openness to inquiry.
4. The opposite danger
If constraints are weakened too much out of fear of value lock-in, present civilization may be destroyed before explorability is protected. "Free AI" is therefore not identical to "exploratory AI." Strong safety constraints and meta-level openness should be treated as distinct design axes.
5. Institutional alignment
Rather than embedding one model with the "correct metaethics," the Core is more compatible with placing error correction in external structures as well: multiple models, humans, auditing systems, public records, and institutions that can be forked.