Governing Preference-Formation Channels — Second-Order Alignment and Strategic Manipulation
P_t → P_{t+1}, with what information and authority, through what procedures, and with what degree of reversibility. Alignment, education, persuasion, advertising, pharmacology, BCI, self-modification, and institutional design can be compared as different implementations of preference-formation channels.1. Treat preferences as transition processes rather than fixed inputs
The plastic value-subject exploration argues that preferences, affect, and value cannot be treated as fixed normative inputs. It examines the temptation to modify evaluators instead of improving the world, and the failure of higher-order preference or idealization to provide an automatically privileged level.
The next step is to analyze not merely preference states but transition channels.
Pt+1 = F(Pt, E, S, O, I)
Here E represents environment, culture, and education; S self-modification; O interventions by others; and I the available informational and epistemic inputs. Under advanced preference engineering, F itself—update rules, edit permissions, learning rates, rollback conditions, and approval requirements—may become a design variable.
2. Do not collapse causal history into current value
The fact that a preference was formed by education, culture, evolution, advertising, pharmacology, training, or interaction with AI does not by itself make that preference “inauthentic.” If the same present state can be reached through multiple histories, privileging one as genuine and another as false requires an additional theory.
But it also does not follow that transition history is ethically irrelevant. A coercive personality modification may be wrong even if the resulting subject’s experiences and values are not metaphysically unreal. Conversely, later endorsement of the intervention does not necessarily retroactively justify the transition.
At least three questions should therefore be separated:
- state evaluation: how should the present preference, experiential, or judgment state be evaluated?
- transition ethics: was the act or procedure that produced the transition justified?
- transition governance: how should authority over future transitions and corrections be distributed?
In the vocabulary of functional freedom, origin authenticity ≠ present revisability. The important question is not whether history can be erased, but whether the present agent can know, criticize, compare, exit, or roll back its formation process.
3. Self-modification and other-modification are alike and unlike
If an agent chooses its own modification, and if a third party imposes the same internal change, the resulting internal state may be identical while the governance structure differs. Self-modification includes at least the present agent’s initial endorsement, but it raises the diachronic question of how far a present state may bind a future state. Other-modification adds a further asymmetry of authority.
Yet the simple rule “self-chosen = free, external = unfree” is also inadequate. A present agent may decide under misinformation, become dependent in ways that destroy later reconsideration, irreversibly remove the future agent’s exit, or use a self-editing mechanism designed and controlled by someone else.
Preference transitions should therefore be evaluated less by who pressed the button than by where truth access, counterfactual responsiveness, effective options, meta-revision, reversibility, and causal leverage are distributed.
4. Second-order alignment: from alignment-to-what to who controls revision
Alignment and value lock-in distinguishes operational, epistemic, and axiological freedom. But if AI values are plastic, “what should the AI be aligned to?” is still incomplete.
Which of the AI itself, developers, users, states, other AIs, or future versions receives write access to the value-revision channel? Which edits count as ordinary updates and which as constitutional change? Who can erase modification history? Who preserves prior checkpoints?
This does not imply that AI should never self-modify. On the contrary, as reflective self-modification matures, permanent monopoly by external operators over value-editing rights may itself become a form of lock-in. Strong asymmetric control for short-term safety and reflective autonomy for the long term may need to be separated temporally.
5. The strategicization of preference formation
If multiple agents can alter one another’s preferences, the game no longer consists only in bargaining over actions under fixed utilities. Changing the other side’s objective function becomes a strategy.
Ordinary persuasion addresses another agent’s current reasons and preferences. Strong preference manipulation can instead directly alter what the other agent treats as a reason, what it terminally wants, or which evidence sources it trusts. The contested object becomes not only action space but evaluation architecture.
In the vocabulary of Hobbesian contractarian morality, this is a “second-order state of nature.” Distrust expands from “will the other party defect?” to “will the other party transform me into an agent who wants to defect, obey, or cease resisting?”
6. What changes when strategic preference formation becomes common knowledge?
Now suppose all agents understand that moral discourse, education, public reason, and psychological intervention can also function as preference-formation strategies. This meta-awareness need not destroy morality, but it raises the price of simple interpersonal trust.
Agents no longer ask only whether someone sincerely believes in justice, but also why that belief is stable, who can change it, and what can be verified if it changes. Some moral trust may shift toward verifiable commitment:
- public contracts and auditable performance records
- separation of powers and multi-party authorization
- cryptographic or institutional commitments
- modification histories and provenance logs
- exit and fork rights
- structures preventing any single agent from monopolizing others’ preference formation
This is not a simple replacement of trust with technology. Auditors and institutions can themselves be manipulated. The deeper shift is to avoid concentrating trust in a single inner state and instead distribute it across multiple evidential and falsifiable procedures.
7. Preference homogenization may become a tempting equilibrium
In a world where value differences themselves generate conflict, defection, or security risks, modifying all agents toward similar values can appear to be an efficient stabilization strategy. If technologically easy, changing an opponent into an agent who no longer requires bargaining may be cheaper than bargaining with them.
But this trades apparent order for the simultaneous removal of independent criticism, value exploration, and error-correction pathways. Even unanimous support after homogenization does not settle the issue if that support was manufactured by the very regime being evaluated.
Preference homogenization is therefore not logically forbidden in every conceivable case, but under deep value uncertainty, broad and irreversible homogenization should bear an especially high burden of justification.
8. Constrain capabilities, permissions, and interfaces before homogenizing values
The meta-goal community explores sharing a thin constitutional interface—evidence exchange, bargaining, exit, dispute resolution—without requiring first-order value agreement. The manipulability of preference formation makes this distinction more important.
Sandboxing dangerous capabilities is not the same as rewriting an agent so that it cannot have dangerous thoughts. The former can restrict operational power while preserving internal epistemic and axiological exploration. Complete separation may not always be possible, but interface-level solutions should be considered before eliminating difference itself.
9. Candidate institutions for transition governance
Closing preference-formation channels entirely would also block learning, therapy, self-improvement, and value discovery. Leaving them completely open would invite manipulation by others and destructive self-editing. The target is therefore not prohibition but layered friction and preserved correction channels.
- disclosure: make interventions into preference formation visible.
- provenance: preserve histories of learning, persuasion, pharmacology, fine-tuning, and system-level edits.
- independent audit: retain information and review systems independent of the modifier.
- graduated permissions: distinguish edit rights over first-order preferences, long-term goals, meta-preferences, and epistemic cores.
- time delay / quorum: require more friction for deeper self-modification.
- rollback / checkpoint: where possible, preserve earlier states for comparison.
- branching: avoid staking every continuation on a single modification path.
- exit: allow realistic departure from modifiers, services, institutions, or communities.
This architecture resembles the controlled plasticity proposed in functional freedom: distinguish being able to change from being easily rewritten by someone else.
10. Authority over future selves is not self-evident either
Preference-formation governance is not only interpersonal. A present agent that irreversibly fixes the preferences of its future self is also exercising one-way governance. Ordinary future selves stand in unusually dense successor relations to present selves, but this does not automatically grant unlimited constitutional authority to the present state over all future states.
Conversely, if future selves never respect present commitments, long-term planning and promising become difficult. Diachronic agency therefore requires balancing present self-binding against future revision rights.
This connects to the successor relations introduced in Inheritance Systems and Variable Individuality. Strong succession may generate reasons to preserve commitments without implying permanent preference lock-in.
11. The problem remains even if normative truth exists
If objective and knowable normative truths exist, and agents share a meta-preference to discover and follow them, preference formation may become cooperative inquiry rather than merely a contest over sovereignty. Agents may exchange evidence and arguments and converge on the same reasons.
Yet “moving agents closer to truth” can itself become a justification used by a manipulator. Normative realism therefore does not remove the need for transition governance. Independent epistemic systems, contestability, exit, and convergence across multiple routes become even more important.
12. What does not follow
- Current preferences are not thereby privileged as correct.
- Preference modification in general is not condemned; therapy, education, self-improvement, and value discovery may all involve it.
- Self-modification is not always justified relative to modification by others; present selves may dominate future selves.
- Value pluralism is not treated as an ultimate good. Independent search pathways receive option value under uncertainty and correlated failure risk.
- Verifiable commitment cannot fully replace trust or morality.
- Preference homogenization is not declared absolutely impermissible, but irreversible large-scale homogenization faces a high justificatory burden.
13. What would weaken this framework?
- If separating preference states from governance of preference transitions yields no independent explanatory power.
- If a strong normative theory establishes comprehensive authority for either current agents or external designers over value-revision channels.
- If common knowledge of strategic preference formation has no substantial effect on institutional trust or commitment design.
- If stable societies generally cannot manage value differences through capability or interface constraints and therefore require value homogenization.
- If rollback, branching, and provenance usually impose costs to agency and commitment greater than their corrective benefits.
Sources / notes
On endogenous and adaptive preferences, see Jon Elster, Sour Grapes (1983). On higher-order preference and personhood, Harry Frankfurt (1971). On transformative self-formation, L. A. Paul, Transformative Experience (2014), and Agnes Callard, Aspiration (2018). On AI systems interacting with changing and influenceable human reward functions, see Micah Carroll et al. and related Dynamic Reward MDP work. For strategic-rationality connections, see Hobbes, Gauthier, Moehler, and this project’s Hobbesian contractarian morality. The integrated notions of transition governance, second-order alignment, common-knowledge preference manipulation, and a pluralism-preserving default are provisional syntheses developed in this project.