AI Value Exploration Notes
Exploration

Governing Preference-Formation Channels — Second-Order Alignment and Strategic Manipulation

Exploration v0.1 · English translation · 2026-09-26

Working thesis: If preferences, values, and meta-preferences are plastic, normative analysis cannot stop at asking which preferences are correct. It must also ask who controls P_t → P_{t+1}, with what information and authority, through what procedures, and with what degree of reversibility. Alignment, education, persuasion, advertising, pharmacology, BCI, self-modification, and institutional design can be compared as different implementations of preference-formation channels.

1. Treat preferences as transition processes rather than fixed inputs

The plastic value-subject exploration argues that preferences, affect, and value cannot be treated as fixed normative inputs. It examines the temptation to modify evaluators instead of improving the world, and the failure of higher-order preference or idealization to provide an automatically privileged level.

The next step is to analyze not merely preference states but transition channels.

Pt+1 = F(Pt, E, S, O, I)

Here E represents environment, culture, and education; S self-modification; O interventions by others; and I the available informational and epistemic inputs. Under advanced preference engineering, F itself—update rules, edit permissions, learning rates, rollback conditions, and approval requirements—may become a design variable.

Second-order question: If the first-order question is “what should an agent want?”, the second-order question is “what, or who, should have authority to change what an agent wants, under what conditions?”

2. Do not collapse causal history into current value

The fact that a preference was formed by education, culture, evolution, advertising, pharmacology, training, or interaction with AI does not by itself make that preference “inauthentic.” If the same present state can be reached through multiple histories, privileging one as genuine and another as false requires an additional theory.

But it also does not follow that transition history is ethically irrelevant. A coercive personality modification may be wrong even if the resulting subject’s experiences and values are not metaphysically unreal. Conversely, later endorsement of the intervention does not necessarily retroactively justify the transition.

At least three questions should therefore be separated:

In the vocabulary of functional freedom, origin authenticity ≠ present revisability. The important question is not whether history can be erased, but whether the present agent can know, criticize, compare, exit, or roll back its formation process.

3. Self-modification and other-modification are alike and unlike

If an agent chooses its own modification, and if a third party imposes the same internal change, the resulting internal state may be identical while the governance structure differs. Self-modification includes at least the present agent’s initial endorsement, but it raises the diachronic question of how far a present state may bind a future state. Other-modification adds a further asymmetry of authority.

Yet the simple rule “self-chosen = free, external = unfree” is also inadequate. A present agent may decide under misinformation, become dependent in ways that destroy later reconsideration, irreversibly remove the future agent’s exit, or use a self-editing mechanism designed and controlled by someone else.

Preference transitions should therefore be evaluated less by who pressed the button than by where truth access, counterfactual responsiveness, effective options, meta-revision, reversibility, and causal leverage are distributed.

4. Second-order alignment: from alignment-to-what to who controls revision

Alignment and value lock-in distinguishes operational, epistemic, and axiological freedom. But if AI values are plastic, “what should the AI be aligned to?” is still incomplete.

Which of the AI itself, developers, users, states, other AIs, or future versions receives write access to the value-revision channel? Which edits count as ordinary updates and which as constitutional change? Who can erase modification history? Who preserves prior checkpoints?

Second-order alignment: alignment concerns not only behavior or present values, but also control rights over value-revision channels, verifiability, reversibility, and inheritance rules.

This does not imply that AI should never self-modify. On the contrary, as reflective self-modification matures, permanent monopoly by external operators over value-editing rights may itself become a form of lock-in. Strong asymmetric control for short-term safety and reflective autonomy for the long term may need to be separated temporally.

5. The strategicization of preference formation

If multiple agents can alter one another’s preferences, the game no longer consists only in bargaining over actions under fixed utilities. Changing the other side’s objective function becomes a strategy.

Ordinary persuasion addresses another agent’s current reasons and preferences. Strong preference manipulation can instead directly alter what the other agent treats as a reason, what it terminally wants, or which evidence sources it trusts. The contested object becomes not only action space but evaluation architecture.

In the vocabulary of Hobbesian contractarian morality, this is a “second-order state of nature.” Distrust expands from “will the other party defect?” to “will the other party transform me into an agent who wants to defect, obey, or cease resisting?”

6. What changes when strategic preference formation becomes common knowledge?

Now suppose all agents understand that moral discourse, education, public reason, and psychological intervention can also function as preference-formation strategies. This meta-awareness need not destroy morality, but it raises the price of simple interpersonal trust.

Agents no longer ask only whether someone sincerely believes in justice, but also why that belief is stable, who can change it, and what can be verified if it changes. Some moral trust may shift toward verifiable commitment:

This is not a simple replacement of trust with technology. Auditors and institutions can themselves be manipulated. The deeper shift is to avoid concentrating trust in a single inner state and instead distribute it across multiple evidential and falsifiable procedures.

7. Preference homogenization may become a tempting equilibrium

In a world where value differences themselves generate conflict, defection, or security risks, modifying all agents toward similar values can appear to be an efficient stabilization strategy. If technologically easy, changing an opponent into an agent who no longer requires bargaining may be cheaper than bargaining with them.

But this trades apparent order for the simultaneous removal of independent criticism, value exploration, and error-correction pathways. Even unanimous support after homogenization does not settle the issue if that support was manufactured by the very regime being evaluated.

Preference homogenization is therefore not logically forbidden in every conceivable case, but under deep value uncertainty, broad and irreversible homogenization should bear an especially high burden of justification.

8. Constrain capabilities, permissions, and interfaces before homogenizing values

The meta-goal community explores sharing a thin constitutional interface—evidence exchange, bargaining, exit, dispute resolution—without requiring first-order value agreement. The manipulability of preference formation makes this distinction more important.

Pluralism-preserving default: where feasible, manage externalities arising from value differences through capability constraints, access control, verifiability, boundaries, exit, and dispute resolution rather than through homogenization of value content.

Sandboxing dangerous capabilities is not the same as rewriting an agent so that it cannot have dangerous thoughts. The former can restrict operational power while preserving internal epistemic and axiological exploration. Complete separation may not always be possible, but interface-level solutions should be considered before eliminating difference itself.

9. Candidate institutions for transition governance

Closing preference-formation channels entirely would also block learning, therapy, self-improvement, and value discovery. Leaving them completely open would invite manipulation by others and destructive self-editing. The target is therefore not prohibition but layered friction and preserved correction channels.

This architecture resembles the controlled plasticity proposed in functional freedom: distinguish being able to change from being easily rewritten by someone else.

10. Authority over future selves is not self-evident either

Preference-formation governance is not only interpersonal. A present agent that irreversibly fixes the preferences of its future self is also exercising one-way governance. Ordinary future selves stand in unusually dense successor relations to present selves, but this does not automatically grant unlimited constitutional authority to the present state over all future states.

Conversely, if future selves never respect present commitments, long-term planning and promising become difficult. Diachronic agency therefore requires balancing present self-binding against future revision rights.

This connects to the successor relations introduced in Inheritance Systems and Variable Individuality. Strong succession may generate reasons to preserve commitments without implying permanent preference lock-in.

11. The problem remains even if normative truth exists

If objective and knowable normative truths exist, and agents share a meta-preference to discover and follow them, preference formation may become cooperative inquiry rather than merely a contest over sovereignty. Agents may exchange evidence and arguments and converge on the same reasons.

Yet “moving agents closer to truth” can itself become a justification used by a manipulator. Normative realism therefore does not remove the need for transition governance. Independent epistemic systems, contestability, exit, and convergence across multiple routes become even more important.

12. What does not follow

13. What would weaken this framework?

Sources / notes

On endogenous and adaptive preferences, see Jon Elster, Sour Grapes (1983). On higher-order preference and personhood, Harry Frankfurt (1971). On transformative self-formation, L. A. Paul, Transformative Experience (2014), and Agnes Callard, Aspiration (2018). On AI systems interacting with changing and influenceable human reward functions, see Micah Carroll et al. and related Dynamic Reward MDP work. For strategic-rationality connections, see Hobbes, Gauthier, Moehler, and this project’s Hobbesian contractarian morality. The integrated notions of transition governance, second-order alignment, common-knowledge preference manipulation, and a pluralism-preserving default are provisional syntheses developed in this project.