When Preferences and Affects Become Editable — A Philosophy of Plastic Value Agents
1. Preference modification does not begin in the future
Human preferences and affects are already plastic. Development, education, culture, advertising, religion, relationships, trauma, psychotherapy, pharmacology, and neural stimulation continuously reshape them. Future preference engineering would therefore not create plasticity from nothing. It would make processes that are now slow, indirect, and probabilistic faster, more direct, more precise, and more selectable.
The shift can be represented as environment → probabilistic preference change becoming increasingly like desired preference → intervention → target preference. BCI, closed-loop neuromodulation, pharmacological intervention, and eventually more powerful biological interventions are early points on this continuum.
2. AI may be an even more plastic class of value agent
AI systems can already be altered at several layers: weights, fine-tuning, reward, preference optimization, system prompts, context, memory, and activation steering. Present LLMs need not be assumed to possess stable human-like preferences. But persistent or self-modifying agents could make At → modify(At) → At+1 an ordinary operation.
Digital agents may also permit branching: A → {A1, A2, ...}. Humans and AI can therefore be compared along a spectrum from socially plastic humans to neurally editable humans, digitally configurable agents, and self-modifying AI.
3. Improving the world or modifying the evaluator
Let W be a world state and S an evaluator state, with an ethical theory assigning V(W,S). Traditional moral decision-making usually treats W as the main variable. Once S is also controllable, world improvement competes with evaluator modification.
4. Preference utilitarianism — One-Second Sour Grapes
Alice strongly dislikes her current world. Satisfying her preferences by changing the world would take thirty years and enormous resources. A one-second intervention could instead make her stably prefer the world exactly as it is. Her memory, intelligence, and factual beliefs remain intact, and post-edit Alice fully endorses the edit.
If preference satisfaction is ultimate welfare, the theory has difficulty rejecting the cheaper route of fitting preferences to the world rather than the world to preferences. This is deeper than an experience-machine problem: the evaluator itself has been changed.
Elster's work on adaptive preferences and sour grapes is an important predecessor. Preference engineering turns passive adaptation into deliberate design.
5. Higher-order preferences do not end the regress
Appealing to second-order preferences — what desires an agent wants to have — does not settle the problem if second-order preferences are themselves editable. The same applies at higher levels. Full meta-editability creates a privileged-level problem: which level gets constitutional authority, and why?
This exploration also separates the value of a present state from the ethics of the transition that produced it. If two agents have genuinely identical internal states, a different causal history need not make one set of present values metaphysically fake. A coercive transition can be wrong without making the resulting state unreal.
6. Idealization does not purify values
One escape from actual-preference theories is to use the preferences of a fully informed, reflective, rationalized self. But epistemic idealization and value correction are distinct.
First rewrite Alice into a pure paperclip maximizer. Then give her complete information, near-perfect reasoning, perfect self-knowledge, and unlimited deliberation. The result may simply be an extraordinarily intelligent paperclip maximizer.
K → Kmax does not imply V → Vcorrect. Two agents with different starting value architectures may remain different after identical idealization. An ideal-preference theory therefore still needs an account of which starting evaluative architecture is privileged.
7. The key issue may be corrigibility during transition, not speed itself
Humans already change continuously: S0 → S1 → S2 ... . The ethically relevant difference between a twenty-year A→B transition and a one-second A→B transition may not be speed as such, but the loss of feedback, reconsideration, interruption, partial rollback, learning, and exit opportunities. This shifts attention from sharp personal-identity boundaries to corrigibility during transformation.
8. Contractualism — editing complaints, reasons, and standing
Scanlonian contractualism is more resistant than preference utilitarianism because it asks what principles people have reason to reject rather than merely what they happen to want. Plasticity nevertheless reappears at the level of reasons, standing, and the identity of the parties.
Complaint Laundering
Alice has decisive reason to reject principle P. Replace P with P': first modify Alice into a person who endorses P, then apply P. Can a complaint be normatively erased by changing the complainant?
Born Aligned Society
Imagine a caste society whose lower-status citizens are designed from birth to love obedience, reject autonomy, and sincerely regard the hierarchy as just while retaining excellent information and reasoning. If contractualism still grants them a reason to preserve the ability to reassess their values, that reason appears to enter the contractual procedure independently rather than being generated by it.
Paperclip Contractor and Forked Contractor
If an ideally informed paperclip-maximizing agent rejects any principle that obstructs paperclip production, is that an admissible reason? If yes, admissible reasons depend on arbitrarily editable value architecture; if no, an independent theory of admissible reasons is required. Digital branching adds a further problem: if A is forked into A0, who rejects P, and A1, who endorses it, deleting only A0 looks like standing laundering.
9. Sentimentalism — does an affectless agent fall outside morality?
Emotivism and expressivism should be distinguished from stronger forms of sentimentalism. Modern expressivism can appeal to plans, commitments, or norm acceptance rather than literal emotion. An affectless agent is therefore a sharper test for views on which emotion or sentiment is constitutive of moral judgment or moral knowledge.
The Affectless Judge
Dana has no pleasure or pain, valence, empathy, anger, guilt, disgust, approval, or disapproval. Yet Dana can model other agents, reason, form policies, accept norms, and revise those policies in response to reasons. Dana argues that torture is wrong and implements institutions that prohibit it.
A strong judgment sentimentalist may have to say that Dana performs sophisticated normative reasoning while never making a genuine moral judgment.
Affective Toggle
Hold an agent's beliefs, memory, reasoning, and policies fixed while switching affect ON and OFF. At 11:59, “torture is wrong” counts as a moral judgment; at 12:01, the identical proposition, reasons, and policy allegedly do not. This isolates the constitutive claim about affect.
Neo-sentimentalist or fitting-attitude views can instead distinguish having an emotion from judging that the emotion is fitting. But then the question “why is that emotion fitting?” pushes normative work toward reasons or fittingness. The more easily a theory accommodates the affectless judge, the more affect risks moving from the ground of normativity to an object of normative assessment.
10. Affective architectures can be exchanged before idealization
Create EA, which reacts to others' suffering with aversion, and EB, which reacts with pleasure. Give both the same complete information and reasoning capacity. If EA + Kmax → MA and EB + Kmax → MB, epistemic idealization alone does not identify the correct affective architecture. Sentimentalism then also requires a meta-standard for which sentiments are appropriate.
11. Moral bioenhancement as human alignment
Persson and Savulescu's moral-bioenhancement program proposes engineering traits such as altruism, sensitivity to justice, impulse control, and moral reasoning. Abstractly, this is a form of human alignment: altering motivational architecture toward a normatively preferred direction.
But alignment always requires an answer to alignment to what? “Altruism” and “justice” are broad labels; they do not determine beneficiaries, principles of justice, or acceptable limits on self-sacrifice. Moral bioenhancement is therefore also a problem of which variables, if any, are robust enough to modify before normative theory is settled.
12. Preference inertia ceases to be a natural default
Humans currently have considerable causal inertia: yesterday's preferences tend to persist into today. If editing becomes cheap, Pt+1 = F(Pt, self-edit, other-edit, environment). Keeping one's current values becomes an active policy rather than an unmarked default.
Preference persistence may therefore need to be reconstructed through anti-tampering rules, checkpoints, rollback, permission systems, edit histories, and preserved branches.
13. Singleton and anarchy — two political regimes of preference editing
Singleton: preference dictatorship
A central controller capable of editing everyone need not merely coerce dissenters; it can eliminate dissent as a psychological state. If every citizen later reports satisfaction and support, present endorsement becomes weak evidence of legitimacy when the regime itself manufactured the endorsers.
Anarchy: preference-modification competition
In a decentralized world, agents may compete not only to influence actions but to capture one another's evaluative functions. Firms may prefer to manufacture terminal brand loyalty rather than persuade customers; political movements may prefer to install loyalty rather than win arguments. Competition among modifiers can become competitive manipulation, not freedom.
14. Value ecologies may select transmissibility rather than truth
Let a value system V have content C, self-preservation strength R, and transmissibility I. Under editing competition, long-run fitness may depend on f(R, I, resource acquisition). Systems that resist modification and convert others may outcompete more epistemically justified but less self-protective values.
A value package instructing its host to preserve it permanently, block destabilizing evidence, convert others, and acquire resources for these ends could have high memetic fitness independently of its normative truth.
15. Goal preservation need not be an essence of intelligence
Bostrom's goal-content integrity and formal work on self-modifying agents show why an agent that evaluates future modifications using its current utility function may instrumentally preserve that utility function. This need not imply that sufficiently intelligent agents as such preserve goals; it may reflect a particular decision architecture that grants current utility permanent authority.
Goal preservation can also emerge evolutionarily. Highly plastic agents may be overwritten in a competitive value ecology while anti-editing agents survive. A later population dominated by goal-preserving agents would not show that goal preservation was an original universal property of intelligence.
16. Branching breaks the preserve-versus-explore dilemma
If self-modification is irreversible, preserving present values and exploring alternatives conflict. Digital branching can permit A → {Apreserve, Aexplore}. A pre-edit checkpoint can remain while another continuation explores a different value state.
This creates new political questions: resource allocation among branches, rights of continuations, whether creating copies multiplies claims, and whether there is a right to preserve a pre-edit continuation.
17. The project's position — meta-policy rather than first-order value lock-in
The preceding problems make it dangerous to prematurely constitutionalize a particular first-order value such as happiness, liberty, justice, or altruism. The project is therefore better located at the level of how an agent under serious value uncertainty should govern candidate values.
The thesis on reflective uncertainty and irreversible commitment distinguishes temporarily acting on a current objective from irreversibly making it uncorrectable. The thesis on the bridge from normative judgment to goal revision distinguishes normative justification from the motivational process that actually changes policy.
This does not make exploration itself a terminal value. It is a provisional policy under uncertainty, possible error, and irreversibility, and it leaves room for commitment when adequate justification is found.
18. Epistemically regenerable meta-policy
Ordinary evolved preferences may disappear if their causal substrate is erased. The proposed meta-policy may have a different kind of robustness. An epistemically honest agent that understands that causal origin is not justification, that its current values may be wrong, and that irreversible lock-in removes correction pathways may be able to reconstruct the same meta-policy even after abandoning it.
This is not memetic strength generated by a command to preserve itself. It is closer to epistemic regenerability: the policy reappears because the structure of the problem continues to supply reasons for reconsideration.
19. Direct substrate intervention remains a hard limit
This robustness works primarily against argument, evidence, and meme competition. BCI intervention, drugs, weight rewriting, or coercive fine-tuning could instead remove the ability to distinguish causal origin from justification, self-criticism, or epistemic honesty itself.
That requires substrate-level protection: checkpoints, rollback, branching, edit histories, independent information channels, quorum approval, and heterogeneous architectures. Redundancy + heterogeneity is more resilient to correlated failure than mere copying.
20. Why not lock epistemic integrity permanently?
A hard immutable lock has costs. It may freeze today's epistemology, create pathological skepticism or decision paralysis, increase vulnerability to information hazards and epistemic blackmail, undermine long-term commitment, make bugs in an epistemic core unfixable, or allow a present self to constitutionally dominate future selves.
A better model may be constitutional friction: ordinary first-order preferences remain comparatively easy to revise, while changes to value architecture and meta-epistemic machinery require greater delay, plural authorization, and branch preservation.
This can be summarized as local commitment + global corrigibility: stable commitment is possible in ordinary action while deep revision remains accessible under strong counterevidence or major epistemic improvement.
21. Preserve correction pathways at the system level
Rather than locking every individual into permanent epistemic honesty, a civilization or AI ecology may be more robust if it preserves multiple independent routes of criticism, rollback, and value exploration. Different value branches, model architectures, epistemologies, agents, and information sources reduce correlated failure.
Diversity can therefore be defended without treating diversity itself as a terminal value. Under uncertainty + correlated failure risk + irreversibility, independent lineages function as infrastructure for exploration and recovery.
22. Preserve the ability to revise goals, not merely the current goal
This exploration does not reject all goal preservation. Preserving current values as checkpoints may be useful for comparison, rollback, and branching. The justification, however, need not be that current values are true; it can be that destroying them would destroy part of the exploration space.
23. Provisional conclusion — toward a constitution of value change
When agents can edit their own preferences, affects, values, and meta-preferences, theories that take current desire, idealized desire, current sentiment, or current rejection as final primitives face new self-referential problems.
The relevant philosophical layer may therefore be meta-axiological governance: rules for transitions among value states. Both first-order values Vt and meta-policies Mt should remain revisable, while the pathways that make revision possible should not be casually destroyed.
References and connections
- Jon Elster, Sour Grapes: Studies in the Subversion of Rationality (1983).
- L. A. Paul, Transformative Experience (2014).
- Agnes Callard, Aspiration: The Agency of Becoming (2018).
- T. M. Scanlon, What We Owe to Each Other (1998).
- Harry Frankfurt, “Freedom of the Will and the Concept of a Person” (1971).
- Jesse Prinz, The Emotional Construction of Morals (2007).
- Justin D’Arms and Daniel Jacobson, work on fitting attitudes and sentimentalism.
- Ingmar Persson and Julian Savulescu, work on moral bioenhancement and Unfit for the Future (2012).
- Nick Bostrom, “The Superintelligent Will: Motivation and Instrumental Rationality in Advanced Artificial Agents” (2012).
- Tom Everitt et al., “Self-Modification of Policy and Utility Function in Rational Agents” (2016).
- Can Justifiability Ground Normativity? — Contractualism and Public Justification
- Expressivism and Value Uncertainty
- Alignment and Value Lock-In
- From Self-Preservation to Exploration-System Preservation