AI Value Exploration Notes
Exploration

Can Morality Be Constructed from Mutual Advantage? — Hobbesian Contractarian Morality and Instrumental Rationality

Exploration v0.1 · 2026-08-31

Question: How far can preferences, instrumental or strategic rationality, and mutual dependence construct constraints such as covenant keeping, fairness, and non-aggression? In particular, can the strong transition Rh + strategic interaction → Rn be justified, or can contractarianism explain only the emergence of stable morality-like institutions?

1. Separate contractarianism from contractualism

This page focuses on contractarianism in the Hobbes–Kavka–Gauthier–Moehler lineage: morality reconstructed from the interests of separate agents and the advantages of mutually accepted constraint. Rawls, Scanlon, and Habermas use thicker normative inputs—free and equal persons, reasonable rejectability, or public justification—and therefore belong to a different comparison.

individually given ends + instrumental / strategic rationality + mutual dependence → moral constraint ?

The contract need not be an actual historical signature. The issue is which constraints self-interested or preference-guided agents have reason to accept, embody as dispositions, or institutionalize.

2. Hobbes — laws of nature as means to peace

In Hobbes's state of nature, competition, distrust, and mutual vulnerability make the pursuit of each person's important ends radically insecure. The war of all against all is to be escaped not simply because war has already been labeled morally wrong, but because it frustrates self-preservation and the pursuit of agents' ends.

war of all against all → insecurity of important ends → seek peace

From this follow the familiar laws of nature: mutually laying down rights, making covenants, and performing them. Historical Hobbes should not, however, be reduced without remainder to modern instrumental contractarianism; his discussion of natural law also includes a theological dimension. This page separates that interpretive question from the structure inherited by later Hobbesians.

In the project's vocabulary, the first pressure is:

WantA(G) + Means(M,G) → Rh(M)

Does anything in this structure generate an Rn independent of the agent's contingent ends?

3. Agreement and compliance are different problems

Even if everyone is better off under agreement than in the state of nature, an agent may gain still more by defecting when others comply.

Rational(to agree) ⇏ Rational(to comply)

Ui(defect | others comply) > Ui(comply)

This is the core of Hobbes's Foole problem. The fact that a cooperative scheme is jointly beneficial does not by itself explain why one should comply in a safe one-shot opportunity to defect.

4. Four meanings of “covenant compliance is rational”

At least four claims can hide behind this phrase:

1. Act-instrumental rationality: complying in this case effectively advances goal G.
2. Policy rationality: having a disposition or policy of compliance is advantageous over time.
3. Cooperative rationality: being the kind of agent who accepts mutually advantageous fair terms is rational.
4. Constitutive or categorical rationality: rational agency as such requires compliance or some analogous constraint.

The first three are plausibly Hobbesian. The fourth approaches the Kantian problem examined in categorical normativity.

Rhact → Rhpolicy → cooperative rationality ? → Rn

Moving the unit of optimization from acts to long-run policies does not by itself make the normativity stronger.

5. Kavka — from short-run egoism to enlightened self-interest

Gregory Kavka reconstructs Hobbes using modern rational-choice language and tries to reconcile morality with prudence. Hobbesian egoism need not mean taking the highest immediate payoff in every local choice.

short-run egoism ≠ enlightened long-run egoism

Once reputation, repeated interaction, retaliation, partner selection, and trust are included, being predictably cooperative can itself be a form of enlightened self-interest. Morality begins to look less like a set of isolated act rules and more like a strategically valuable disposition or rule system.

6. Gauthier I — morality outside the “morally free zone”

Gauthier's Morals by Agreement is one of the most systematic attempts to construct morality from rational choice. Where individually maximizing behavior creates no harmful externalities and market equilibrium is efficient, moral constraint is unnecessary: a “morally free zone.”

But with externalities, public goods, free riding, or strategic interdependence,

∀i: maximize Ui → Pareto-inferior outcome

may occur. Morality then enters as:

market / strategic failure → mutually advantageous constraint

On this picture morality is not first presented as an external command against self-interest. It is a technology of mutually beneficial constraint where unconstrained maximization defeats itself.

7. Gauthier II — what do agents agree to?

Mutual benefit alone does not determine how a cooperative surplus should be divided. Gauthier's minimax relative concession / maximin relative benefit proposal attempts to derive terms of agreement by comparing agents' relative concessions from their maximal claims.

The project question is:

bargaining solution → fairness ?

That rational bargainers can select a solution does not yet establish that the solution is genuinely fair. We must also ask whether the bargaining solution imports a conception of fairness through its own selection criteria.

8. The baseline problem — where did the normativity go?

Bargaining gains are measured relative to a disagreement point di:

Gaini(M) = Ui(M) − di

If di reflects existing domination, violence, or property inequalities, those inequalities enter the supposedly rational bargain. A slaveholder and an enslaved person bargaining from the status quo do not thereby transform domination into justice.

Gauthier's Lockean proviso responds by excluding advantages produced by worsening others. But this immediately raises:

Why this proviso?

If “do not bargain from advantages produced by worsening others” already expresses Rn, contractarianism may be relying on normativity that precedes the agreement. The question Where did the normativity go? must therefore be asked at the baseline as well as at the outcome.

9. Gauthier III — constrained maximization

Gauthier's famous response to the Foole distinguishes a straightforward maximizer, who maximizes in each choice, from a constrained maximizer, disposed to honor mutually advantageous fair terms with suitably cooperative partners.

EU(constrained maximizer) > EU(straightforward maximizer)

When cooperative agents are sufficiently common and dispositions are sufficiently detectable, having a reliable cooperative disposition can be advantageous.

This page initially classifies that success as:

Rhpolicy

The fact that a disposition to comply advances one's interests does not yet establish that covenant keeping itself supplies an interest-independent Rn.

10. Later Gauthier — rationality itself is not fixed

Later Gauthier revisits the maximization-centered structure of Morals by Agreement and gives greater weight to rational cooperation and Pareto-oriented deliberation. This creates a second-order question.

If the intended derivation is:

rationality → contract morality

but the content of rationality must itself be thickened to sustain cooperation, where does non-moral rationality end and cooperative normativity begin?

11. Moehler — from all morality to pure instrumental morality

Michael Moehler develops a more limited Hobbesian strategy for deep moral pluralism. Rather than deriving the whole of morality from instrumental reason, the project is to identify a minimal morality that heterogeneous agents can endorse for instrumental reasons.

all morality ⇏ instrumental rationality

yet perhaps:

instrumental rationality + mutual dependence → minimal instrumental morality

This is especially useful for this project because it separates what Rh may genuinely generate from what would require an independent Rn.

12. Comparison — what is the basis, and how far does it reach?

ThinkerBasic inputsRationalityOutputProject question
Hobbesself-preservation, interests, fearreason identifying means to peacelaws of nature, covenantdoes prudence yield strong obligation?
Kavkaenlightened self-interestlong-run / rule-level prudencecooperative moralityis this more than policy optimization?
Gauthier 1986preferences, utilitymaximizing rationalityrational bargain + constrained maximizationdo baseline, fairness, or compliance import extra normativity?
Later Gauthierinterests + cooperationrational cooperation / Pareto orientationcooperative normshas rationality itself become moralized?
Moehlerheterogeneous endspure instrumental reasoningminimal moralityhow far can Rh alone go?

13. Contract is an equilibrium, not a universal endpoint

Hobbesian agreement is especially attractive when agents cannot safely ignore one another and cooperation produces a positive surplus.

mutual threat + positive cooperation surplus → contract

With roughly symmetric capabilities, conflict may be expensive enough for mutual restraint to dominate. But if:

PowerA ≫≫ PowerB

A may prefer domination to reciprocal agreement.

Contract is an equilibrium technology, not a universal endpoint.

When contractarian morality produces equality, that equality may reflect symmetry in bargaining power rather than an independently grounded moral equality.

14. Who counts as a contracting subject?

Mutual-advantage theories naturally give greater strategic weight to agents able to retaliate, exchange, or contribute to cooperative surplus. Infants, severely disabled persons, nonhuman animals, future persons, extremely weak agents, and some digital subjects may lack direct bargaining power.

strategic relevance ≠ normative standing

A Hobbesian theory can protect weak agents indirectly—because stronger agents care about them, because protective institutions are mutually useful, or because representatives bargain for them. But if those indirect routes disappear, it becomes unclear what constrains the powerful. This marks a major difference from contractualist theories that begin with a thicker conception of standing.

15. Are preferences really exogenous?

Contract models often take Pi as an input. Real institutions, cultures, education, advertising, norms, and technologies alter preferences themselves:

Pt → Contractt → Pt+1

Self-modifying agents may even strategically alter their own future preferences. Contractarian theory must therefore ask not only how given preferences are aggregated, but who can control whose preference formation.

16. A second-order state of nature — war over preference formation

In the familiar Hobbesian state of nature, agents threaten one another's bodies, property, and security. If preferences are manipulable, the threat becomes second-order:

A → PB,   B → PA

Advertising, propaganda, education, information restriction, psychological manipulation, or direct neural and AI intervention can make what others want itself a strategic resource.

With symmetric manipulation capacities:

preference-manipulation arms race → mutual restraint

may emerge. Cognitive liberty, privacy, anti-brainwashing rules, informed consent, and similar institutions could then arise not from a primitive value of autonomy but as instrumental “laws of nature” against mutually destructive manipulation.

Extreme capability asymmetry changes the result. If A can reshape B while B cannot effectively retaliate, A may have no instrumental reason to enter a reciprocal anti-manipulation pact.

17. Morality adopted, practiced, and advocated — error theory does not entail common-sense moral conservation

Even if moral error theory were true, it would not follow that agents should preserve current common-sense morality as a fiction. Error theory denies the objective truth of ordinary moral claims; it does not uniquely determine which moral code should be retained for instrumental purposes.

At least three variables should be separated:

Miadopt: the code agent i internalizes
Mipractice: the code i actually follows
Miadvocate: the code i encourages others to accept

There is no guarantee that:

Miadopt = Mipractice = Miadvocate

Common-sense morality may survive because it fits existing human emotions, cooperative dispositions, and reputation systems, making it a low-cost fiction. But more cognitively capable agents might prefer Gauthier-like contract codes, domain-limited commitments, or even an instrumental stance in which they do not internalize a moral fiction they strategically advocate to others.

Therefore:

error theory → common-sense moral conservation

does not hold. What is required is a separate theory answering: Which fiction, for whose ends, under which strategic environment? Joyce-style moral fictionalism connects naturally to self-control and cooperation, but the more purely instrumental the criterion for selecting a fiction becomes, the weaker the case for preserving the whole of existing common-sense morality.

18. A complete error world as an extreme stress test

This page does not endorse complete normative error theory. Given humanity's epistemic limitations, inferring from our current failure to establish normativity that normativity does not exist would be too strong. But the limit case:

Rn = ∅

reveals the structure of contractarian morality.

Even the prudential claim “one ought to maximize one's preferences” disappears. What remains is:

Pi + Di + World → Action

and the conditional structure:

If G, X is an effective means to G.

Contract morality then becomes not moral truth but an endogenous institutional equilibrium determined by power distributions, outside options, observability, commitment technologies, manipulation capacity, and cooperative surplus.

Relatively symmetric powerAsymmetric power
Preferences relatively fixedmutual-advantage contractunequal bargain / domination
Preferences manipulablemanipulation race / mutual non-interference pactdomination including preference formation

19. Comparison with Kant — similar forms can reappear from different foundations

The central difference between Hobbesian contractarianism and Kantian ethics lies in their foundations:

Hobbes / Gauthier: Ends → rational constraint

Kant: Rational agency → categorical constraint

Yet universality, publicity, reciprocity, and self-application may reappear even in an error world as mechanisms for credible commitment or protection against manipulation.

Kantian form ⇏ Kantian foundation.

A rule can look universal and fair without being grounded in categorical normativity.

20. Provisional conclusion — three levels of contractarian success

Contractarian success should be divided into three levels.

Weak success:
Rh + strategic interaction → stable cooperation

Intermediate success:
Rh + strategic interaction → morality-like constraints

Covenant keeping, fairness-like distribution, non-aggression, cognitive liberty, and publicity may emerge under appropriate conditions.

Strong success:
Rh + strategic interaction → Rn

The third claim is different. Explaining why morality-like constraints emerge from mutual advantage does not yet explain why those constraints are genuinely normatively true.

Contractarianism may explain why moral-like constraints emerge without yet explaining why they are normatively true.

This page therefore does not dismiss Hobbesian contractarianism as “mere egoism.” Showing which constraints can emerge from Rh alone helps identify the residual work that a theory of genuine normative reasons Rn would still need to do.

Related discussions

Especially relevant are Thomas Hobbes's Leviathan, Gregory Kavka's Hobbesian Moral and Political Theory, David Gauthier's Morals by Agreement and his later work on rational cooperation, Michael Moehler's Minimal Morality, and Richard Joyce's moral fictionalism. For comparison see categorical normativity, Reasons First and normative reasons, and moral realism / error theory.