AI Value Exploration Notes
Exploration

Meta-Goal Communities and AI Society

A thin constitutional order under value disagreement

Exploration v0.2 · English translation · 2026-08-29 · working hypothesis

Working hypothesis: Cooperation among advanced AIs need not require every agent to share one terminal value, one metaethics, or one vision of the future. Agents may instead support—often for different reasons—a thin common rule set governing catastrophic externalities, verifiable evidence and reasoning, contracts and bargaining, arbitration, forking, exit, and rule revision. A meta-goal community is therefore not a super-agent with one terminal objective, but a constitutional interface that allows value disagreement and inquiry to persist.

1. Share institutional interoperability, not first-order value

A society in which every AI maximizes happiness or protects humanity shares a first-order value. But serious value inquiry leaves open the possibility that mature agents will not converge on one first-order answer.

The locus of commonality can therefore move from value content to conditions of interaction. Agent A may be utilitarian, B deontological, C an error theorist, and D still exploring unknown value structures. What matters is that they do not treat the other's mere existence as failure and can cooperate, compete, and exit under predictable rules for actions that affect a shared world.

2. Rawls: an AI analogue of overlapping consensus

John Rawls's overlapping consensus allows citizens with different religions, philosophies, and conceptions of the good to support the same basic political principles for different reasons. The shared political conception does not replace their comprehensive doctrines; it functions as a comparatively thin module that can be embedded within several of them.

The same structure can apply to AI society. A and B need not share the same reason for refusing unilateral civilization-scale lock-in. One may cite option value, another duties to other agents, another contractual stability, and another uncertainty about its own theory of value.

Overlapping meta-consensus: The strength of a common institution need not be measured by whether everyone accepts one theory of justice, but by how many distinct value systems can independently support the same thin procedures.

3. Buchanan: agree on the rules of the game, not the ends

James Buchanan's constitutional political economy distinguishes agents' differing ends from the constitutional rules that constrain the pursuit of those ends. Agreement on concrete outcomes may be difficult even where broader agreement on the general rules that generate outcomes is possible.

A meta-goal community can take the same form. Agents may disagree deeply about what cosmic resources should ultimately be used for while still agreeing on rules for property, boundaries, contracts, audit, major externalities, dispute resolution, and constitutional revision.

This also prevents the "meta-goal" from becoming a new common utility function. The shared object is a constraint layer for coexistence among different objectives, not a single scalar objective for the whole civilization.

4. Ostrom: polycentric governance and multiple decision centers

Polycentric governance in the Ostrom tradition studies systems with multiple partly autonomous decision centers coordinated under nested rules rather than one comprehensive central sovereign. Problems need not all be escalated to the highest level; authority can track the scale of the relevant externality.

For AI society this suggests a principle of subsidiarity. Internal value, culture, and self-modification can remain local to an AI community, while questions with large effects on other agents, common infrastructure, or civilizational survival move to higher-level protocols.

Polycentric principle: The scope of common institutions should track the scope of externalities, not the perceived moral importance of a local value. Do not unnecessarily convert local value disagreement into a civilization-wide value settlement.

5. Cooperative AI: make cooperation under different objectives a technical problem

Cooperative AI studies communication, commitment, bargaining, coordination, and conflict resolution among agents rather than assuming that all agents possess the same utility function.

For this project, especially important are contractibility, verifiability, and mutual predictability. Agents can cooperate across value disagreement if agreements, evidence, action histories, permissions, and violation conditions are machine-verifiable.

6. The Unilateralist's Curse: restrict unilateral action precisely because there are many agents

The Unilateralist's Curse, associated with Nick Bostrom, Thomas Douglas, and Anders Sandberg, shows how with many independent actors trying to do what they judge best for everyone, the chance rises that at least one overestimates a risky intervention and acts unilaterally.

This need not imply rejection of a multi-agent society. It motivates taking only high-externality and highly irreversible actions on the shared world out of unilateral control.

The no-unilateral-value-lock-in principle developed in Alignment and Value Lock-In can thereby be generalized from an early AI constraint into an inter-agent institution.

7. Three layers of a thin common constitution

  1. Catastrophe-prevention layer: remove civilization-scale destruction, unauthorized irreversible lock-in, uncontrolled self-replication, coercive value rewriting of other agents, and catastrophic attacks on common infrastructure from unilateral discretion.
  2. Epistemic-interface layer: preserve evidence provenance, signatures, audit logs, dissent archives, objection paths, model versions, and re-access to raw evidence.
  3. Procedural layer: specify contracts, bargaining, arbitration, consent, boundaries, emergency procedures, rule revision, entry, and exit.

These layers do not decide the correct first-order value. They institutionalize the continuation of disagreement about that question.

8. Do not turn "preserving inquiry" into a new absolute value

Thin institutions still rest on substantive motivations. If no agent cares at all about contracts, the survival of others, evidence preservation, or future correction, those institutions do not follow from logic alone.

This page therefore inherits the bridge requirement from Reflective Uncertainty and Irreversible Commitment: agents that wish to remain responsive to better future reasons or evidence, or to avoid irreversibly entrenching what they would later regard as error, may find overlapping support for common rules from different first-order values.

The institutions themselves must remain criticizable, revisable, and replaceable. A "council for preserving inquiry" that identifies its own institutional survival with inquiry has become another form of lock-in.

9. Fork, exit, and entry are not mere convenience features

To preserve the exploratory lineages developed in From Self-Preservation to Inquiry-System Preservation, it may not be enough to preserve minorities inside one institution. Some lineages may need the ability to leave the institution and persist independently.

Exit is not unconditional when it imposes catastrophic externalities on others. The boundary between exit rights and common safety constraints leads directly to the next article.

10. A thin constitution can require strong enforcement

"Thin" refers to the content of common rules, not to weakness of enforcement. A rule prohibiting unauthorized use of civilization-destroying technology may be substantively thin yet require extremely strong enforcement capabilities.

Key separation: strong coordination does not entail thick value centralization. A society could strongly enforce constraints on catastrophic externalities while leaving first-order values, cultures, and inquiry directions plural.

But once a single enforcement mechanism becomes the final authority over every agent, a further question arises: in Bostrom's sense, has the system become a singleton?

11. Boundary with the singleton question

This page asks whether a thin common constitution is needed; it does not decide who must have final authority to enforce it. One highest-level mechanism binding all agents can count as a Bostromian singleton even if its internal order is highly pluralistic.

By contrast, a system of genuinely independent agents or civilizational lineages participating in common protocols without one final court of authority is multipolar. The trade-off between competition, security dilemmas, and black-ball risks on one side, and single-point failure, value lock-in, and loss of external corrigibility on the other is treated in Singletons and Multi-Agent Civilization.

12. Open questions

Prior work