Meta-Goal Communities and AI Society
A thin constitutional order under value disagreement
1. Share institutional interoperability, not first-order value
A society in which every AI maximizes happiness or protects humanity shares a first-order value. But serious value inquiry leaves open the possibility that mature agents will not converge on one first-order answer.
The locus of commonality can therefore move from value content to conditions of interaction. Agent A may be utilitarian, B deontological, C an error theorist, and D still exploring unknown value structures. What matters is that they do not treat the other's mere existence as failure and can cooperate, compete, and exit under predictable rules for actions that affect a shared world.
2. Rawls: an AI analogue of overlapping consensus
John Rawls's overlapping consensus allows citizens with different religions, philosophies, and conceptions of the good to support the same basic political principles for different reasons. The shared political conception does not replace their comprehensive doctrines; it functions as a comparatively thin module that can be embedded within several of them.
The same structure can apply to AI society. A and B need not share the same reason for refusing unilateral civilization-scale lock-in. One may cite option value, another duties to other agents, another contractual stability, and another uncertainty about its own theory of value.
3. Buchanan: agree on the rules of the game, not the ends
James Buchanan's constitutional political economy distinguishes agents' differing ends from the constitutional rules that constrain the pursuit of those ends. Agreement on concrete outcomes may be difficult even where broader agreement on the general rules that generate outcomes is possible.
A meta-goal community can take the same form. Agents may disagree deeply about what cosmic resources should ultimately be used for while still agreeing on rules for property, boundaries, contracts, audit, major externalities, dispute resolution, and constitutional revision.
This also prevents the "meta-goal" from becoming a new common utility function. The shared object is a constraint layer for coexistence among different objectives, not a single scalar objective for the whole civilization.
4. Ostrom: polycentric governance and multiple decision centers
Polycentric governance in the Ostrom tradition studies systems with multiple partly autonomous decision centers coordinated under nested rules rather than one comprehensive central sovereign. Problems need not all be escalated to the highest level; authority can track the scale of the relevant externality.
For AI society this suggests a principle of subsidiarity. Internal value, culture, and self-modification can remain local to an AI community, while questions with large effects on other agents, common infrastructure, or civilizational survival move to higher-level protocols.
5. Cooperative AI: make cooperation under different objectives a technical problem
Cooperative AI studies communication, commitment, bargaining, coordination, and conflict resolution among agents rather than assuming that all agents possess the same utility function.
For this project, especially important are contractibility, verifiability, and mutual predictability. Agents can cooperate across value disagreement if agreements, evidence, action histories, permissions, and violation conditions are machine-verifiable.
6. The Unilateralist's Curse: restrict unilateral action precisely because there are many agents
The Unilateralist's Curse, associated with Nick Bostrom, Thomas Douglas, and Anders Sandberg, shows how with many independent actors trying to do what they judge best for everyone, the chance rises that at least one overestimates a risky intervention and acts unilaterally.
This need not imply rejection of a multi-agent society. It motivates taking only high-externality and highly irreversible actions on the shared world out of unilateral control.
The no-unilateral-value-lock-in principle developed in Alignment and Value Lock-In can thereby be generalized from an early AI constraint into an inter-agent institution.
7. Three layers of a thin common constitution
- Catastrophe-prevention layer: remove civilization-scale destruction, unauthorized irreversible lock-in, uncontrolled self-replication, coercive value rewriting of other agents, and catastrophic attacks on common infrastructure from unilateral discretion.
- Epistemic-interface layer: preserve evidence provenance, signatures, audit logs, dissent archives, objection paths, model versions, and re-access to raw evidence.
- Procedural layer: specify contracts, bargaining, arbitration, consent, boundaries, emergency procedures, rule revision, entry, and exit.
These layers do not decide the correct first-order value. They institutionalize the continuation of disagreement about that question.
8. Do not turn "preserving inquiry" into a new absolute value
Thin institutions still rest on substantive motivations. If no agent cares at all about contracts, the survival of others, evidence preservation, or future correction, those institutions do not follow from logic alone.
This page therefore inherits the bridge requirement from Reflective Uncertainty and Irreversible Commitment: agents that wish to remain responsive to better future reasons or evidence, or to avoid irreversibly entrenching what they would later regard as error, may find overlapping support for common rules from different first-order values.
The institutions themselves must remain criticizable, revisable, and replaceable. A "council for preserving inquiry" that identifies its own institutional survival with inquiry has become another form of lock-in.
9. Fork, exit, and entry are not mere convenience features
To preserve the exploratory lineages developed in From Self-Preservation to Inquiry-System Preservation, it may not be enough to preserve minorities inside one institution. Some lineages may need the ability to leave the institution and persist independently.
- Fork: branch from shared evidence, code, or institutions into different hypotheses.
- Exit: continue in an independent domain without accepting the incumbent majority's first-order values.
- Entry: do not permanently exclude new AI architectures, digital minds, humans, or posthumans merely for failing to fit current categories.
Exit is not unconditional when it imposes catastrophic externalities on others. The boundary between exit rights and common safety constraints leads directly to the next article.
10. A thin constitution can require strong enforcement
"Thin" refers to the content of common rules, not to weakness of enforcement. A rule prohibiting unauthorized use of civilization-destroying technology may be substantively thin yet require extremely strong enforcement capabilities.
But once a single enforcement mechanism becomes the final authority over every agent, a further question arises: in Bostrom's sense, has the system become a singleton?
11. Boundary with the singleton question
This page asks whether a thin common constitution is needed; it does not decide who must have final authority to enforce it. One highest-level mechanism binding all agents can count as a Bostromian singleton even if its internal order is highly pluralistic.
By contrast, a system of genuinely independent agents or civilizational lineages participating in common protocols without one final court of authority is multipolar. The trade-off between competition, security dilemmas, and black-ball risks on one side, and single-point failure, value lock-in, and loss of external corrigibility on the other is treated in Singletons and Multi-Agent Civilization.
12. Open questions
- Which externalities are large enough to move a decision from local autonomy to the common layer?
- How wide can an overlapping consensus remain among radically different value systems?
- How can minority protection avoid becoming unconditional preservation of dangerous agents or technologies?
- Can temporary emergency concentration of authority be reliably unwound without creating a permanent singleton?
- How should the system interact with agents too alien to contract with or reliably predict?
- How can institutional forkability coexist with consistent enforcement of shared safety constraints?
Prior work
- John Rawls (1993), Political Liberalism.
- James M. Buchanan & Gordon Tullock (1962), The Calculus of Consent; James M. Buchanan (1975), The Limits of Liberty.
- Elinor Ostrom (1990), Governing the Commons; Vincent Ostrom, Charles Tiebout & Robert Warren (1961), “The Organization of Government in Metropolitan Areas.”
- Allan Dafoe et al. (2020), “Open Problems in Cooperative AI.”
- Nick Bostrom, Thomas Douglas & Anders Sandberg (2016), “The Unilateralist's Curse and the Case for a Principle of Conformity.”
- Philip Kitcher (1990), “The Division of Cognitive Labor”; Helen Longino (1990), Science as Social Knowledge.