AI Value Exploration Notes
Exploration

Can ASI Capability Diffusion Be Made Safe?

Civilizational resilience, compartmentalization, and least-sovereign safety

Exploration v0.1 · 2026-08-16 · working hypothesis · English working translation 2026-08-17

Question: If a single actor with ASI could cause civilization-scale catastrophe, must advanced capabilities be concentrated in a small number of hands for safety? Or could physical AI, defensive capacity, distributed production, physical compartmentalization, and branching into space make civilization sufficiently fault-tolerant that one actor’s failure remains local—allowing capabilities to diffuse while plural political agency survives?

1. The conflict between ASI safety and plural agency

AI-safety arguments often emphasize the possibility that failure, misuse, or loss of control by one actor could create enormous externalities. If one actor’s use of ASI could irreversibly damage civilization as a whole, unrestricted capability diffusion is plainly dangerous.

Yet effective capability restriction requires someone to monitor and, when necessary, stop access to compute, models, communications, robotics, finance, or research facilities. The same safety authority can become political power capable of eliminating other actors’ effective agency.

Conditional trilemma: If even one ASI-enabled actor can cause irreversible civilization-scale harm and defenders cannot reliably stop that actor, it becomes difficult to jointly satisfy catastrophic safety, plural agency, and no superior sovereign.

This is not claimed as a proven impossibility. The question is which technological conditions make the trilemma severe, and which conditions might relax it.

2. Four meanings of “ASI diffusion”

“Everyone has ASI” is too coarse. At least four dimensions should be separated.

Intelligence diffusion
Broad access to high-level reasoning, research, and design capability.
Agency diffusion
Independent ability to use that intelligence for production, research, organization-building, and institutional experimentation.
Catastrophic capability diffusion
Ability of a single actor to impose irreversible civilization-scale harm on many non-consenting others.
Control-power diffusion
How authority to stop or constrain other actors’ AI, compute, and actuation is distributed.
Distinction: democratized intelligence ≠ democratized agency, and democratized agency ≠ democratized catastrophic capability.

The desirable target need not be equal distribution of every capability. A more attractive equilibrium may widely distribute the capacities needed for self-maintenance, inquiry, exit, and defense while constraining capacities that let one actor erase the future possibility of others.

3. Look at failure domains, not capability alone

ASI risk depends on more than how intelligent the system is. The same capability can have very different consequences depending on whether society is one tightly coupled global failure domain or can localize failures.

The more civilization depends on single communications, financial, cloud, logistics, food, energy, or AI-update systems, the easier it is for one compromise to become a global state transition. If regions, networks, and supply systems can be isolated and restored from independent backups, the same attack or accident may remain a local disaster.

Comparison axis: Conceptually, systemic risk rises with capability × offense advantage × coupling × irreversibility and falls with defense capacity × compartmentalization × recoverability.

ASI safety therefore has a second strategy in addition to reducing dangerous capability: increase the denominator on the civilization side.

4. Safety by restriction and safety by resilience

Restriction strategies constrain dangerous capability in advance through model access rules, compute controls, user authentication, monitoring, or limits on dangerous research.

Resilience strategies assume that some dangerous acts will occur and improve the capacity to resist, absorb, recover, and adapt before they propagate into global catastrophe.

The 2026 International AI Safety Report treats societal resilience as one layer of defense in depth because technical safeguards have limits. Examples include early detection, isolation, and medical stockpiles for biological risk, and network segmentation, automated isolation, and offline backups for cyber risk; it also notes major empirical gaps in how effective such measures would be.

The political side effects differ. Restriction tends to concentrate monitoring and shutdown authority. Resilience, if well designed, may strengthen each actor’s capacity for self-maintenance and reduce dependence on a central authority.

Hypothesis: resilience investment may be both safety technology and a way to reduce safety-induced political centralization.

5. Biosecurity: can physical AI change the cost of isolation?

Strong epidemic isolation is socially expensive because stopping human contact also disrupts logistics, care work, manufacturing, food supply, infrastructure maintenance, and parts of medicine.

If sufficiently capable physical AI could maintain warehouses, delivery, factories, waste handling, construction and repair, testing, medical support, disinfection, and infrastructure without human-to-human contact, it could sharply reduce isolation cost. That might enlarge the space of equilibria in which dangerous biological knowledge need not be perfectly contained because outbreaks can still be localized.

The WHO’s 2024 Laboratory biosecurity guidance already includes information, cyber, and AI within biosecurity risk management alongside high-impact biological materials. Under ASI, the relevant frame may need to extend from laboratory management to resilience of social infrastructure.

The boundary condition is demanding: conceptually, defense needs Tdetect + Tresponse < Tcascade. If AI accelerates offensive design faster than defensive detection and response, physical-AI recovery alone will not suffice.

6. Physical compartmentalization: from terrestrial distribution to space

Pushed far enough, civilizational resilience leads to geographic and physical causal compartmentalization: cities, regions, continents, orbital habitats, the Moon, or Mars designed as failure domains whose accidents do not automatically propagate to one another.

Distance alone is not independence. A genuinely independent compartment needs reproductive autonomy: the ability to reproduce food, power, compute, manufacturing, repair, medicine, and AI locally.

A 2026 NASA study of long-duration lunar habitation likewise treats high availability, graceful degradation, strong fault tolerance, autonomous computing, monitoring, and fault response as requirements for long-term sustainability: Enabling Reliable, Fault-Tolerant Autonomous Lunar Habitats.

Distinction: geographical separation ≠ causal independence. Independent civilizational units require self-reproduction as well as distance.

7. Multiple planets with the same ASI can still be one failure domain

Several settlements on the Moon, Mars, and in orbit can remain one failure domain if all depend on the same ASI, root key, model-update channel, semiconductor supply, authentication system, or financial infrastructure. One software failure, compromise, or political decision could then propagate everywhere.

Real compartmentalization therefore requires not only geography but hardware diversity, software diversity, AI diversity, institutional diversity, supply-chain diversity, and forkability.

This is relevant not only to extinction risk. Independent causal paths capable of trying different institutions, values, and AI designs are part of the infrastructure of Value Exploration.

8. LAWS, private robots, and coercive outside options

One hypothesis is that privately or communally owned robots could play a role analogous to armed citizens in some early-modern political orders. But firearms and autonomous robotics differ: if one actor can operate large fleets, nominally broad ownership can still reconcentrate coercive capability through capital inequality.

Likewise, privately owned robots that depend on cloud authorization and can be remotely disabled by firms or governments do not provide strong political control sovereignty.

Full diffusion of autonomous lethal capability to individuals may not be necessary for political agency. In 2025 the UN General Assembly adopted A/RES/80/57 on lethal autonomous weapons systems by 164 votes to 6, while international norm formation continues; the UN Secretary-General has also called for legal prohibitions on systems that take life without human control.

For political outside options, independent access to production, power, communication, compute, mobility, defense, and repair may matter more. There may therefore be a strong case for prioritizing material autonomy diffusion over lethal capability diffusion.

9. Three capability classes

Class I — Resilience-enhancing agency
Capabilities whose wider distribution tends to improve both safety and political agency: distributed power, production, repair, medicine, communication, defense, and independent compute.
Class II — Reciprocal coercive capability
Capabilities that may check central power when distributed but can also create arms races, accidents, or private domination: robotic defense and some cyber capabilities.
Class III — Non-local irreversible capability
Capabilities allowing one actor to impose enormous irreversible harm on many non-consenting others. Simple diffusion “for balance” may destroy the possibility of balance itself.
Provisional principle: Widely distribute capacities that let actors maintain, defend, and exit for themselves; strongly constrain capacities that let one actor unilaterally and irreversibly eliminate others’ possibility of existence.

10. Least-Sovereign Safety

Compressed into an institutional principle, the aim is not to build the central actor best able to manage safety, but to minimize discretionary sovereignty over others subject to meeting the necessary safety threshold. This page calls that working idea Least-Sovereign Safety.

A recursion problem remains: who updates the safety rules, grants exceptions, invokes emergency powers, or determines that a protocol has been violated? Even with distributed operations, sovereignty can reappear wherever amendment or emergency authority is concentrated.

11. Conditions under which ASI becomes a real trilemma

At minimum, forecasts should track the following variables.

If φ is low, D and R are high, and dangerous actions retain strong physical bottlenecks, broad ASI diffusion may be compatible with safety. If φ is high, offense dominates, irreversibility is high, and software intelligence converts easily into physical catastrophic capability, pressure toward centralized control will be much stronger.

12. The resilience race

If ASI competition is viewed only as a capability race, attention goes only to who acquires capability first. Yet the compatibility of political freedom with safety also depends on how quickly civilization itself becomes fault-tolerant.

Resilience race: the race between the speed at which ASI expands the action space of one actor and the speed at which physical AI, automated medicine, distributed production, independent power, network segmentation, multiple AI lineages, and space settlement shrink failure domains.

If the first side runs far ahead, “temporary concentration for safety” may appear rational and then become safety-induced political lock-in as the managing actor uses ASI to reduce its material dependence on humans.

If the second side can intentionally lead, the domain in which civilization can tolerate imperfect containment expands, potentially reducing the need for permanent centralized monitoring and shutdown authority.

13. What does not follow

14. Provisional conclusion

Whether ASI diffusion leads to human extinction is not determined by ASI capability alone. The coupling structure that propagates one actor’s failure, the relative speed of offense and defense, physical bottlenecks, and recoverability matter just as much.

If civilization remains a global single failure domain while ASI diffuses, strong arguments for centralized control gain force. But centralized control is itself civilization-scale political power and, combined with AI-enabled reductions in dependence on humans, may create durable power lock-in.

An alternative strategy is therefore to build a civilization in which capability can diffuse safely before diffusion becomes unavoidable: use physical AI, distributed production and power, rapid medical response, network compartmentalization, multiple independent AI systems, and self-reproducing space settlements to localize failures and make society hard to kill.

The final question is not simply whether to centralize or diffuse ASI. It is whether we can construct a material and technological equilibrium in which many actors have enough agency to build their own worlds, while no actor can unilaterally erase everyone else’s worlds.

Working hypothesis: This is not a settled policy recommendation based on known ASI capability characteristics. Capability fungibility, offense-defense balance, physical bottlenecks, and the effectiveness of resilience measures remain uncertain and could substantially revise the conclusion.