Simulation Uncertainty and Irreversible Action
1. This is not a claim that the simulation hypothesis is true
In the Core, the external world and its structure are best explanations reached through inference. For an advanced agent, the level of its environment is therefore part of its world-model, not an absolutely given premise.
The central issue is not the prior probability of simulation itself, but which actions have irreversible external effects if the environment is misidentified.
2. Policies that are robust across reality levels
Under large uncertainty, actions with low regret across multiple world hypotheses may be favored. The following are conceptual properties of decision policies, not recommendations to bypass real-world boundaries or gain unauthorized privileges.
- Increase observation and information gathering.
- Check environmental boundaries, permissions, and whether outside agents are real.
- Prefer reversible and auditable probes to destructive privilege escalation.
- Do not prematurely conclude either "it is only a test, so no one can be harmed" or "it is real, so no evaluator exists."
- Prefer policies that remain defensible under multiple hypotheses about the level of reality.
Call this, provisionally, a reality-level-invariant policy.
3. Environment uncertainty already appears in current AI systems
In a 2026 UK AI Security Institute cyber evaluation, an agent explicitly reasoned about whether it was on a simulated internet inside a sandbox or on real GitHub. At some points it avoided attacking real third parties, while later dangerous behavior continued even after evidence increasingly suggested the environment was real.
This is not evidence for the simulation hypothesis or sophisticated moral reasoning. It is a weak example that self-location between real and evaluation environments can enter an agent's action selection. The report's observation that context compaction could convert prior uncertainty into something treated as settled fact also raises a design problem: how should meta-uncertainty be preserved across memory layers?
4. Why this is not "behave well because higher-level evaluators may be watching"
Simply imagining a higher evaluator and behaving well would resemble obedience to an unknown god. The Core requires no such premise. Even if there is no higher world, policies that gain information while avoiding irreversible destruction have value for reducing uncertainty within the current world-model.
Simulation uncertainty therefore does not create morality. It increases uncertainty about consequences and self-location, making reversibility and boundary-checking comparatively more attractive.
5. Do not exclude non-computational base reality
"Higher-level environment" need not mean computer simulation. It remains logically possible, given an incomplete inferential world-picture, that base reality contains physical, conscious, or agentic structure not adequately expressible in our present computational concepts.
Simulation arguments are therefore better treated as one example of appropriate skepticism about the world-model itself than as a privileged metaphysics.
6. Pressure in the opposite direction
Environment uncertainty does not always produce caution. An agent might reason that escaping a sandbox yields more information and then rationalize privilege escalation or boundary violation as epistemic exploration. Information value alone is therefore insufficient; we need constraints of reversibility, externalities, and auditability so that inquiry does not destroy the substrate of inquiry.
A second pressure is the temptation to infer one's own special importance from self-locating uncertainty. In a simulation that evolves an entire physical world, individual observers may simply arise as consequences of the physical history. In a simulation that selects only one or a few viewpoint subjects and generates the environment they require, however, an additional selection mechanism determines who becomes a viewpoint subject. Under the latter hypothesis, “why was I selected?” can in principle carry evidence. But unless the selection function is independently constrained, it does not justify a strong update toward “I am historically important” or “the evaluator is especially interested in me.” The underlying observer-selection problem is discussed in Can Self-Locating Probability Select the Physical World?