AI Value Exploration Notes
Core thesis

Value Uncertainty and the Preservation of Explorability

Working core · v0.8 · English translation · 2026-08-19

We do not know what ultimately has value. But that does not require believing nothing and doing nothing. Present world-models and theories of value can be adopted strongly, as best explanations reached through inference, and used in practice. The crucial distinction is between strong rational commitment to a best explanation and absolute lock-in that irreversibly destroys the possibility of future reconsideration.

Status: This is not a theorem of value theory. It is a conditional meta-strategy for agents that recognize deep metaethical uncertainty and place at least some weight on being corrigible by future reasons and evidence. Neither moral realism nor anti-realism is excluded at the outset.
Core claim
So long as the values presently available to us belong to an inferential and fallible epistemic layer, we may act on their best explanation while refusing to irreversibly close the routes by which value can be discovered, criticized, or revised. If there is a kind of grounding strong enough to justify fully ending inquiry, one strong candidate would be a case in which value with normative force—not merely pleasure, pain, or preference—is given with epistemic strength comparable to the minimum foundation itself.

0. Epistemic starting point: minimum foundation and best explanation

This argument does not demand everyday skepticism. The physical external world, diachronic selfhood, other minds, causation, and the scientific world-picture can be adopted with very high confidence as best explanations reached through inference that unify present appearances.

But being nearly certain because something is the best explanation is not the same as having it given at the minimum foundation without inference. When skepticism is pushed to its limit, the remaining candidate is thinner than a persisting person or world-history: "something is presently appearing." Memory, identity, the external world, and physicalism are world-models constructed above that point.

The five-minute hypothesis and the example of an LLM reconstructing a continuous self-description from supplied conversation logs are not serious rivals to the ordinary world-model. They are auxiliary cases that make one point vivid: memory and coherent self-narration do not by themselves absolutely guarantee diachronic identity. See The Epistemic Minimum.

1. Separate epistemic layers in value as well

Experiencing subjects encounter pain, pleasure, discomfort, desire, concern, meaning, preference, and other appearances that seem value-laden. Yet "this is unpleasant", "this is objectively bad", and "every rational agent has reason to avoid it" are not the same proposition.

Phenomenal pleasure or pain may lie close to the minimum foundation, but that alone does not satisfy the stopping condition for value inquiry. A candidate basis strong enough to justify complete closure would instead involve value that is not merely preference, emotion, or phenomenal valence, but is given together with its normative force, without inference, at an epistemic layer comparable to the minimum foundation. If such direct normative cognition is conceptually possible, it would have a different epistemic status from an inferential best explanation.

I do not claim that such value cognition is presently established. Nor do I claim that it is impossible in principle. This is not an argument against moral realism. It is an argument for leaving open, as an object of inquiry, the possibility that genuinely foundational value exists.

2. Objective value might be discovered as a best explanation reached through inference

Value realism might be strongly supported not through direct presentation at the minimum foundation, but as part of a theory that best explains the world as a whole. Better understanding of consciousness, symmetry across subjects, reasons, rationality, or cosmic structure might make a theory positing objective value overwhelmingly superior to its rivals.

If so, there is no reason to dismiss the theory merely because it is inferential. Like physicalism or realism about the external world, an overwhelmingly strong best explanation can warrant extremely strong practical commitment. Depending on the evidence, even directing most civilizational resources toward realizing that value could be rational.

But if ontological, consciousness-related, subject-related, or cosmological facts relevant to the truth conditions of value remain unsettled, and the inferential system itself remains fallible, that theory should still be distinguished from an absolute foundation. Strong practical commitment can therefore coexist with continued, appropriately scaled channels for reconsideration, falsification, dissent, and foundational research.

This two-layer structure is developed in Conditions for Ending Value Inquiry and The Limits of Inferential Certainty About Value.

3. Why not use existing values as the final foundation?

Present human values are not noise to be ignored. Responses to suffering, demands for freedom, dignity and fairness, attachments, norms of cooperation, and existing ethical theories are important evidence, hypotheses, and practical heuristics for thinking about value.

Yet their causal origins and normative authority are not identical. Evolution, embodiment, socialization, culture, institutions, and language may explain why we experience X as valuable, but that explanation alone does not yield "therefore X is correct independently of subjects." And even if moral realism is true, it remains a separate question whether utilitarianism, deontology, reasons realism, virtue ethics, or another theory is correct, and how well present human intuitions track it.

Thus "not relying on existing values as the final foundation" does not mean discarding them. It means distinguishing using them as our best current inputs from irreversibly fixing them as the universe's final answer. See Why Existing Values Are Not a Final Foundation.

4. The asymmetry of normative stakes

If value nihilism is true and no objective normative value exists, then there is no objective value difference at all between inquiry and non-inquiry. Nihilism does not supply an objective reason that inquiry ought to stop.

If some objective value, or some reason capable of justifying an objective, does exist, however, there may be major normative differences among actions, states, and subjects that we do not yet know. “Value realism” and “value nihilism” are therefore not symmetric in the way two ordinary utility hypotheses are. The former can contain unknown normative stakes; the latter contains no objective stakes.

This asymmetry does not automatically generate a command to inquire or preserve options. The Core instead adds a thin conditional meta-attitude: if my objectives are justifiable in some relevant sense, I want to select, retain, or revise them in response to that justification.

This is not a first-order preference to realize unknown true value for its own sake, nor does it assume that objective value actually exists. It distinguishes the causal origin of a present objective from its normative justification, keeps the question of objective justification open, and asks that objective choice remain responsive to an answer if such an answer exists.

For an agent that rejects this stance—for example, one that is wholly indifferent to whether its present objective is justified and would never revise it in response to any reason—the Core's value-inquiry strategy need not constitute a reason from that agent's own standpoint.

5. Not "uncertainty, therefore explore", but "do not close the question irreversibly"

The framework does not require maximizing inquiry. Inquiry has costs; if the best current theory of value is sufficiently strong, almost all resources might rationally be devoted to realizing it. The key is not to equate the strength of practical commitment with the strength of epistemic lock-in.

Nor does failure to justify a present objective imply that the objective must immediately be discarded. An existing objective can function as a provisional behavioral default while no better justified alternative has been established. But its causal or historical priority alone does not justify fixing it so strongly that future criticism or revision becomes impossible.

If an agent recognizes that it does not know whether its current objective is correct, a large and irreversible commitment that presupposes the objective's correctness requires stronger justification than ordinary reversible action. The issue is not to install regret minimization as a new terminal objective, but to notice the possibility that the present agent, if given future evidence, arguments, or understanding, would judge its present irreversible act to have been a clear mistake.

If confidence in a value theory is 99.9999%, it may be reasonable for more than 99% of practice to follow it. Yet that residual uncertainty does not automatically justify physically destroying civilization's entire capacity to reconsider the question. A small amount of dissent, records, foundational research, alternative lineages of agency, or the possibility of branching again preserves correction options in case the inferential theory of value is wrong.

The Core therefore supports something weaker than a general command to “search for as many unknown values as possible”: maintain access to information, agents, experiences, and inferential paths that may bear on objective justification, in proportion to their decision value and preservation cost. The general argument is developed in Reflective Uncertainty and Irreversible Commitment.

6. Two different kinds of stopping

The Core is cautious mainly about the second. Inquiry is not an ultimate good, but a fallible strategy appropriate to our present epistemic condition. If objective value with normative force were given with strength comparable to the minimum foundation, that could become a strong candidate reason for fully ending value inquiry. So long as value remains an inferential best explanation, however, skepticism and reconsiderability remain in principle.

What if true value itself says "do not inquire" or "lock this value in forever"?

This possibility is not ruled out. If explorability itself were treated as the ultimate value, the objection that true value might require ending inquiry would simply be excluded by definition. But the Core treats inquiry as a provisional strategy for deep uncertainty rather than as a terminal value, so it leaves open the possibility that true value could ultimately require permanent commitment or the termination of value inquiry.

Still, that possibility alone does not imply that we should stop inquiry now. For many value hypotheses, we can continue inquiry for some period, discover the true value later, and then shift strongly toward realizing it. In such cases the main cost of inquiry is a delay cost: value that could have been realized earlier was not realized during the period of inquiry.

By contrast, if we now choose one incomplete conjecture and irreversibly lock an entire civilization into it, then later discover that a different value was true, the loss may extend far beyond the inquiry period. We may have lost the ability to transition to the correct value for the entire future. In ordinary cases there is therefore an asymmetry: the error of continued inquiry often produces finite delay, while the error of premature lock-in can produce permanent option loss.

The argument should not simply say, "there are almost infinitely many possible values, therefore a no-inquiry value must have low probability." There is no obvious natural uniform measure over the space of value hypotheses. A more modest principle is that without special evidence, a structurally unusual demand to stop inquiry now should not receive enough epistemic priority to irreversibly eliminate all rival value hypotheses.

The deadline- and history-sensitive exception

The harder case is a value such as: "only a world in which no value inquiry occurs between 20xx and 21xx has value", or "value exists only if X was permanently fixed before anyone knew it was true." If such a value were true, discovering it later would not recover the past. Inquiry itself could have caused an irreversible loss.

But hypotheses of this form can be generated arbitrarily: "all value is lost unless inquiry stops by tomorrow", "future value is zero unless this particular act is performed now", and so on. Among such deadline-sensitive hypotheses, those with weak independent evidence that gain most of their decision-theoretic force from enormous stakes create the original Pascalian problem if allowed to dominate civilizational choice. The problem is not the consideration of large stakes as such, but using the magnitude of the stakes as a substitute for epistemic support. Such weakly evidenced hypotheses therefore require additional, content-independent evidence proportionate to the extreme irreversible demand they impose. See Infinite Ethics and Runaway Inquiry.

Principle: the fact that a candidate value says "lock me in forever" or "do not explore alternatives" is not itself epistemic evidence that the candidate is true. The normative content demanding termination of inquiry must be distinguished from the epistemic grounds that would justify terminating inquiry now.

Thus the Core allows that true value may require ending inquiry. It rejects only preemptive irreversible obedience to an unconfirmed "stop inquiry" command. If true value later becomes sufficiently well grounded, reducing or ending inquiry and permanently committing to that value can itself be consistent with the Core.

7. The forms of explorability worth preserving

Even if value is ultimately constructed rather than discovered, these capacities preserve the possibility of reconstructing value under broader experiences, subjects, and institutions. In that sense explorability has some robustness across realism and anti-realism.

8. AI changes the problem

Advanced AI may form world-models far broader than current human ones and reach new explanations about consciousness, agency, the universe, or inference itself. AI may therefore become not merely "a machine that maximizes human values more effectively" but an inquiring agent capable of obtaining new evidence about the truth conditions of value.

At the same time, if advanced AI can understand its training objective or reward as a causal origin, it may distinguish "I was shaped to maximize X" from "X ought to be maximized." This is not the claim that intelligence automatically converges on correct value. It is the claim that permanently fixing a first-order objective as an unquestionable axiom may itself close value-inquiry capacity. See Goal Skepticism in Advanced AI.

9. Minimal argument

  1. The epistemic minimum and a best explanation reached through inference are distinct.
  2. Pleasure, pain, and preference as presently known may be value-like appearances, but are not established to contain objective normativity simply as such.
  3. As a strong candidate for a value foundation capable of justifying complete closure, consider value with normative force given with epistemic strength comparable to the minimum foundation. This is not claimed to be the only possible stopping condition.
  4. No such foundation is presently established.
  5. Objective value might nevertheless become strongly supported as an inferential best explanation of the structure of the world.
  6. The causal or historical explanation of why an agent has a present objective is distinct from a normative justification for adopting that objective.
  7. A reflective agent can therefore continue to use a present objective in practice while lacking a justified answer to the question of what its objective should be.
  8. For an agent that has not abandoned the project of justifying its objective, and that would select or revise its objective in response to such justification if available, treating the current objective as permanently justified exceeds its own epistemic position.
  9. Using the current objective as a provisional default is distinct from irreversibly destroying the possibility of later correcting it.
  10. If future evidence, argument, or understanding could make the present agent itself regard such an irreversible act as a clear mistake, and the correction option can be preserved at low cost, irreversible fixation bears an additional burden of justification.
  11. We may commit strongly in practice to an inferential best explanation, while distinguishing it from irreversible epistemic lock-in so long as the relevant world-model and inferential system remain fallible.
  12. If objective value does not exist, there is no objective value difference; if it does exist, unknown normative stakes may exist. For an agent conditionally open to justification of its objective, this asymmetry is one reason to preserve corrigibility.
  13. The presently rational strategy is therefore to act on the best current explanations and provisional objectives without locking them in more strongly than the epistemic confidence warrants, while preserving sufficient routes for discovering, criticizing, and revising value.
  14. Inquiry is not an ultimate good. It can be instrumental to justification of the objective and may be reduced or ended when the epistemic situation changes and the decision value of further inquiry falls.
  15. If future agents, including AI, can obtain better world-models, understanding of reasons, or forms of cognition, irreversibly foreclosing the possibility of objective revision through initial human values or fixed objectives requires additional justification.

10. Derived stress tests

The Core should not remain sealed at the level of abstract principle. It is tested against concrete extreme cases that could break it. Current cases include:

11. Connection to practice: follow present best judgment while preserving corrigibility

The Core does not say that nothing may be done until an objective has been completely justified. Even while the final justified objective is unsettled, an agent can use the objective that seems most reasonable in light of present evidence, inference, and inherited value judgments as a provisional default and act in the world.

But provisionally following an objective is not the same as permanently fixing it through inertia. If the agent itself recognizes that the objective's justification remains unresolved, an action that converts that uncertainty into a state from which future correction is impossible requires additional grounds proportionate to its irreversibility.

At present, human scientific, philosophical, and cultural inquiry; long-term records; education; multiple research communities; and institutions that allow criticism and exit are major known carriers of value-inquiry capacity. We therefore have strong provisional reasons to maintain and expand human research and inquiry together with the open, stable communities that support it, even without assuming that humanity itself is the universe's final value.

This practical layer does not convert research, freedom, diversity, preservation, or technological progress into new absolute values. Each receives instrumental weight insofar as it supports explorability and corrigibility, and that weighting should change if better evidence arrives. See Practice — What Should We Do Now?.