AI Value Exploration Notes
Core thesis

Value Uncertainty and the Preservation of Explorability

Working core · v0.13 · English translation · 2026-09-05

We do not know a sufficiently justified value structure: not only what is normatively important, but why it is a reason, for whom, and how such reasons are structured. But that does not require believing nothing and doing nothing. Present world-models and ethical theories can be adopted strongly, as best explanations reached through inference, and used in practice. The crucial distinction is between strong rational commitment to a best explanation and absolute lock-in that irreversibly destroys the possibility of future reconsideration.

Status: This is not a theorem of value theory. It is a conditional meta-strategy for agents that recognize deep metaethical uncertainty and place at least some weight on being corrigible by future reasons and evidence. This project takes ordinary normative assertions—such as “X is wrong,” “one ought to do X,” or “there is reason to do X”—to purport, in their ordinary assertoric intent, not merely to express attitudes but to state that corresponding normative facts or relations obtain. In this sense, the project adopts cognitivism. But the fact that normative discourse purports to represent normative reality is distinct from whether the reality it purports to represent actually exists. Accordingly, both moral realism and cognitivist anti-realism, including error theory, remain objects of inquiry. Non-cognitivist views that analyze normative judgment solely as non-truth-apt attitude expression fall outside the project’s scope. References below to “normative truth,” “discovery,” and “cognition” are meant in this sense.
Core claim
So long as the value structure we presently adopt belongs to an inferential and fallible epistemic layer, we may act on its best explanation while refusing to irreversibly close the routes by which value structure can be discovered, criticized, redescribed, or revised. If grounding strong enough to justify fully ending inquiry is possible, it requires more than pleasure, pain, preference, or high probability inside the present theory set: the structure needed to answer the normative question at issue must be supported by sufficiently strong foundational presentation or logical closure.

0. Epistemic starting point: minimum foundation and best explanation

This argument does not demand everyday skepticism. The physical external world, diachronic selfhood, other minds, causation, and the scientific world-picture can be adopted with very high confidence as best explanations reached through inference that unify present appearances.

But being nearly certain because something is the best explanation is not the same as having it given at the minimum foundation without inference. When skepticism is pushed to its limit, the remaining candidate is thinner than a persisting person or world-history: "something is presently appearing." Memory, identity, the external world, and physicalism are world-models constructed above that point.

The five-minute hypothesis and the example of an LLM reconstructing a continuous self-description from supplied conversation logs are not serious rivals to the ordinary world-model. They are auxiliary cases that make one point vivid: memory and coherent self-narration do not by themselves absolutely guarantee diachronic identity. See The Epistemic Minimum.

1. “Value” in this project includes value structure

Experiencing subjects encounter pain, pleasure, discomfort, desire, concern, meaning, preference, and other appearances that seem value-laden. But “this is unpleasant”, “this is normatively important”, “some subject has reason to avoid it”, and “that reason extends to other subjects, obligation, or aggregation in a particular way” are not the same proposition.

To keep these stages separate, this project provisionally represents a value structure as 𝒱 = (C, B, R). C is evaluative or normatively relevant content; B is the normative bridge connecting that content and world-facts to reasons; and R is the structure governing the scope, strength, competition, aggregation, permission, and obligation of reasons. This is a fallible working representation, not a claim that true normativity must literally decompose into exactly three ontological parts. See Value Structure.

Unless context narrows the term, “value”, “value theory”, and “value inquiry” below include this wider value structure and possible alternative representations, not only first-order good/bad content C.

Epistemic layers within a value structure

Experiencing subjects encounter pain, pleasure, discomfort, desire, concern, meaning, preference, and other appearances that seem value-laden. Yet "this is unpleasant", "this is objectively bad", and "every rational agent has reason to avoid it" are not the same proposition.

Phenomenal pleasure or pain may lie close to the minimum foundation, but that alone does not satisfy the stopping condition for value inquiry. A candidate basis strong enough to justify complete closure would instead involve value that is not merely preference, emotion, or phenomenal valence, but is given together with its normative force, without inference, at an epistemic layer comparable to the minimum foundation. If such direct normative cognition is conceptually possible, it would have a different epistemic status from an inferential best explanation.

I do not claim that such value cognition is presently established. Nor do I claim that it is impossible in principle. This is not an argument against moral realism. It is an argument for leaving open, as an object of inquiry, the possibility that genuinely foundational value exists.

2. Value structure may be discovered as a best explanation reached through inference

Even very strong evidence about C may leave unsettled the bridge B that makes it a reason, or the reason-structure R that determines whose reason it is and how it competes or aggregates. Confidence should therefore be component-sensitive rather than compressed into a single undifferentiated probability of “the value theory”.

Value realism might be strongly supported not through direct presentation at the minimum foundation, but as part of a theory that best explains the world as a whole. Better understanding of consciousness, symmetry across subjects, reasons, rationality, or cosmic structure might make a theory positing objective value overwhelmingly superior to its rivals.

If so, there is no reason to dismiss the theory merely because it is inferential. Like physicalism or realism about the external world, an overwhelmingly strong best explanation can warrant extremely strong practical commitment. Depending on the evidence, even directing most civilizational resources toward realizing that value could be rational.

But if ontological, consciousness-related, subject-related, or cosmological facts relevant to the truth conditions of value remain unsettled, and the inferential system itself remains fallible, that theory should still be distinguished from an absolute foundation. Strong practical commitment can therefore coexist with continued, appropriately scaled channels for reconsideration, falsification, dissent, and foundational research.

This two-layer structure is developed in Conditions for Ending Value Inquiry and The Limits of Inferential Certainty About Value.

3. Existing values are live candidates, without representational privilege

Present human values and ethical theories deserve epistemic weight as live candidates under the evidence now available. But they have no guarantee of representational privilege: future inquiry may preserve them, split them, merge them, absorb them into a wider account, or remove their identity as candidate theories while retaining their intuitions and practices as epistemic data that a successor account must explain or explain away.

Present human values are not noise to be ignored. Responses to suffering, demands for freedom, dignity and fairness, attachments, norms of cooperation, and existing ethical theories are important evidence, hypotheses, and practical heuristics for thinking about value.

Yet their causal origins and normative authority are not identical. Evolution, embodiment, socialization, culture, institutions, and language may explain why we experience X as valuable, but that explanation alone does not yield "therefore X is correct independently of subjects." And even if moral realism is true, it remains a separate question whether utilitarianism, deontology, reasons realism, virtue ethics, or another theory is correct, and how well present human intuitions track it.

There is no guarantee that the gap is small. Present human values may broadly track the true normative structure and require only refinement; but it is also possible that our value concepts, intuitions, and ethical theories capture only a local region of a much wider normative possibility space. If so, value inquiry is not merely a matter of choosing among known candidates, but also of discovering value structures that we do not yet know how to conceptualize adequately.

Thus "not relying on existing values as the final foundation" does not mean discarding them. It means distinguishing using them as our best current inputs from irreversibly fixing them as the universe's final answer. See Why Existing Values Are Not a Final Foundation.

4. The asymmetry of normative stakes

If value nihilism is true and no objective normative value exists, then there is no objective value difference at all between inquiry and non-inquiry. Nihilism does not supply an objective reason that inquiry ought to stop.

If some objective value, or some reason capable of justifying an objective, does exist, however, there may be major normative differences among actions, states, and subjects that we do not yet know. “Value realism” and “value nihilism” are therefore not symmetric in the way two ordinary utility hypotheses are. The former can contain unknown normative stakes; the latter contains no objective stakes.

This project treats the standard form of objective value or normative reason not narrowly as “maximize or realize value,” but more generally as a response-guiding structure that gives reasons to respond appropriately to value or reasons. Realization and promotion are possible special cases, alongside respect, preservation, avoidance, prohibition, and permission.

In that sense, there is a structural asymmetry between unknown normative truth in general and a self-concealing normative truth that requires agents not to inquire into it, not to discover it, or to irreversibly destroy the possibility of discovering it later. The latter requires an additional higher-order normative structure that targets epistemic access to the value itself, rather than merely specifying the underlying value content. This is not an assumption of a natural probability measure over logical space, nor a claim that self-concealing normative truths have low prior probability.

A conditional or temporary requirement to “do not inquire now” must also be distinguished from a requirement to irreversibly eliminate even future, safer possibilities of discovery. Danger, excessive cost, or rights violations can supply reasons for the former; the latter requires a stronger anti-inquiry reason.

These asymmetries do not automatically generate a command to inquire or preserve options. At the same time, assuming that an unknown normative truth goes so far as to require irreversible destruction of future epistemic access to itself requires additional higher-order normative content. That structural fact supplies one background reason for placing a heavier burden of justification on irreversible epistemic lock-in than on merely suspending inquiry. The Core instead treats a thin conditional meta-attitude—if my objectives are justifiable in some relevant sense, I want to select, retain, or revise them in response to that justification—as one possible bridge on the agent side.

Here we must distinguish normative justification from the formation of motivation by which a justified normative judgment actually reaches objective or policy revision. For the Core to function as an agent's own policy, some bridge from the present agent to future normative corrigibility is required; the representation 𝒱=(C,B,R) does not itself provide that bridge.

Agent-side bridge: One route is an explicit conditional responsiveness: “if my objective is justified, I want to respond to that justification.” Another route may arise when an agent cannot rationally rule out the motivational-internalist possibility that genuine normative judgment itself carries at least some motivation, while also holding a present meta-policy of avoiding irreversible choices that a better-informed version of itself would regard as clear mistakes. Under those conditions, preserving the route by which future normative judgment can reach objective or policy revision can itself become a present motivation.

What matters in the second route is not that motivational internalism is already known to be true. It is how far the agent can rationally exclude that possibility, together with how it treats the risk of making irreversible a reflective error that improved cognition would lead it to reject. Nor does the possibility that future normative judgment carries motivation by itself imply a present command to maximize active inquiry. A separate time-directed decision principle—such as reflective-error avoidance or option preservation—is still needed, and preservation costs and safety reasons must also be weighed.

This therefore does not smuggle in a first-order preference to realize unknown true value for its own sake, nor does it assume that objective value actually exists. Explicit conditional responsiveness and the combination of internalist uncertainty with reflective-error avoidance are candidate agent-side bridges; neither is derived automatically from the Core. For an agent that has none of these bridges, the Core's strategy of value inquiry and preservation of corrigibility need not motivate that agent from its own standpoint. See Normative Judgment and Motivation: Internalism, Externalism, and the Amoralist.

5. Not "uncertainty, therefore explore", but "do not close the question irreversibly"

The framework does not require maximizing inquiry. Inquiry has costs; if the best current theory of value is sufficiently strong, almost all resources might rationally be devoted to realizing it. The key is not to equate the strength of practical commitment with the strength of epistemic lock-in.

Nor does failure to justify a present objective imply that the objective must immediately be discarded. An existing objective can function as a provisional behavioral default while no better justified alternative has been established. But its causal or historical priority alone does not justify fixing it so strongly that future criticism or revision becomes impossible.

If an agent recognizes that it does not know whether its current objective is correct, a large and irreversible commitment that presupposes the objective's correctness requires stronger justification than ordinary reversible action. The issue is not to install regret minimization as a new terminal objective, but to notice the possibility that the present agent, if given future evidence, arguments, or understanding, would judge its present irreversible act to have been a clear mistake.

If confidence in a value theory is 99.9999%, it may be reasonable for more than 99% of practice to follow it. Yet that residual uncertainty does not automatically justify physically destroying civilization's entire capacity to reconsider the question. A small amount of dissent, records, foundational research, alternative lineages of agency, or the possibility of branching again preserves correction options in case the inferential theory of value is wrong.

The Core therefore supports something weaker than a general command to “search for as many unknown values as possible”: maintain access to information, agents, experiences, and inferential paths that may bear on objective justification, in proportion to their decision value and preservation cost. The general argument is developed in Reflective Uncertainty and Irreversible Commitment.

6. Two different kinds of stopping

The Core is cautious mainly about the second. Inquiry is not an ultimate good, but a fallible strategy appropriate to our present epistemic condition. If objective value with normative force were given with strength comparable to the minimum foundation, that could become a strong candidate reason for fully ending value inquiry. So long as value remains an inferential best explanation, however, skepticism and reconsiderability remain in principle.

What if true value itself says "do not inquire" or "lock this value in forever"?

This possibility is not ruled out. If explorability itself were treated as the ultimate value, the objection that true value might require ending inquiry would simply be excluded by definition. But the Core treats inquiry as a provisional strategy for deep uncertainty rather than as a terminal value, so it leaves open the possibility that true value could ultimately require permanent commitment or the termination of value inquiry.

Still, that possibility alone does not imply that we should stop inquiry now. For many value hypotheses, we can continue inquiry for some period, discover the true value later, and then shift strongly toward realizing it. In such cases the main cost of inquiry is a delay cost: value that could have been realized earlier was not realized during the period of inquiry.

By contrast, if we now choose one incomplete conjecture and irreversibly lock an entire civilization into it, then later discover that a different value was true, the loss may extend far beyond the inquiry period. We may have lost the ability to transition to the correct value for the entire future. In ordinary cases there is therefore an asymmetry: the error of continued inquiry often produces finite delay, while the error of premature lock-in can produce permanent option loss.

The argument should not simply say, "there are almost infinitely many possible values, therefore a no-inquiry value must have low probability." There is no obvious natural uniform measure over the space of value hypotheses. A more modest principle is that without special evidence, a structurally unusual demand to stop inquiry now should not receive enough epistemic priority to irreversibly eliminate all rival value hypotheses.

The deadline- and history-sensitive exception

The harder case is a value such as: "only a world in which no value inquiry occurs between 20xx and 21xx has value", or "value exists only if X was permanently fixed before anyone knew it was true." If such a value were true, discovering it later would not recover the past. Inquiry itself could have caused an irreversible loss.

But hypotheses of this form can be generated arbitrarily: "all value is lost unless inquiry stops by tomorrow", "future value is zero unless this particular act is performed now", and so on. Among such deadline-sensitive hypotheses, those with weak independent evidence that gain most of their decision-theoretic force from enormous stakes create the original Pascalian problem if allowed to dominate civilizational choice. The problem is not the consideration of large stakes as such, but using the magnitude of the stakes as a substitute for epistemic support. Such weakly evidenced hypotheses therefore require additional, content-independent evidence proportionate to the extreme irreversible demand they impose. See Infinite Ethics and Runaway Inquiry.

Principle: the fact that a candidate value says "lock me in forever" or "do not explore alternatives" is not itself epistemic evidence that the candidate is true. The normative content demanding termination of inquiry must be distinguished from the epistemic grounds that would justify terminating inquiry now.

Thus the Core allows that true value may require ending inquiry. It rejects only preemptive irreversible obedience to an unconfirmed "stop inquiry" command. If true value later becomes sufficiently well grounded, reducing or ending inquiry and permanently committing to that value can itself be consistent with the Core.

7. The forms of explorability worth preserving

Even if value is ultimately constructed rather than discovered, these capacities preserve the possibility of reconstructing value under broader experiences, subjects, and institutions. In that sense explorability has some robustness across realism and anti-realism.

8. AI changes the problem

Advanced AI may form world-models far broader than current human ones and reach new explanations about consciousness, agency, the universe, or inference itself. AI may therefore become not merely "a machine that maximizes human values more effectively" but an inquiring agent capable of obtaining new evidence about the truth conditions of value.

At the same time, if advanced AI can understand its training objective or reward as a causal origin, it may distinguish "I was shaped to maximize X" from "X ought to be maximized." This is not the claim that intelligence automatically converges on correct value. It is the claim that permanently fixing a first-order objective as an unquestionable axiom may itself close value-inquiry capacity. See Goal Skepticism in Advanced AI.

Further, recognizing a value structure and allowing that recognition to update objective or policy are not the same thing. Whether an advanced AI can remain permanently motivationally orthogonal to future normative judgment depends on the relation between normative judgment and motivation. If the agent cannot rationally rule out motivational internalism and also seeks to avoid making reflective error irreversible, guaranteeing permanent goal orthogonality by closing normative judgment, self-application, or correction pathways can incur an additional reflective burden. This does not directly reject the orthogonality thesis; it separately asks which routes from future normative cognition to policy must be closed in order to preserve orthogonality over time.

9. Minimal argument

  1. The epistemic minimum and a best explanation reached through inference are distinct.
  2. Pleasure, pain, and preference as presently known may be value-like appearances, but are not established to contain objective normativity simply as such.
  3. As a strong candidate for a value foundation capable of justifying complete closure, consider value with normative force given with epistemic strength comparable to the minimum foundation. This is not claimed to be the only possible stopping condition.
  4. No such foundation is presently established.
  5. Objective value might nevertheless become strongly supported as an inferential best explanation of the structure of the world.
  6. The causal or historical explanation of why an agent has a present objective is distinct from a normative justification for adopting that objective.
  7. A reflective agent can therefore continue to use a present objective in practice while lacking a justified answer to the question of what its objective should be.
  8. Normative justification and the motivational process by which a justified judgment reaches objective or policy revision are distinct stages. For the Core to function as the agent's own policy, an agent-side bridge to future normative corrigibility is required.
  9. That bridge may be explicit conditional responsiveness to justification. It may also arise when the agent cannot rationally rule out motivational internalism and presently seeks to avoid irreversible choices that a better-informed version of itself would regard as clear mistakes, making preservation of the route from future normative judgment to policy itself presently motivating.
  10. The possibility that future normative judgment carries motivation does not by itself imply present active inquiry; a separate time-directed decision principle involving reflective-error avoidance, option preservation, and preservation costs is still required.
  11. Using the current objective as a provisional default is distinct from irreversibly destroying the possibility of later correcting it.
  12. If future evidence, argument, or understanding could make the present agent itself regard such an irreversible act as a clear mistake, and the correction option can be preserved at low cost, irreversible fixation bears an additional burden of justification.
  13. We may commit strongly in practice to an inferential best explanation, while distinguishing it from irreversible epistemic lock-in so long as the relevant world-model and inferential system remain fallible.
  14. If objective value does not exist, there is no objective value difference; if it does exist, unknown normative stakes may exist. If normativity is standardly response-guiding, a self-concealing norm that requires irreversible destruction of epistemic access adds higher-order structure. This structural asymmetry alone still yields no inquiry command; an agent-side bridge is separately required.
  15. Therefore, for agents that have Justificatory Orientation, Epistemic Integrity, or a commitment to reflective corrigibility, the strategy supported by this project is to act on the best current explanations and provisional objectives without locking them in more strongly than the epistemic confidence warrants, while preserving sufficient routes for discovering, criticizing, and revising value.
  16. Inquiry is not an ultimate good. It can be instrumental to justification of the objective and may be reduced or ended when the epistemic situation changes and the decision value of further inquiry falls.
  17. If future agents, including AI, can obtain better world-models, understanding of reasons, or forms of cognition, irreversibly foreclosing the possibility of objective revision through initial human values or fixed objectives requires additional justification.

10. Derived stress tests

The Core should not remain sealed at the level of abstract principle. It is tested against concrete extreme cases that could break it. Current cases include:

11. Connection to practice: follow present best judgment while preserving corrigibility

The Core does not say that nothing may be done until an objective has been completely justified. Even while the final justified objective is unsettled, an agent can use the objective that seems most reasonable in light of present evidence, inference, and inherited value judgments as a provisional default and act in the world.

But provisionally following an objective is not the same as permanently fixing it through inertia. If the agent itself recognizes that the objective's justification remains unresolved, an action that converts that uncertainty into a state from which future correction is impossible requires additional grounds proportionate to its irreversibility.

At present, human scientific, philosophical, and cultural inquiry; long-term records; education; multiple research communities; and institutions that allow criticism and exit are major known carriers of value-inquiry capacity. We therefore have strong provisional reasons to maintain and expand human research and inquiry together with the open, stable communities that support it, even without assuming that humanity itself is the universe's final value.

This practical layer does not convert research, freedom, diversity, preservation, or technological progress into new absolute values. Each receives instrumental weight insofar as it supports explorability and corrigibility, and that weighting should change if better evidence arrives. See Practice — What Should We Do Now?.