Decision Policy — Acting Under Incomplete and Revisable Values
1. Scope — choice rather than belief
Epistemic Policy concerns what to believe, how strongly to believe it, and which world models, hypothesis spaces, and epistemic methods to adopt. This page concerns which action, policy, or allocation to choose given that epistemic state and the agent's current value state. Value theory asks what counts as a value or reason; decision theory asks how to act from such inputs when they are incomplete.
Decision theory should therefore not be used as a substitute for value discovery. The fact that a preference can be represented by an expected-utility function does not establish that the preference is normatively correct. Likewise, internal coherence of a decision rule does not make the rule an ultimate normative principle.
This page is also conditional on the agent having an agent-side bridge of the kind used in the Core—for example, a conditional disposition to revise goals or policy if they become better justified, or a present concern not to make irreversible choices that a better-informed version of itself would regard as a clear mistake. Value uncertainty by itself does not automatically generate motivation toward inquiry, preservation, or corrigibility. Claims below that E, P, or optionality can supply present reasons should be read under this scope condition.
epistemic state + value state + feasible actions + constraints → decision policy → action. Uncertainty does not imply practical paralysis.2. Baseline case — do not reject expected utility
If states s, actions a, subjective probabilities P(s), and a sufficiently cardinal utility U(a,s) are available, the familiar rule EU(a)=Σs P(s)U(a,s) is a powerful decision criterion. Representation theorems in the von Neumann–Morgenstern and Savage traditions clarify conditions under which preferences can be represented as expected-utility maximization.
The issue is not “EU or anti-EU.” If probabilities and value scales are sufficiently trustworthy and differences among outcomes are commensurable on the relevant scale, there is no reason to avoid expected utility or ordinary optimization. The real question is whether the agent is entitled to use that numerical representation.
The standard model presupposes an adequately represented action space, state space, outcome mapping, probability model, value scale, and comparison or aggregation rule. The project is mainly concerned with cases in which some of these are unsettled.
sufficiently represented decision model → use the strongest justified optimization available.3. How much quantification is justified? — representational entitlement
Assigning numbers does not license arbitrary arithmetic. If only ordinal ranking is justified, A > B may be meaningful while U(A)-U(B) or cross-person averages are not. Interval structure, ratio structure, interpersonal comparability, and intertheoretic comparability each require additional assumptions.
Nor must rational preference always be complete. Some pairs may be genuinely unresolved rather than tied. Treating incomparability as indifference creates trade-offs that the current value model does not support.
Value representation may therefore be scalar, interval- or set-valued, multi-attribute, ordinal, partially ordered, or explicitly incomparable. Even an additive multiattribute form such as V(a)=Σi wi vi(a) requires substantive independence and trade-off structure rather than following automatically from “multiple values exist.”
4. Represented value uncertainty — what if several theories are live?
Even when candidate value theories are already represented, simple expectation across them may be ill-defined. Credences in T1 and T2 do not suffice for P(T1)U1(a)+P(T2)U2(a) if the theories' choiceworthiness scales have independently arbitrary units. This is the intertheoretic comparability problem in the moral-uncertainty literature.
Where comparison is sufficiently grounded, expected choiceworthiness may be used. Moral-uncertainty work also includes social-choice analogies and ordinal aggregation for cases with weaker cardinal comparability. Separately, imprecise decision theory can retain sets of probability or utility representations and apply different rules such as E-admissibility or maximality. In MCDA, Robust Ordinal Regression distinguishes necessary preference, which holds across every compatible model, from possible preference, which holds under at least one.
These are not one theory, but they share a useful lesson: the absence of a single precise scale does not erase all decision structure. Incomparability is information about the current representation, not merely a failure to decide. Use the comparable structure to narrow the choice set and use robustness, information acquisition, or reversibility for what remains.
uncertain theory ≠ automatically commensurable theory.5. Unrepresented value — do not give an imaginary weight to an unknown theory
Bayesian Updating and Open Hypothesis Spaces argues that a hypothesis not yet present in the hypothesis space need not be a zero-prior hypothesis; its probability may simply be undefined in the current representation. Value inquiry has the same problem. A not-yet-represented V* outside the current candidate set 𝒱t should not be treated as a fictitious “unknown theory party” with, say, five percent of the current practical vote.
For the agents within this page's scope, the possibility of unconceived value can affect E (search for new value candidates), P (preserve the capacity to reassess), caution around irreversible acts, and future optionality rather than adding a fictional utility term to X. Represented value uncertainty can directly inform first-order action comparison; unrepresented possibilities mainly constrain the openness of the decision architecture.
This response has a close precedent in Cordasco's 2026 account of evaluative flexibility and endogenous menus. If future evaluative outlooks may change, Cordasco argues that agents have reason not only to preserve current options but to support meta-rules that generate and filter future options. This project generalizes that idea from life-options to normative hypothesis spaces: P includes preserving the capacity to generate new value concepts, hypotheses, and representational forms—generative openness.
Thus preparation for unrepresented value is not merely preserve current options, but also preserve option / hypothesis generation capacity. For the broader comparison, see From Value Uncertainty to Normative Discovery.
unrepresented value possibility → exploration / preservation / generative openness / optionality, not → fictitious scalar utility.6. Representation constrains decision rules but does not uniquely determine one
The preceding sections do not motivate one universal rule. Expected utility is appropriate when probabilities and cardinal utilities are adequate; set-valued beliefs or utilities point toward families of imprecise or robust rules; partial comparison permits dominance or partial ordering; deep model uncertainty motivates RDM or regret-based tools; future learning motivates VOI or real-options reasoning; adaptation through time motivates DAPP; and revisable values require reflective policies.
But a representation does not automatically select one uniquely correct rule. Even with a set of admissible utilities, rules such as E-admissibility, maximality, and different dominance criteria may remain available. The representation therefore constrains an admissible rule family rather than producing a single decision rule.
Choosing within that family may depend on further decision criteria such as dominance, conservatism, regret, dynamic consistency, computational cost, or optionality. This second-stage choice is itself not treated here as finally settled.
representation → admissible rule family. Do not use a rule that requires operations the justified representation cannot support, but do not assume representation alone uniquely selects a rule.7. E / P / X — three functional dimensions, not exclusive action classes
Allocating Inquiry Under Unresolved Normative Uncertainty distinguishes E: active inquiry, P: preservation of inquiry options, and X: realization of present values. These are neither three terminal goods nor three mutually exclusive boxes into which every action must be placed. They are better understood as three decision-relevant functions that an action or institution may serve simultaneously.
a → { E(a), P(a), X(a) }. One action may contribute to all three.Maintaining a research university, for example, can produce inquiry now through E, preserve future inquiry through education, institutions, and archives via P, and realize present educational, productive, or welfare value through X. When allocating budgets or attention, the same distinctions can be represented more coarsely as a resource allocation (Et,Pt,Xt).
Justificatory relevance, value of information, irreversibility of inquiry opportunities, inquiry risk, present-value opportunity cost, and preservation cost can change the weight placed on each dimension. The E/P distinction especially permits policies that pursue X strongly, preserve P, and temporarily reduce E when active inquiry is dangerous or expensive.
8. Deep uncertainty — prefer robustness when the represented model itself is fragile
When uncertainty concerns not merely parameter values but which futures are plausible, how causal models operate, or which outcome metrics matter, the problem enters the domain of Decision Making under Deep Uncertainty. Robust Decision Making (RDM) does not optimize for one best-estimate future; it runs candidate policies across many plausible futures, discovers conditions under which they fail, and iteratively redesigns them to reduce those vulnerabilities.
RDM should not be reduced to worst-case maximin. The point is not to optimize against the most extreme imaginable state, but to seek acceptable performance across a wide range of plausible world and value representations, using satisficing, regret, multi-metric evaluation, and vulnerability discovery as needed.
RDM can directly stress-test only represented world models, value models, and performance criteria. It cannot run an unconceived V* ∉ 𝒱t as a scenario. Residual unrepresented possibilities must instead be addressed indirectly through P, reopenability, alternative research pathways, and preserved records.
represented deep uncertainty → stress-test / robustness; unrepresented possibility → preservation / reopenability.9. Should we seek more information? — VOI and the risk of inquiry
Active inquiry E is itself a decision. Value of Information (VOI) asks how much an agent should pay for information that may improve later choices, often by comparing the optimal expected value before and after receiving that information.
Value inquiry may not fit the standard structure directly. First, inquiry can generate a new value model or comparison axis rather than merely reveal which known state obtains. “Meta-VOI” is a useful analogy here, but it does not automatically create a common scale on which currently unconceived values can be priced. Second, inquiry itself may be hazardous, as with dangerous experiments or high-capability AI. Third, inquiry can destroy subjects, samples, raw data, or institutional conditions needed for future inquiry.
Information is therefore not monotonically valuable. For agents within this page's scope, decision relevance, expected epistemic improvement, inquiry cost, present-value opportunity cost, inquiry risk, and effects on future inquiry capacity all matter.
epistemic gain alone ≠ sufficient reason for active inquiry.10. Irreversibility and the value of waiting — from real options to reflective optionality
When an investment or commitment is costly to reverse, can be delayed, and waiting is expected to bring information, not acting now can preserve a valuable option. Real-options analysis treats future opportunities to invest, expand, contract, switch, or abandon as options whose loss is part of the opportunity cost of irreversible action.
For this project, the future evaluator may not share the current evaluator's value function. It is therefore useful first to define a structurally neutral concept: reflective optionality, the capacity of a future agent with more information, greater cognitive ability, or better value understanding to reassess and alter a present decision or value representation.
How much that reflective optionality is worth to the present agent is not automatic. An agent with the agent-side bridge assumed in section 1 may assign it derivative reflective option value because it preserves responsiveness to future, better-justified judgment.
reflective optionality + agent-side bridge → possible reflective option value. Future optionality is not itself assumed to be an unconditional terminal value.Cordasco's preference for flexibility is a close prudential argument for this kind of reflective option value. It derives an instrumental reason to preserve future options from uncertainty about one's future evaluative outlook. His peer-reviewed 2026 paper, however, focuses the flexible set on futures recognizable under present higher-order values and reachable through feasible uptake routes. This page instead defines reflective optionality first as a structural capacity and derives its present value through an agent-side bridge; present recognizability is therefore not treated as a criterion of normative admissibility.
Waiting is not free. Present value may be lost; technical opportunities may disappear; competitors may move first; and inaction may itself create irreversible losses. The relevant comparison is therefore between the irreversibility of acting and the irreversibility of waiting.
11. Choose a pathway, not only a one-shot act — DAPP
Dynamic Adaptive Policy Pathways (DAPP) addresses deep uncertainty by committing to near-term actions while designing pathways for switching policies as observations change. It uses adaptation tipping points—conditions under which a policy no longer meets specified objectives—together with signposts and triggers for moving to another action.
This structure directly combines strong present action with future corrigibility. Represent a policy as π: future evidence/state → future action: pursue X strongly now; increase E if anomaly A appears; strengthen P as irreversible threshold B approaches.
DAPP itself does not generate unconceived future values or objectives. An adaptation tipping point already presupposes some objective against which policy performance is judged. What this project mainly borrows is the architecture of monitoring, triggers, reassessment, and path switching, including preserving the procedure by which the policy itself can later be reconsidered.
12. Why reversibility must sometimes be designed early — Collingridge
Collingridge's control dilemma observes that early in a technology's development it is relatively easy to change but difficult to predict its consequences; later, consequences become clearer but sunk costs, skills, infrastructure, institutions, and dependence make change difficult. In shorthand, knowledge ↑ while changeability ↓ can move in opposite directions.
“Wait until we understand it, then make the irreversible decision” therefore fails as a general policy. The capacity to change may disappear while understanding is still developing. This creates a reason to build staged deployment, rollback, monitoring, branching, independent alternatives, and exit paths into systems before their long-run effects are fully known.
But the problem is not merely lack of factual information. If objectives or value structures themselves are unsettled, as in this project, more future knowledge alone need not uniquely select the policy. The system must preserve not only learning about consequences but also the capacity to reassess objectives.
The problem is especially acute for AI and value lock-in, where economic, political, and security dependence can make later correction far more costly than early redesign.
13. Reflective decision — choose not only actions, but channels of value revision
Ordinary dynamic-choice theory already asks whether a present agent should use commitment against predictable future preference change or instead defer to future judgment. This project goes further because Vt → Vt+1 may result from new normative evidence, direct affective modification, manipulation by another agent, cognitive enhancement, or acquisition of a new conceptual scheme.
Future preference is not automatically authoritative, but current preference is not automatically privileged either. Permanently binding future evaluators and surrendering to every future preference change each require additional justification.
The decision target is therefore not only which future value Vt+1 to choose now. The agent may also permit, preserve, or constrain a value-revision channel c ∈ 𝒞t through which values can change. Schematically, c: (Vt, Kt, At) → (Vt+1, Kt+1, At+1), where K represents cognitive or epistemic capacities and A the capacity for future reassessment and self-revision.
- Reason-mediated revision
- Value judgment changes through new evidence, argument, self-criticism, or reflection.
- Direct preference modification
- Desire, affect, reward systems, or preferences are directly rewritten.
- Externally induced modification
- Another agent's persuasion, training, manipulation, or environmental design changes the value state.
- Cognitive transformation
- Changes in intelligence, perception, conceptual scheme, or conscious architecture alter value representation or comparison capacity.
These causal labels do not by themselves settle which transformations count as legitimate learning and which count as corruption. Direct modification could sometimes implement a justified correction, while apparently reason-mediated processes could be systematic manipulation. Whether a transition is learning, corruption, manipulation, or value discovery is itself part of the unresolved reflective decision problem.
The meta-policy M from Reflective Uncertainty and Irreversible Commitment connects here. When justified objective remains unsettled, the agent must decide not only how to use present values provisionally but which value-revision channels to keep open, which to constrain, and which cognitive or self-modification capacities to preserve.
choose not only actions under values, but channels by which values may change, while current value use ≠ current value privilege ≠ future-value surrender.14. Extension to multiple agents — a boundary of single-agent Decision Policy
The core of this page concerns a single agent or one integrated decision process. Once humans, AIs, future agents, and distinct value-inquiry lineages coexist, additional layers arise from social choice, judgment aggregation, bargaining, mechanism design, and constitutional design.
This page retains only two boundary claims. First, preserving multiple agents does not automatically generate a correct collective value. Second, the fact that a collective procedure selects A does not automatically authorize coercion of dissenters into A.
plurality ≠ automatic aggregation, and collective choice ≠ coercive authorization.The distinction from Normative Bridges from Future Value to Present Action remains: reason ≠ permission ≠ obligation ≠ coercive authorization. Detailed design of aggregation, bargaining, exit, forks, and decentralization belongs to a separate theory of collective decision.
15. Overall structure — strong action, no false precision, preserved correction
The resulting decision policy is not a fixed algorithm. Current representation constrains an admissible rule family; further decision criteria help select a policy within it; and actions then change the evidence, option set, value state, revision channels, and sometimes the evaluator itself.
{ epistemic state, value state, comparability, feasible actions, constraints }t → admissible rule family → decision policyt → actiont → { new evidence, new options, new value state, new revision channels }t+1Optimize strongly where comparison is justified; do not force incomparable values into a fabricated scalar. Stress-test represented deep uncertainty; answer unrepresented possibilities through preservation and reopenability. Evaluate information when future learning matters; compare the irreversibility of acting and waiting; and build adaptive pathways through time.
When values themselves are revisable, the decision target expands from “what should I do under my current values?” to “which channels of future value revision should remain available?” Agents within this page's scope may use current values strongly while demanding additional reason before destroying the capacity for better future reassessment.
This decision policy is itself revisable, just like the Epistemic Policy. Better future comparison structures, decision theories, or cognitive capacities may replace it.
What this page does not claim
- That expected-utility theory is false. It remains central in adequately represented standard cases.
- That one representation uniquely determines one correct decision rule.
- That incomplete value must always be handled by maximin, minimax regret, or RDM.
- That incomparability is permanent; inquiry may discover new comparison structure.
- That RDM or DAPP can directly represent or optimize unconceived values.
- That reversibility, inquiry, optionality, or pluralism is itself an ultimate value.
- That irreversible action should always be postponed; inaction, competition, and expiring opportunities may also be irreversible.
- That future values automatically outrank present values, or vice versa.
- That any value-revision channel is legitimate or corrupt merely because of its causal form.
- That social plurality automatically produces correct collective decisions.
Selected references
- Stanford Encyclopedia of Philosophy, “Decision Theory” — expected utility, incomplete preference, imprecise probability and utility.
- SEP, “Normative Theories of Rational Choice: Expected Utility”.
- SEP, “Normative Theories of Rational Choice: Rivals to Expected Utility” — imprecise representations and multiple choice rules.
- SEP, “Moral Decision-Making Under Uncertainty” — expected choiceworthiness, intertheoretic comparison, and ordinal/social-choice approaches.
- Carlo Ludovico Cordasco, “Abstraction as Flexibility: The Veil of Evaluative Uncertainty” (2026) — evaluative uncertainty, preference for flexibility, and endogenous option menus.
- Ralph L. Keeney, “Utility Independence and Preferences for Multiattributed Consequences” (1971).
- Robert J. Lempert et al., RAND overview of Decision Making under Deep Uncertainty / Robust Decision Making.
- David Manheim, Value of Information for Policy Analysis (2018).
- Robert S. Pindyck, “Irreversibility, Uncertainty, and Investment” (1990).
- Marjolijn Haasnoot et al., “Dynamic Adaptive Policy Pathways” (2013).
- David Collingridge, The Social Control of Technology (1980).
- SEP, “Dynamic Choice”.
- SEP, “Social Choice Theory”.
Related pages
- Epistemic Policy — Best Explanation and Open-Ended Inquiry
- Value Structure — Content, Normative Bridges, and Reason Structure
- Allocating Inquiry Under Unresolved Normative Uncertainty
- Conditions for Ending Value Inquiry
- Reflective Uncertainty and Irreversible Commitment
- Normative Bridges from Future Value to Present Action
- From Value Uncertainty to Normative Discovery
- Bayesian Updating and Open Hypothesis Spaces
- Practice — What Should We Do Now?