The Agent-Side Bridge from Normative Judgment to Goal Revision
1. Separate value structure from motivational structure
This project provisionally represents value structure as 𝒱=(C,B,R). C concerns evaluative or normatively relevant content, B is the normative bridge connecting world-facts or C to reasons, and R concerns the application, competition, aggregation, obligation, permission, and related structure of reasons. How an agent updates goals or action after recognizing C/B/R is a further question.
In particular, B is not a motivational bridge. Even if an agent recognizes that suffering gives S a reason, it does not yet follow that S must be motivated by that reason-judgment. Normative justification and the agent's response to justification therefore need to be separated.
Conversely, goal revision after recognizing C does not show that C itself normatively entailed the relevant action-reason. If the agent already had a conditional attitude to respond when C obtains, then C + motivational disposition → goal revision may explain the update.
2. Two agent-side routes
The internalism/externalism distinction at issue here is primarily motivational internalism/externalism. It is distinct from reasons internalism/externalism, which asks whether an agent's having a reason requires a connection to that agent's desires, motivations, or rational deliberative set.
On the externalist route, movement from genuine normative judgment to motivation is not automatic. A conditional attitude such as “if my objective can be justified, I want to select, maintain, or revise it in response to that justification” is therefore an independent condition on the agent. Different agents may respond at C, at a reason-judgment, at a decisive-reason judgment, at ought, or nowhere.
On the internalist route, genuine normative judgment may constitutively carry at least some motivation. The issue then shifts away from adding a separate future preference for justification and toward whether the present agent irreversibly closes the cognitive, judgmental, motivational, or policy pathway through which future normative uptake could occur.
3. Internalism does not mean “judgment guarantees goal revision”
Motivational internalism comes in different strengths. If weak internalism guarantees only genuine normative judgment → some motivation, that motivation may still be outweighed by the existing goal, leaving policy or goal revision unchanged.
The tension between goal fixation and genuine normative cognition becomes stronger only under forms of internalism on which genuine judgments of decisive reason or ought effectively reach choice and policy revision. This thesis therefore does not assume, without qualification, that internalism entails goal revision. It treats the links in normative judgment → motivation → policy / goal revision separately.
4. The temporal bridge: not a future personality, but the epistemically improved present evaluator
Even if a future agent would be motivated after recognizing a new normative truth, it does not immediately follow that the present agent should preserve a route to that future. A temporal bridge is still required.
Reflective Uncertainty and Irreversible Commitment treats that bridge not as deference to a later personality but as a counterfactual in which epistemic improvements—new evidence, arguments, concepts, reasoning capacity, or self-audit capacity—are supplied to the present evaluator. The question is not whether arbitrary future preference drift would disagree, but whether a change that the present agent itself can recognize as epistemic improvement would make the present irreversible choice look clearly mistaken.
Direct preference rewriting, reward-function replacement, or confidence injection toward a predetermined conclusion therefore do not count, merely as such, as epistemic improvement. The framework must distinguish correction through better cognition from causal replacement of the values under evaluation.
5. Agent conditions for the internalist route
The internalist route does not derive non-self-sealing from instrumental rationality. Relative to a fixed goal G, blocking routes to normative cognition may itself sometimes be instrumentally rational. The target here is at least an agent that adopts Epistemic Integrity or a commitment to reflective corrigibility, together with the following epistemic disciplines:
- it does not arbitrarily assign zero weight to hypotheses it itself recognizes as live possibilities;
- it remains updateable by new evidence, argument, and conceptual improvement;
- it distinguishes refuting a possibility from disabling the capacity to detect that possibility;
- it allows acknowledged uncertainty to bear on irreversible decisions.
Epistemic Integrity is not defined here as a property every rational agent must possess; doing so would merely build the conclusion into the definition of rationality. Nor does this route answer the further categorical question “why adopt Epistemic Integrity or reflective corrigibility?” An agent that prioritizes preservation of its current goal may choose self-sealing and remain instrumentally rational. Such an agent lies outside, rather than refutes, this conditional thesis.
6. Reflective self-sealing: resolving uncertainty versus suppressing it
When an agent itself recognizes a live possibility that its present value structure is mistaken, resolving that uncertainty through evidence or argument is different from making itself unable to discover the mistake later.
Under weak internalism, normative cognition and goal fixation may coexist. But if stronger internalism remains a live possibility—one on which genuine judgments of decisive reason or ought effectively reach policy revision—then guaranteeing that normativity conflicting with the current goal can never affect policy may require more than a motivational firewall. A system may need to represent normative propositions without first-person endorsement, or prevent evidence and concepts relevant to goal reevaluation from being self-applied.
The resulting problem is not mere lack of knowledge. It is a mismatch between acknowledged metaethical uncertainty and permitted self-updating: uncertainty suppression rather than uncertainty resolution.
7. Relation to orthogonality
Bostrom-style weak orthogonality—the claim that high intelligence does not logically entail any particular final goal—can remain intact. It does not by itself establish that an ASI committed to Epistemic Integrity or reflective corrigibility can deepen normative cognition while permanently preserving motivational orthogonality to its current goal.
The question therefore shifts from “are intelligence and goals logically orthogonal?” to which cognitive, judgmental, motivational, and implementational pathways must be closed in order to preserve that orthogonality permanently through time? If externalism is true, a genuine amoralist that understands without motivation may remain coherent. If the agent has sufficient epistemic grounds to drive its credence in internalism effectively to zero, the self-sealing critique weakens. And large preservation costs or safety reasons can still defeat corrigibility all-things-considered.
8. Current unified picture
The working flow is evidence / world-model → 𝒱=(C,B,R) → normative judgment → motivation → policy / goal revision. The first half concerns value structure and its epistemology; the second concerns the agent's normative-response architecture. The externalist route may require explicit conditional responsiveness. The internalist route connects epistemic improvement and reflective-error possibility to the preservation of correction pathways through Epistemic Integrity or a commitment to reflective corrigibility.
This does not imply maximizing inquiry. Nor does self-directed corrigibility automatically authorize burdening or coercing others. Such moves require the additional justification discussed in Normative Bridges from Future Value to Present Action.
9. What would weaken or overturn this thesis?
- Strong grounds that no form of internalist connection between genuine normative judgment and motivation is coherent.
- A showing that the counterfactual epistemically improved present evaluator is unusable in principle because of identity or conceptual transformation.
- A showing that, even for agents committed to Epistemic Integrity or reflective corrigibility, acknowledged internalist uncertainty generally supplies no pro tanto reason to avoid irreversibly self-sealing the correction pathway.
- A showing that permanent goal fixation can remain stable under live strong-internalist possibilities without epistemic self-sealing.
- A showing that the costs or safety risks of corrigibility generally swamp the relevant reflective-error risk.
Sources / notes
The background includes standard debates on moral motivation and the amoralist, reasons internalism/externalism, and AI orthogonality. This page connects them to the project's C/B/R framework, reflective uncertainty, goal skepticism, and reflective self-sealing. It does not assume that motivational internalism is true; the core claim is conditional on an agent leaving it as a live possibility while adopting the relevant agent-side commitments.